Guide
Score Inputs
Quality score weights star movement, commits, releases, contributors, and issue response; agent score adds documentation, deployment, popularity, and community fit.
Git.Top Guide
Inspect quality scores, confidence, freshness, and risk signals for open-source project selection.
Guide
Quality score weights star movement, commits, releases, contributors, and issue response; agent score adds documentation, deployment, popularity, and community fit.
Guide
Use /api/quality, project-level quality_signal_confidence, and sync status together before treating a recommendation as high-confidence.
Data Source
1047 loaded projects. Generated at 2026-08-05T22:42:02.093Z.
Project
Evaluation package that allows benchmarking of agentic AIs from various sources and frameworks by producing statistical results which can be compared across different use cases and datasets.
Project
A Model Context Protocol (MCP) server that provides programmatic control over MuseScore!
Project
Agentic development harness for Claude Code — SPEC-driven plan/run/sync, TRUST 5 quality gates, model+effort routing, and Claude×GLM multi-LLM cost control. Single Go binary, 16 languages, zero deps.
Project
Agent skills for Qdrant vector search: scaling, performance optimization, search quality, monitoring, deployment, model migration, version upgrades, and SDK usage across Python, TypeScript, Rust, Go, .NET, Java
Collection
📚 A curated list of papers & technical articles on AI Quality & Safety
Project
Production-ready development workflows for Claude Code, powered by specialized AI agents.
Project
Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
Project
A Python client to interact with Arize API
Project
A framework for standardizing evaluations of large foundation models, beyond single-score reporting and rankings.
EvalBench is a flexible framework designed to measure the quality of generative AI (GenAI) workflows around database specific tasks.
Project
LLM evaluation framework for Elixir: evaluate and test LLM outputs, detect hallucinations, measure response quality
Project
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards