Guide
What The Score Means
Git.Top Score summarizes maintenance, community, documentation, stability, adoption, deployment fit, and agent readability into a decision aid instead of a popularity counter.
Git.Top Guide
Understand Git.Top Score, Agent Score, quality confidence, and the signals behind project selection.
Guide
Git.Top Score summarizes maintenance, community, documentation, stability, adoption, deployment fit, and agent readability into a decision aid instead of a popularity counter.
Guide
Inspect the score breakdown, confidence, freshness, and quality evidence before comparing projects or sending an agent-readable recommendation.
Data Source
1067 loaded projects. Generated at 2026-08-13T13:33:56.552Z.
Project
Evaluation package that allows benchmarking of agentic AIs from various sources and frameworks by producing statistical results which can be compared across different use cases and datasets.
Project
A Model Context Protocol (MCP) server that provides programmatic control over MuseScore!
Project
Agentic development harness for Claude Code — SPEC-driven plan/run/sync, TRUST 5 quality gates, model+effort routing, and Claude×GLM multi-LLM cost control. Single Go binary, 16 languages, zero deps.
Project
Agent skills for Qdrant vector search: scaling, performance optimization, search quality, monitoring, deployment, model migration, version upgrades, and SDK usage across Python, TypeScript, Rust, Go, .NET, Java
Collection
📚 A curated list of papers & technical articles on AI Quality & Safety
Project
Development workflows for Claude Code that keep broad exploration focused on the outcome you approved.
Project
Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
Project
A Python client to interact with Arize API
Project
A framework for standardizing evaluations of large foundation models, beyond single-score reporting and rankings.
EvalBench is a flexible framework designed to measure the quality of generative AI (GenAI) workflows around database specific tasks.
Project
LLM evaluation framework for Elixir: evaluate and test LLM outputs, detect hallucinations, measure response quality
Project
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards