Guide
Score Inputs
Quality score weights star movement, commits, releases, contributors, and issue response; agent score adds documentation, deployment, popularity, and community fit.
Git.Top Guide
Inspect quality scores, confidence, freshness, and risk signals for open-source project selection.
Guide
Quality score weights star movement, commits, releases, contributors, and issue response; agent score adds documentation, deployment, popularity, and community fit.
Guide
Use /api/quality, project-level quality_signal_confidence, and sync status together before treating a recommendation as high-confidence.
Data Source
1213 loaded projects. Generated at 2026-09-20T00:16:32.012Z.
VocalVerse: A powerful vocal evaluation framework powered by the Qwen LLMs
Project
Evaluation package that allows benchmarking of agentic AIs from various sources and frameworks by producing statistical results which can be compared across different use cases and datasets.
Project
A Model Context Protocol (MCP) server that provides programmatic control over MuseScore!
Project
Agentic development harness for Claude Code — SPEC-driven plan/run/sync, TRUST 5 quality gates, model+effort routing, and Claude×GLM multi-LLM cost control. Single Go binary, 16 languages, zero deps.
Project
Agent skills for Qdrant vector search: scaling, performance optimization, search quality, monitoring, deployment, model migration, version upgrades, and SDK usage across Python, TypeScript, Rust, Go, .NET, Java
Collection
📚 A curated list of papers & technical articles on AI Quality & Safety
Project
Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
Project
Development workflows for Claude Code that keep broad exploration focused on the outcome you approved.
Official Monte Carlo toolkit for AI coding agents. Skills and plugins that bring data and agent observability — monitoring, triaging, troubleshooting, health checks — into Claude Code, Cursor, and more.
Project
A Python client to interact with Arize API
Project
A framework for standardizing evaluations of large foundation models, beyond single-score reporting and rankings.
EvalBench is a flexible framework designed to measure the quality of generative AI (GenAI) workflows around database specific tasks.