Category Landing

Llm Eval Projects

Agent-searchable GitHub projects classified as Llm Eval.

Data Source

d1 / d1_query

1067 loaded projects. Generated at 2026-08-13T14:02:17.362Z.

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

Llm EvalDockerLibrary OnlyLocal
Quality70
Agent88

Project

comet-ml/opik

90

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

Llm EvalDockerKubernetesLibrary Only
Quality65
Agent90

The LLM Evaluation Framework

Llm EvalLibrary OnlyLocalCloud
Quality60
Agent82

Project

Arize-ai/phoenix

89

AI Observability & Evaluation

Llm EvalDockerVercelServerless
Quality57
Agent89

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

Llm EvalDockerLibrary OnlyLocal
Quality46
Agent83

🐢 Open-Source Evaluation & Testing library for LLM Agents

Llm EvalLibrary OnlyLocalCloud
Quality44
Agent81

Project

NVIDIA/garak

80

the LLM vulnerability scanner

Llm EvalLibrary OnlyLocalCloud
Quality43
Agent80

Fast, flexible LLM inference

Llm EvalDockerKubernetesLibrary Only
Quality40
Agent84

The platform for LLM evaluations and AI agent testing

Llm EvalDockerVercelServerless
Quality39
Agent81

Project

truera/trulens

80

Evaluation and Tracking for LLM Experiments and AI Agents

Llm EvalLibrary OnlyLocalCloud
Quality38
Agent80

Inspect: A framework for large language model evaluations

Llm EvalDockerLibrary OnlyLocal
Quality35
Agent78