DeepEval alternatives: 8 open-source AI and LLM evaluation tools
Metrics as of , from the GitHub or GitLab API of each repository. Refreshed monthly.
About DeepEval
DeepEval is a listed AI and LLM evaluation tool: Python framework for unit testing LLM applications, agents and RAG pipelines with LLM-as-a-judge and local metrics.
Order: Listed tools in the same category, with tools in the same primary language (GitHub API) first, then by GitHub stars. Each line gives one fact from the tool's own documentation where it differs from what DeepEval's documentation states, with its source.
8 alternatives to DeepEval
OpenAI Evals: Model providers: Models on the OpenAI API; completion functions in evals/registry/completion_fns or any CompletionFn implementation. source: Docs: How to run evals
Ragas: Model providers: OpenAI, Anthropic, Google directly; Azure OpenAI, AWS Bedrock, Google Vertex AI and others through LiteLLM. source: Docs: Customise models
Language Model Evaluation Harness: Model providers: Hugging Face transformers, vLLM, SGLang, GGUF via llama.cpp, NeMo, Megatron-LM; OpenAI, Anthropic, LiteLLM, local API servers. source: README
garak: Model providers: Hugging Face, Replicate, OpenAI, AWS Bedrock, LiteLLM, Cohere, Groq, NIM, ggml/GGUF, REST endpoints. source: README
Giskard: Model providers: Provider SDK extras such as openai and anthropic; default judge model openai/gpt-4o-mini. source: README
TruLens: Model providers: OpenAI, Azure OpenAI, LiteLLM, Google Gemini, AWS Bedrock, Snowflake Cortex, HuggingFace, LangChain models, OrcaRouter. source: README
HELM: Model providers: Models from various providers through one interface, such as OpenAI, Anthropic Claude, Google Gemini. source: README
Inspect: Model providers: OpenAI, Anthropic, Google, Grok, Mistral, DeepSeek; AWS Bedrock, Azure AI; Groq, Together AI; local models. source: Docs: Model Providers
Repository metrics
| OpenAI Evals | 19,4921 | 2023-04-061v0.1.1 | 2026-04-141 | MIT | Python | slow |
|---|---|---|---|---|---|---|
| DeepEval | 18,3931 | 2026-09-221python-v4.2.4 | 2026-09-221 | Apache-2.0 | Python | active |
| Ragas | 15,8241 | 2026-01-131v0.4.3 | 2026-02-241 | Apache-2.0 | Python | slow |
| Language Model Evaluation Harness | 14,0551 | 2026-08-311v0.4.13 | 2026-09-141 | MIT | Python | active |
| garak | 9,3301 | 2026-09-091v0.17.0 | 2026-09-161 | Apache-2.0 | Python | active |
| Giskard | 5,8341 | 2026-09-141giskard-checks/v1.0.4 | 2026-09-211 | Apache-2.0 | Python | active |
| TruLens | 3,5701 | 2026-09-031trulens-2.14.0 | 2026-09-221 | MIT | Python | active |
| HELM | 2,9211 | 2026-04-301v0.5.16 | 2026-06-051 | Apache-2.0 | Python | slow |
| Inspect | 2,8371 | 2025-11-281release/2025-11-28 | 2026-09-221 | MIT | Python | active |
1 Fetched from the GitHub or GitLab API on . Hover a value for its own date.