Suggest a tool

DeepEval alternatives: 8 open-source AI and LLM evaluation tools

Metrics as of , from the GitHub or GitLab API of each repository. Refreshed monthly.

About DeepEval

DeepEval is a listed AI and LLM evaluation tool: Python framework for unit testing LLM applications, agents and RAG pipelines with LLM-as-a-judge and local metrics.

Order: Listed tools in the same category, with tools in the same primary language (GitHub API) first, then by GitHub stars. Each line gives one fact from the tool's own documentation where it differs from what DeepEval's documentation states, with its source.

8 alternatives to DeepEval

  • OpenAI Evals: Model providers: Models on the OpenAI API; completion functions in evals/registry/completion_fns or any CompletionFn implementation. source: Docs: How to run evals

  • Ragas: Model providers: OpenAI, Anthropic, Google directly; Azure OpenAI, AWS Bedrock, Google Vertex AI and others through LiteLLM. source: Docs: Customise models

  • Language Model Evaluation Harness: Model providers: Hugging Face transformers, vLLM, SGLang, GGUF via llama.cpp, NeMo, Megatron-LM; OpenAI, Anthropic, LiteLLM, local API servers. source: README

  • garak: Model providers: Hugging Face, Replicate, OpenAI, AWS Bedrock, LiteLLM, Cohere, Groq, NIM, ggml/GGUF, REST endpoints. source: README

  • Giskard: Model providers: Provider SDK extras such as openai and anthropic; default judge model openai/gpt-4o-mini. source: README

  • TruLens: Model providers: OpenAI, Azure OpenAI, LiteLLM, Google Gemini, AWS Bedrock, Snowflake Cortex, HuggingFace, LangChain models, OrcaRouter. source: README

  • HELM: Model providers: Models from various providers through one interface, such as OpenAI, Anthropic Claude, Google Gemini. source: README

  • Inspect: Model providers: OpenAI, Anthropic, Google, Grok, Mistral, DeepSeek; AWS Bedrock, Azure AI; Groq, Together AI; local models. source: Docs: Model Providers

Repository metrics

DeepEval and its alternatives. Sorted by GitHub stars, descending, descending. Select a column heading to change the sort.
OpenAI Evals19,49212023-04-061v0.1.12026-04-141MITPythonslow
DeepEval18,39312026-09-221python-v4.2.42026-09-221Apache-2.0Pythonactive
Ragas15,82412026-01-131v0.4.32026-02-241Apache-2.0Pythonslow
Language Model Evaluation Harness14,05512026-08-311v0.4.132026-09-141MITPythonactive
garak9,33012026-09-091v0.17.02026-09-161Apache-2.0Pythonactive
Giskard5,83412026-09-141giskard-checks/v1.0.42026-09-211Apache-2.0Pythonactive
TruLens3,57012026-09-031trulens-2.14.02026-09-221MITPythonactive
HELM2,92112026-04-301v0.5.162026-06-051Apache-2.0Pythonslow
Inspect2,83712025-11-281release/2025-11-282026-09-221MITPythonactive

1 Fetched from the GitHub or GitLab API on . Hover a value for its own date.

Comparisons with DeepEval