Suggest a tool

DeepEval

Metrics as of , from the GitHub or GitLab API of each repository. Refreshed monthly.

What DeepEval is

DeepEval is an open-source Python framework for evaluating large language model systems. The README compares it to Pytest for unit testing LLM applications. It evaluates whole applications, complete agent trajectories and individual agent steps such as tool calls and retrieval. Metric groups cover custom criteria through G-Eval and DAG, agents, RAG, multi-turn conversations, MCP and multimodal output. Metrics use LLM-as-a-judge, statistical methods or NLP models that can run locally. Test files run through the deepeval test run CLI command. Integrations include LangChain, LangGraph, CrewAI, LlamaIndex and OpenAI Agents. It also runs models against benchmarks such as MMLU, HellaSwag and GSM8K. A separate platform, Confident AI, integrates with it for comparing iterations and sharing reports.

Written from the project's README, read .

Category
AI and LLM evaluation
License
Apache-2.0
Language
Python
Changelog
Releases on GitHub

Repository metrics

Repository metrics
Stars18,3931
Forks1,9661
Open issues and PRs6351
Contributors3371
Commits in the last 90 days5551
Last commit2026-09-221
Last releasepython-v4.2.4, 2026-09-221
Fetched

1 Fetched from the GitHub or GitLab API on . Hover a value for its own date.

Status

activeLast commit within 90 days of the fetch date.

Computed from the last commit date and the archive flag on the fetch date. See the status rules.

Alternatives

Listed AI and LLM evaluation tools, same primary language first, then by GitHub stars. Each line gives one fact from the tool's documentation where it differs from DeepEval's, with its source. See also DeepEval alternatives.

  • OpenAI Evals: Model providers: Models on the OpenAI API; completion functions in evals/registry/completion_fns or any CompletionFn implementation. source: Docs: How to run evals

  • Ragas: Model providers: OpenAI, Anthropic, Google directly; Azure OpenAI, AWS Bedrock, Google Vertex AI and others through LiteLLM. source: Docs: Customise models

  • Language Model Evaluation Harness: Model providers: Hugging Face transformers, vLLM, SGLang, GGUF via llama.cpp, NeMo, Megatron-LM; OpenAI, Anthropic, LiteLLM, local API servers. source: README

  • garak: Model providers: Hugging Face, Replicate, OpenAI, AWS Bedrock, LiteLLM, Cohere, Groq, NIM, ggml/GGUF, REST endpoints. source: README

  • Giskard: Model providers: Provider SDK extras such as openai and anthropic; default judge model openai/gpt-4o-mini. source: README

  • TruLens: Model providers: OpenAI, Azure OpenAI, LiteLLM, Google Gemini, AWS Bedrock, Snowflake Cortex, HuggingFace, LangChain models, OrcaRouter. source: README

  • HELM: Model providers: Models from various providers through one interface, such as OpenAI, Anthropic Claude, Google Gemini. source: README

  • Inspect: Model providers: OpenAI, Anthropic, Google, Grok, Mistral, DeepSeek; AWS Bedrock, Azure AI; Groq, Together AI; local models. source: Docs: Model Providers

Comparisons

How to install

pip install -U deepeval
From README, read .

Questions

Is DeepEval open source?

Yes. DeepEval is released under Apache-2.0, an OSI-approved license, as reported by the GitHub API on 2026-09-22.

Is DeepEval maintained?

On 2026-09-22, the last commit to the default branch was on 2026-09-22, so the listed status is active. Rule: Last commit within 90 days of the fetch date.

How many GitHub stars does DeepEval have?

18,393 stars on 2026-09-22, from the GitHub API. The number is refreshed at each monthly update.

What language is DeepEval written in?

The repository's primary language, as reported by the GitHub API, is Python.

How do I install DeepEval?

The README gives this command: pip install -U deepeval

Sources

  1. GitHub REST API: repository, read
  2. GitHub REST API: commits, read
  3. GitHub REST API: latest release, read
  4. GitHub REST API: contributors, read
  5. README, read