Suggest a tool

DeepEval vs Ragas: open-source AI and LLM evaluation tools compared

Metrics as of , from the GitHub or GitLab API of each repository. Refreshed monthly.

Summary

DeepEval: Python framework for unit testing LLM applications, agents and RAG pipelines with LLM-as-a-judge and local metrics.

Ragas: Python toolkit for evaluating LLM applications with LLM-based and traditional metrics and test set generation.

Both are listed under AI and LLM evaluation. The table gives repository metrics from the GitHub API with their fetch date, then the documented facts used for every comparison in this category, each linked to the README or docs page it comes from. A cell reads "not documented" when the fact was not found in the project's documentation (1 of 12 fact cells here); that does not mean the feature is absent.

Side by side

DeepEval and Ragas: metrics and documented facts, in the fixed row order used for AI and LLM evaluation comparisons.
FactDeepEvalRagas
Stars18,393115,8241
Forks1,96611,7181
Contributors33712461
Last release2026-09-2212026-01-131
Last commit2026-09-2212026-02-241
Commits in 90 days555101
LicenseApache-2.01Apache-2.01
Primary languagePython1Python1
Statusactive1slow1
Model providersOpenAI, Azure OpenAI, Ollama, Anthropic, Gemini, LiteLLM; custom models through DeepEvalBaseLLM source: Docs: Introduction to LLM MetricsOpenAI, Anthropic, Google directly; Azure OpenAI, AWS Bedrock, Google Vertex AI and others through LiteLLM source: Docs: Customise models
Evaluation metricsG-Eval, DAG; agentic, RAG, multi-turn, MCP and multimodal metrics; hallucination, bias, toxicity, JSON correctness source: READMEContext Precision, Context Recall, Faithfulness, Response Relevancy, Tool Call Accuracy, BLEU, ROUGE, Aspect Critic and others source: Docs: List of available metrics
Test definition formatCode (Python test files with assert_test, run by deepeval test run; Pytest-style) source: READMECode (Python metrics API, for example DiscreteMetric); ragas quickstart project templates source: README
CI integrationdeepeval test run in any shell-step CI; GitHub Actions example; GitLab CI, CircleCI, Jenkins named source: Docs: Unit Testing in CI/CDnot documentednot found in the project's documentation as of
Report formatsTestRun JSON files (results_folder); dashboard export as html or md (file_type); optional SQLite store source: Docs: Flags and ConfigsExperiment results saved as timestamped CSV files in experiments/ source: Docs: Experimentation
Install methodpip (pip install -U deepeval) source: READMEpip (pip install ragas), or pip install from the Git repository source: README

1 Fetched from the GitHub or GitLab API on . Hover a value for its own date.

When each fits

Written from each project's documented scope, not from preference. Neither tool is ranked.

DeepEval

Fits projects that unit test LLM apps such as AI agents, RAG pipelines and chatbots; its README describes a framework similar to Pytest but specialized for LLM apps. source: README

Ragas

Fits projects that evaluate LLM applications, including RAG systems, and generate test datasets; its README describes LLM-based and traditional metrics. source: README

More on these tools

Sources

  1. GitHub REST API: repository, read
  2. GitHub REST API: repository, read
  3. GitHub REST API: contributors, read
  4. GitHub REST API: contributors, read
  5. GitHub REST API: latest release, read
  6. GitHub REST API: latest release, read
  7. GitHub REST API: commits, read
  8. GitHub REST API: commits, read
  9. Docs: Introduction to LLM Metrics, read
  10. Docs: Customise models, read
  11. README, read
  12. Docs: List of available metrics, read
  13. README, read
  14. Docs: Unit Testing in CI/CD, read
  15. Docs: Flags and Configs, read
  16. Docs: Experimentation, read

Documented facts collected 2026-09-22. See the methodology for how metrics and facts are gathered.