DeepEval vs Ragas: open-source AI and LLM evaluation tools compared
Metrics as of , from the GitHub or GitLab API of each repository. Refreshed monthly.
Summary
DeepEval: Python framework for unit testing LLM applications, agents and RAG pipelines with LLM-as-a-judge and local metrics.
Ragas: Python toolkit for evaluating LLM applications with LLM-based and traditional metrics and test set generation.
Both are listed under AI and LLM evaluation. The table gives repository metrics from the GitHub API with their fetch date, then the documented facts used for every comparison in this category, each linked to the README or docs page it comes from. A cell reads "not documented" when the fact was not found in the project's documentation (1 of 12 fact cells here); that does not mean the feature is absent.
Side by side
| Fact | DeepEval | Ragas |
|---|---|---|
| Stars | 18,3931 | 15,8241 |
| Forks | 1,9661 | 1,7181 |
| Contributors | 3371 | 2461 |
| Last release | 2026-09-221 | 2026-01-131 |
| Last commit | 2026-09-221 | 2026-02-241 |
| Commits in 90 days | 5551 | 01 |
| License | Apache-2.01 | Apache-2.01 |
| Primary language | Python1 | Python1 |
| Status | active1 | slow1 |
| Model providers | OpenAI, Azure OpenAI, Ollama, Anthropic, Gemini, LiteLLM; custom models through DeepEvalBaseLLM source: Docs: Introduction to LLM Metrics | OpenAI, Anthropic, Google directly; Azure OpenAI, AWS Bedrock, Google Vertex AI and others through LiteLLM source: Docs: Customise models |
| Evaluation metrics | G-Eval, DAG; agentic, RAG, multi-turn, MCP and multimodal metrics; hallucination, bias, toxicity, JSON correctness source: README | Context Precision, Context Recall, Faithfulness, Response Relevancy, Tool Call Accuracy, BLEU, ROUGE, Aspect Critic and others source: Docs: List of available metrics |
| Test definition format | Code (Python test files with assert_test, run by deepeval test run; Pytest-style) source: README | Code (Python metrics API, for example DiscreteMetric); ragas quickstart project templates source: README |
| CI integration | deepeval test run in any shell-step CI; GitHub Actions example; GitLab CI, CircleCI, Jenkins named source: Docs: Unit Testing in CI/CD | not documentednot found in the project's documentation as of |
| Report formats | TestRun JSON files (results_folder); dashboard export as html or md (file_type); optional SQLite store source: Docs: Flags and Configs | Experiment results saved as timestamped CSV files in experiments/ source: Docs: Experimentation |
| Install method | pip (pip install -U deepeval) source: README | pip (pip install ragas), or pip install from the Git repository source: README |
1 Fetched from the GitHub or GitLab API on . Hover a value for its own date.
When each fits
Written from each project's documented scope, not from preference. Neither tool is ranked.
DeepEval
Fits projects that unit test LLM apps such as AI agents, RAG pipelines and chatbots; its README describes a framework similar to Pytest but specialized for LLM apps. source: README
Ragas
Fits projects that evaluate LLM applications, including RAG systems, and generate test datasets; its README describes LLM-based and traditional metrics. source: README
More on these tools
Sources
- GitHub REST API: repository, read
- GitHub REST API: repository, read
- GitHub REST API: contributors, read
- GitHub REST API: contributors, read
- GitHub REST API: latest release, read
- GitHub REST API: latest release, read
- GitHub REST API: commits, read
- GitHub REST API: commits, read
- Docs: Introduction to LLM Metrics, read
- Docs: Customise models, read
- README, read
- Docs: List of available metrics, read
- README, read
- Docs: Unit Testing in CI/CD, read
- Docs: Flags and Configs, read
- Docs: Experimentation, read
Documented facts collected 2026-09-22. See the methodology for how metrics and facts are gathered.