Suggest a tool

Lighteval

Metrics as of , from the GitHub or GitLab API of each repository. Refreshed monthly.

What Lighteval is

Lighteval is a toolkit from Hugging Face for evaluating large language models across multiple backends. It evaluates models served remotely or already loaded in memory. Results are saved sample by sample for debugging and inspection. The task catalogue covers knowledge, math, code, chat, multilingual and core language understanding benchmarks. Users can add custom tasks and custom metrics. The CLI offers entry points for inspect-ai, Accelerate, Nanotron, vLLM, SGLang, inference endpoints and custom models. A Python API runs evaluation pipelines on in-memory models. The README states that Windows is untested and unsupported.

Written from the project's README, read .

Category
AI and LLM evaluation
License
MIT
Language
Python
Changelog
Releases on GitHub

Repository metrics

Repository metrics
Stars2,5471
Forks5581
Open issues and PRs4121
Contributors1241
Commits in the last 90 days41
Last commit2026-09-171
Last releasev0.13.0, 2025-11-241
Fetched

1 Fetched from the GitHub or GitLab API on . Hover a value for its own date.

Status

activeLast commit within 90 days of the fetch date.

Computed from the last commit date and the archive flag on the fetch date. See the status rules.

Alternatives

Listed AI and LLM evaluation tools, same primary language first, then by GitHub stars. Each line gives one fact from the tool's documentation where it differs from Lighteval's, with its source.

  • OpenAI Evals: Model providers: Models on the OpenAI API; completion functions in evals/registry/completion_fns or any CompletionFn implementation. source: Docs: How to run evals

  • DeepEval: Model providers: OpenAI, Azure OpenAI, Ollama, Anthropic, Gemini, LiteLLM; custom models through DeepEvalBaseLLM. source: Docs: Introduction to LLM Metrics

  • Ragas: Model providers: OpenAI, Anthropic, Google directly; Azure OpenAI, AWS Bedrock, Google Vertex AI and others through LiteLLM. source: Docs: Customise models

  • Language Model Evaluation Harness: Model providers: Hugging Face transformers, vLLM, SGLang, GGUF via llama.cpp, NeMo, Megatron-LM; OpenAI, Anthropic, LiteLLM, local API servers. source: README

  • garak: Model providers: Hugging Face, Replicate, OpenAI, AWS Bedrock, LiteLLM, Cohere, Groq, NIM, ggml/GGUF, REST endpoints. source: README

  • Giskard: Model providers: Provider SDK extras such as openai and anthropic; default judge model openai/gpt-4o-mini. source: README

  • TruLens: Model providers: OpenAI, Azure OpenAI, LiteLLM, Google Gemini, AWS Bedrock, Snowflake Cortex, HuggingFace, LangChain models, OrcaRouter. source: README

  • HELM: Model providers: Models from various providers through one interface, such as OpenAI, Anthropic Claude, Google Gemini. source: README

How to install

pip install lighteval
From README, read .

Questions

Is Lighteval open source?

Yes. Lighteval is released under MIT, an OSI-approved license, as reported by the GitHub API on 2026-09-22.

Is Lighteval maintained?

On 2026-09-22, the last commit to the default branch was on 2026-09-17, so the listed status is active. Rule: Last commit within 90 days of the fetch date.

How many GitHub stars does Lighteval have?

2,547 stars on 2026-09-22, from the GitHub API. The number is refreshed at each monthly update.

What language is Lighteval written in?

The repository's primary language, as reported by the GitHub API, is Python.

How do I install Lighteval?

The README gives this command: pip install lighteval

Sources

  1. GitHub REST API: repository, read
  2. GitHub REST API: commits, read
  3. GitHub REST API: latest release, read
  4. GitHub REST API: contributors, read
  5. README, read