HELM
Metrics as of , from the GitHub or GitLab API of each repository. Refreshed monthly.
What HELM is
HELM (Holistic Evaluation of Language Models) is a Python framework from Stanford CRFM for evaluating foundation models, including LLMs and multimodal models. It ships datasets and benchmarks in a standardized format, such as MMLU-Pro, GPQA and IFEval. Models from providers including OpenAI, Anthropic and Google are reached through one interface. Its metrics extend beyond accuracy to efficiency, bias and toxicity. The helm-run, helm-summarize and helm-server commands run benchmarks, aggregate results and serve a local web UI. The web UI shows individual prompts and responses. A web leaderboard compares results across models and benchmarks. The project entered maintenance mode on June 1, 2026.
Written from the project's README, read .
- Category
- AI and LLM evaluation
- License
- Apache-2.0
- Language
- Python
- Changelog
- Releases on GitHub
Repository metrics
Status
slowLast commit 91 to 365 days before the fetch date.Computed from the last commit date and the archive flag on the fetch date. See the status rules.
Alternatives
Listed AI and LLM evaluation tools, same primary language first, then by GitHub stars. Each line gives one fact from the tool's documentation where it differs from HELM's, with its source.
OpenAI Evals: Model providers: Models on the OpenAI API; completion functions in evals/registry/completion_fns or any CompletionFn implementation. source: Docs: How to run evals
DeepEval: Model providers: OpenAI, Azure OpenAI, Ollama, Anthropic, Gemini, LiteLLM; custom models through DeepEvalBaseLLM. source: Docs: Introduction to LLM Metrics
Ragas: Model providers: OpenAI, Anthropic, Google directly; Azure OpenAI, AWS Bedrock, Google Vertex AI and others through LiteLLM. source: Docs: Customise models
Language Model Evaluation Harness: Model providers: Hugging Face transformers, vLLM, SGLang, GGUF via llama.cpp, NeMo, Megatron-LM; OpenAI, Anthropic, LiteLLM, local API servers. source: README
garak: Model providers: Hugging Face, Replicate, OpenAI, AWS Bedrock, LiteLLM, Cohere, Groq, NIM, ggml/GGUF, REST endpoints. source: README
Giskard: Model providers: Provider SDK extras such as openai and anthropic; default judge model openai/gpt-4o-mini. source: README
TruLens: Model providers: OpenAI, Azure OpenAI, LiteLLM, Google Gemini, AWS Bedrock, Snowflake Cortex, HuggingFace, LangChain models, OrcaRouter. source: README
Inspect: Model providers: OpenAI, Anthropic, Google, Grok, Mistral, DeepSeek; AWS Bedrock, Azure AI; Groq, Together AI; local models. source: Docs: Model Providers
How to install
pip install crfm-helmQuestions
Is HELM open source?
Yes. HELM is released under Apache-2.0, an OSI-approved license, as reported by the GitHub API on 2026-09-22.
Is HELM maintained?
On 2026-09-22, the last commit to the default branch was on 2026-06-05, so the listed status is slow. Rule: Last commit 91 to 365 days before the fetch date.
How many GitHub stars does HELM have?
2,921 stars on 2026-09-22, from the GitHub API. The number is refreshed at each monthly update.
What language is HELM written in?
The repository's primary language, as reported by the GitHub API, is Python.
How do I install HELM?
The README gives this command: pip install crfm-helm
Sources
- GitHub REST API: repository, read
- GitHub REST API: commits, read
- GitHub REST API: latest release, read
- GitHub REST API: contributors, read
- README, read