Public Standard
The Public Standard for Real World AI Performance
Generic benchmarks only go so far. getEvals evaluates models on the real tasks each industry relies on.
getEvals Index
A single measure of AI's potential economic impact — agentic model performance across finance, coding, and legal tasks, weighted by each sector's share of U.S. GDP.
RSI Index
Can a model do the research that builds the next model? Autonomous AI research scored against human records.
Web Search Index
Comparing native provider search against independent web-search tools on legal-research and finance-analysis tasks.
getEvals Multimodal Index
Weighted performance across finance, coding, and education tasks. Showing the potential impact that models can have on the economy.
Legal Research Bench
Evaluating agents on legal research tasks across diverse areas of US law.
Harvey's Legal Agent Benchmark
Tests an agent's ability to complete legal work using documents, spreadsheets, presentations, and file-system tools.
LegalBench
Evaluating language models on a wide range of open source legal reasoning tasks.
Finance Agent v2
Evaluating agents on core financial analyst tasks.
Excel Modeling Benchmark
Evaluating agents on Excel-based financial modeling tasks used in investment banking and private equity.
MortgageTax
Evaluating reading and understanding tax certificates as images.
TaxEval v2
A getEvals-created set of questions and responses to tax questions.
GPQA Diamond
Graduate-level Google-Proof Q&A benchmark evaluating models on questions that require deep reasoning.
MMLU Pro
Academic multiple-choice benchmark covering 14 subjects including STEM, humanities, and social sciences.
MMMU Pro
Multimodal multi-task benchmark spanning 30 subjects in 6 major disciplines.
Vibe Code Bench v1.1
Can models build web applications from scratch?
Code Migration
Can language models reimplement working programs in another language?
ProgramBench
Can language models rebuild programs from scratch?
Terminal-Bench 2.1
State-of-the-art set of difficult terminal-based tasks.
SWE-bench Verified
Solving production software engineering tasks.
LiveCodeBench
Our implementation of the LiveCodeBench benchmark.
IOI
International Olympiad in Informatics.
Powered by PumaAI