Independent Evaluation, Unbiased Benchmarks
Testing AI on Real-World Tasks
We benchmark the world's leading AI models on economically valuable tasks such as finance, software, and frontier risk like cybersecurity, recursive self improvement and mental health. We run all of our own evaluations and create many of our benchmarks in-house.
Showing leading models on the getEvals Index. Best model from each lab is highlighted in the full leaderboard.
Powered by PumaAI