Cost-per-correct-answer benchmarking
Run your own cases across every model, score the answers, and rank by cost per correct answer — so a cheap model that passes beats an expensive one. The ranking engine is open-source; you only pay for the hosted runs you make.
Sign in · 5 free credits · How it works · Explore benchmarks · Open-source engine on GitHub
The open engine — local CLI/UI, model catalog, estimates, scoring, and cost-per-correct ranking logic — is yours to inspect or run yourself.
This site runs that same engine with managed provider keys, GitHub sign-in, credits, saved analyses, and MCP access for agents.