Find the cheapest model that still passes.
TokenHunger benchmarks every model on your own task, ranks them by cost per correct answer, and lets you pay only for the runs you make — from someone who doesn't sell you the model.
The problem
Model prices vary by 10–100×. The expensive model isn't always the right one, and the cheapest model isn't always good enough. The only way to know is to test the models on the work you actually do — not a generic leaderboard, and not the advice of whoever happens to make one of the models.
What TokenHunger does
You bring a task and a handful of cases (an input and the answer you expect). TokenHunger runs each case across the models you pick, checks the answers, measures tokens and cost, then ranks the models by cost ÷ correct answers. A cheap model that passes beats an expensive one that also passes.
- Estimate for free. Anyone gets the cost table — no signup.
- Sign in with GitHub and get 5 free credits to start running.
- Run. Credits are reserved, the benchmark runs, and the unused portion of the hold is refunded — you pay only for actual spend.
- Decide. Ship the lean model. Buy more credits to keep going.
What makes it different
Leaderboards rank models on generic benchmarks. Eval platforms hand you a dashboard of twenty metrics and a YAML file to fill in. Routers decide for you at runtime, after you've already picked a quality bar. TokenHunger does one thing:
- One decision metric — cost per correct answer. Not a wall of charts. The cheapest model that clears your bar, ranked first.
- Your task, not a generic benchmark. Results come from the work you actually do, scored against the answers you actually expect.
- Neutral by design. We don't make a model and we don't take a cut from one. TokenHunger runs on its own separate billing and database, so the ranking has no thumb on the scale. Advice from a company that also sells you the model isn't advice.
- Zero setup. No API keys, no infrastructure, no config. Sign in, paste your cases, run.
- Pay only for the runs you make. Prepaid credits, not a subscription — built for deciding once, not for standing infrastructure.
How it's built
TokenHunger is the commercial cloud layer on top of costbench, an open-source benchmarking engine. The engine is open and license-separate; this site adds managed model access, accounts, credits and billing so you don't have to wire up provider keys yourself. It also speaks MCP, so agents and tools like QualityMax can drive benchmarks directly.
Who's behind it
TokenHunger is built and operated by an independent team under the QualityMax umbrella. Questions, feedback, or press — we'd love to hear from you on the contact page.