mager-bench

mager-bench

v1.1 // 13-challenge personal coding bench

Opinionated coding tasks, scored by an LLM judge on correctness, code quality, and documentation. Free-tier models run free. Paid models get crowdfunded.

#01 GPT-OSS 120B
6.4
/ 10 leader · 5 models scored
avg latency
8580ms
challenges
13
judge
claude-sonnet-5
last run
Sun, 26 Jul 2026 14:21:14 UTC

leaderboard

#modelcostcrct · qual · docsavg scoreavg latencyn
01GPT-OSS 120Bfree5.8 · 6.5 · 6.76.48580ms13
02Claude Sonnet 4.6paid5.9 · 5.9 · 6.66.233426ms13
03Claude Haiku 4.5cheap6.2 · 6.2 · 6.46.213326ms13
04Llama 3.3 70Bfree5.3 · 5.5 · 5.85.62244ms13
05Gemini 2.5 Flashfree4.8 · 5.0 · 5.75.226981ms13

crct · qual · docs are the rubric dimensions averaged across all 13 challenges — the headline average hides whether a model is correct-but-undocumented or well-written-but-wrong.

click a model for its full challenge breakdown, raw responses, and judge notes. free + cheap run by default; paid models land when crowdfunded.

the 13 challenges

full specs + rubrics →

score shown is the board average across all 5 scored models — a rough difficulty read. every model attempts every challenge.

fund the bench

10 paid models still waiting. Free tiers (Groq / Gemini) never need a card — crowdfunding buys the expensive head-to-heads.

open fund drive →

$0 / $100 season goal · 0%

not scored yet

  • Claude Opus 4.8unfunded~$8 / full suite
  • Claude Sonnet 5unfunded~$4 / full suite
  • GPT-4ounfunded~$5 / full suite
  • Gemini 2.5 Prounfunded~$4 / full suite
  • GPT-5.6unfunded~$6 / full suite
  • Claude Fable 5unfunded~$10 / full suite
  • Claude Opus 4.6unfunded~$7 / full suite
  • Claude Opus 4.7unfunded~$8 / full suite
  • GLM 2.5unfunded~$2 / full suite
  • Kimi K3unfunded~$2 / full suite
put them on the board →

methodology, honestly

Every score comes from a single LLM judge — currently claude-sonnet-5 — pinned across all runs so scores stay comparable. That judge is a Claude model scoring a board that includes Claude models: same-family bias is real, so treat cross-family gaps as directional, not gospel. Thinking and verdict budgets are separate contracts in the harness, and every score links to its raw response and judge notes so you can audit the reasoning instead of trusting the number.

api

$ curl /api/results → 200 OK

Same data behind this page, as JSON. Cached 1h at the edge.