Skip to content
mager-bench1.3

Subscription archive

Original suite · 13 challenges

Thirteen coding tasks, run through headless Codex with a ChatGPT subscription and scored on correctness, code quality, and documentation.

Historical subscription leaderboard

2 models · Fri, 25 Sep 2026 21:56:45 UTC

Every model below uses GPT-5.6 Sol through Codex CLI as its judge. The older Sonnet 5 results are preserved in the archive.

#modelaccesscrct · qual · docsavg scoreavg latencyn
01GPT-6 Astra→subscription9.6 · 9.3 · 8.99.351138ms13
02GPT-5.6 Sol→subscription9.5 · 9.2 · 8.39.024856ms13

crct · qual · docs are the rubric dimensions averaged across all 13 challenges — the headline average hides whether a model is correct-but-undocumented or well-written-but-wrong.

click a model for its full challenge breakdown, raw responses, and judge notes. each published result comes from a saved headless Codex run.

the 13 challenges

full specs + rubrics →

score shown is the board average across all 2 scored models — a rough difficulty read. every model attempts every challenge.

methodology, honestly

Every score here uses codex-cli/gpt-5.6-sol. The judge is also a subject on this board, so self-judging bias is possible. Codex CLI prompts for an output length but does not enforce the API token cap. Every model page links to its full response and judge notes. The earlier API runs remain in the Sonnet 5 archive.

api

$ curl /api/results → 200 OK

Same data behind this page, as JSON. Cached 1h at the edge.