Subscription archive
Original suite · 13 challengesThirteen coding tasks, run through headless Codex with a ChatGPT subscription and scored on correctness, code quality, and documentation.
Historical subscription leaderboard
2 models · Fri, 25 Sep 2026 21:56:45 UTCEvery model below uses GPT-5.6 Sol through Codex CLI as its judge. The older Sonnet 5 results are preserved in the archive.
| # | model | access | crct · qual · docs | avg score | avg latency | n |
|---|---|---|---|---|---|---|
| 01 | GPT-6 Astra→ | subscription | 9.6 · 9.3 · 8.9 | 9.3 | 51138ms | 13 |
| 02 | GPT-5.6 Sol→ | subscription | 9.5 · 9.2 · 8.3 | 9.0 | 24856ms | 13 |
crct · qual · docs are the rubric dimensions averaged across all 13 challenges — the headline average hides whether a model is correct-but-undocumented or well-written-but-wrong.
click a model for its full challenge breakdown, raw responses, and judge notes. each published result comes from a saved headless Codex run.
the 13 challenges
full specs + rubrics →Classic FizzBuzz — tests basic correctness and code style
Binary search implementation with full docs
Write a small HTTP API client class with error handling and docs
Write a README for a CLI tool — tests documentation ability directly
Refactor messy code and explain each change
Write pytest tests for a provided function — tests edge-case thinking and assertion quality
Find and fix 3 bugs in broken Python code — tests careful reading and correctness reasoning
Write async Python for concurrent HTTP fetching with timeout and retry
Write a complex SQL query with CTEs, window functions, and aggregations
Write idiomatic Go table-driven tests — tests knowledge of Go testing conventions
Write idiomatic Elixir ExUnit tests — tests knowledge of Elixir testing conventions
Build a Doom-style raycasting FPS engine in a single HTML file — the mager-bench signature challenge
Build a Vegas-style slot machine in a single HTML file — reels, pay table, betting, wins
score shown is the board average across all 2 scored models — a rough difficulty read. every model attempts every challenge.
methodology, honestly
Every score here uses codex-cli/gpt-5.6-sol. The judge is also a subject on this board, so self-judging bias is possible. Codex CLI prompts for an output length but does not enforce the API token cap. Every model page links to its full response and judge notes. The earlier API runs remain in the Sonnet 5 archive.
api
$ curl /api/results → 200 OK
Same data behind this page, as JSON. Cached 1h at the edge.