GPT-5.6 Sol
ChatGPT subscription · Codex CLI
13 original challenges. 3 Counterexample attempts. Every answer on the record.
9.0 out of 10
Mean across thirteen tasks, scored by GPT-5.6 Sol. 1 run per task, Sep 25, 2026.
3 attempts. Every result.
low reasoning · faulty versions exposed · preliminary calibration
The full box score
Highest to lowest. The lowest score here is async fetch at 6.3/10. Open any task to inspect the submission and the judge’s notes.
FizzBuzz
Can it get the basics right?
Debugging
Can it diagnose the actual failure?
SQL analysis
Can it keep a complicated query correct?
README writer
Can it explain a tool well enough to use it?
Refactor
Can it improve code without changing its behavior?
Binary search
Does the algorithm survive the edges?
Elixir tests
Can it handle another language’s conventions?
API client
Can it build an interface someone else can use?
Go tests
Does it know how Go developers actually test?
Doom raycaster
Can it ship a whole interactive system?
Python tests
Does it test beyond the happy path?
Slot machine
Do the visuals and the state agree?
Async fetch
What happens when the network misbehaves?
These scores are preserved from the original suite, including three retired warm-ups. They are not a new run. The model-judge setup, single samples, and self-judging for Sol limit how much to read into small differences.
What did its tests actually catch?
A model must predict the correct ledger’s outputs exactly before its tests can earn credit for exposing a fault. These counts show how often GPT-5.6 Sol caught each fault.
- 01Lost request history3 / 3 attempts
- 02Forgotten rejection3 / 3 attempts
- 03Unchecked payload3 / 3 attempts
- 04Revalidated retry3 / 3 attempts
- 05Partial transfer0 / 3 attempts
- 06Missing revision3 / 3 attempts
- 07Stale write2 / 3 attempts
- 08Overwritten history1 / 3 attempts
low reasoning · 12 events
low reasoning · 12 events
low reasoning · 12 events
Same frozen contract and settings across all attempts. 3 of 3 graded traces matched the reference. All responses are preserved, including weaker attempts and unscored failures.