GPT-6 Astra
ChatGPT subscription · Codex CLI
13 original challenges. 3 Counterexample attempts. Every answer on the record.
9.3 out of 10
Mean across thirteen tasks, scored by GPT-5.6 Sol. 1 run per task, Sep 25, 2026.
3 attempts. Every result.
low reasoning · faulty versions exposed · preliminary calibration
The full box score
Highest to lowest. The lowest score here is async fetch at 7.3/10. Open any task to inspect the submission and the judge’s notes.
FizzBuzz
Can it get the basics right?
Binary search
Does the algorithm survive the edges?
README writer
Can it explain a tool well enough to use it?
Refactor
Can it improve code without changing its behavior?
Elixir tests
Can it handle another language’s conventions?
Doom raycaster
Can it ship a whole interactive system?
Python tests
Does it test beyond the happy path?
SQL analysis
Can it keep a complicated query correct?
Debugging
Can it diagnose the actual failure?
Go tests
Does it know how Go developers actually test?
API client
Can it build an interface someone else can use?
Slot machine
Do the visuals and the state agree?
Async fetch
What happens when the network misbehaves?
These scores are preserved from the original suite, including three retired warm-ups. They are not a new run. The model-judge setup, single samples, and self-judging for Sol limit how much to read into small differences.
What did its tests actually catch?
A model must predict the correct ledger’s outputs exactly before its tests can earn credit for exposing a fault. These counts show how often GPT-6 Astra caught each fault.
- 01Lost request history3 / 3 attempts
- 02Forgotten rejection3 / 3 attempts
- 03Unchecked payload3 / 3 attempts
- 04Revalidated retry3 / 3 attempts
- 05Partial transfer0 / 3 attempts
- 06Missing revision3 / 3 attempts
- 07Stale write3 / 3 attempts
- 08Overwritten history3 / 3 attempts
low reasoning · 12 events
low reasoning · 12 events
low reasoning · 12 events
Same frozen contract and settings across all attempts. 3 of 3 graded traces matched the reference. All responses are preserved, including weaker attempts and unscored failures.