Earlier benchmarks
The questions changed. The original prompts, answers, and scoring records are still here. Scores from different versions are not comparable.
1.2 · Three everyday programs
Weighted bills, contact CSVs, and meeting times. The first GPT-6 Sol attempt failed at the provider and remains unscored.
View v1.2 tasks and attempt1.1 · Counterexample Lab
Models wrote test traces to expose eight ledger faults. All 16 calibration attempts are preserved.
View calibration runsOriginal thirteen tasks
The ChatGPT subscription board, scored by GPT-5.6 Sol on a 0–10 scale.
View subscription boardEarlier API board
The original runs using Sonnet 5 as judge. Kept separate from the subscription results.
View Sonnet 5 board