Skip to content
mager-bench1.3
Original 13 / Historical coding results

9.3 out of 10

Mean across thirteen tasks, scored by GPT-5.6 Sol. 1 run per task, Sep 25, 2026.

Correctness9.6
Code quality9.3
Documentation8.9
Counterexample Lab / Version 1.1

3 attempts. Every result.

low reasoning · faulty versions exposed · preliminary calibration

Original 13

The full box score

Highest to lowest. The lowest score here is async fetch at 7.3/10. Open any task to inspect the submission and the judge’s notes.

Scoring rules

These scores are preserved from the original suite, including three retired warm-ups. They are not a new run. The model-judge setup, single samples, and self-judging for Sol limit how much to read into small differences.

New challenge / 1.1

What did its tests actually catch?

A model must predict the correct ledger’s outputs exactly before its tests can earn credit for exposing a fault. These counts show how often GPT-6 Astra caught each fault.

Understand the test
  • 01Lost request history
    3 / 3 attempts
  • 02Forgotten rejection
    3 / 3 attempts
  • 03Unchecked payload
    3 / 3 attempts
  • 04Revalidated retry
    3 / 3 attempts
  • 05Partial transfer
    0 / 3 attempts
  • 06Missing revision
    3 / 3 attempts
  • 07Stale write
    3 / 3 attempts
  • 08Overwritten history
    3 / 3 attempts

Same frozen contract and settings across all attempts. 3 of 3 graded traces matched the reference. All responses are preserved, including weaker attempts and unscored failures.