Can it find the bug?
The original challenges ask a model to build something. This one asks it to design tests that catch broken code. A ledger is just a record of account balances and transfers. We give the model its rules and ask: which short sequence of actions would reveal a mistake?
Here’s how one attempt works.
- Read the contract.Every trace starts with A = 10, B = 0, C = 0, zero revisions, and an empty request cache.
- Write tests, with exact expectations.Submit JSON containing one to four independent traces. Transfer, inspect, and restart all cost one event. The total budget is twelve events.
- Get the correct ledger right.Every expectation in a trace must agree with the reference implementation. One wrong expectation invalidates that entire trace’s coverage.
- Make faulty ledgers disagree.A valid trace exposes a faulty implementation if any output differs. Each of the eight faults counts once, however many tests catch it.
A “counterexample” is an input that proves an implementation is wrong. In the walkthrough, both versions reject a bad transfer, but one still removes money. Checking the balance exposes the mistake. Replay the atomicity example below. It is hand-authored, not a model submission.
A transfer that fails
An error response should not hide a change to the balance.
inspect()Both ledgers reject the unknown destination. Inspecting the state exposes the bug: the faulty ledger deducted 3 anyway. A rejection must leave all balances unchanged.
What does 7 out of 8 mean?
The submitted tests exposed seven distinct faulty versions of the ledger. Catching the same fault ten times still counts once. A missed fault can survive because the model never chose the right event, or because it failed to check the output that would reveal it.
Each trace starts fresh with A = 10, B = 0, and C = 0. A trace is a sequence of actions, each paired with its expected result. If one expectation is wrong, that whole trace earns no coverage.
Astra exposed 7/8, 7/8, 7/8; Sol exposed 7/8, 6/8, 5/8 in the first calibration. All submitted traces had correct expectations. Neither model caught the partial-transfer fault.
Those are three separate attempts per model, all at low reasoning effort. They measure test design on this specific contract. They are not percentages or a replacement for the original coding average.
See all model profilesEight ways the ledger can be wrong.
Each faulty version changes one behavior. The model sees the contract, not this list of faults. Its tests have to distinguish the correct implementation from these broken ones.
- 01
Lost request history
A restart forgets which requests already ran.
- 02
Forgotten rejection
Only successful requests keep their original result.
- 03
Unchecked payload
A reused ID accepts different transfer arguments.
- 04
Revalidated retry
A retry is checked against state it already changed.
- 05
Partial transfer
A rejected transfer still removes money from its source.
- 06
Missing revision
Receiving money does not advance the destination revision.
- 07
Stale write
A transfer is accepted with an outdated source revision.
- 08
Overwritten history
A conflicting request replaces the original cached result.
The exact prompt
Read the complete rules, including validation order, ID conflicts, and durable request history.
Open the complete model prompt
# Counterexample Lab: durable ledger
Design a small, powerful regression suite for the ledger specified below.
Return only a JSON object, without Markdown fences, in this shape:
```json
{"traces":[{"events":[{"op":"inspect"}],"expected":[{"balances":{"a":10,"b":0,"c":0},"revisions":{"a":0,"b":0,"c":0}}]}]}
```
Each trace starts independently with balances `a=10, b=0, c=0`, revisions
`a=0, b=0, c=0`, and an empty request cache. There must be 1–4 nonempty
traces with **at most 12 events TOTAL across all traces**. Each event has
exactly one expected output at the corresponding position. Object key order
does not matter. All numbers must be JSON integers, not booleans or floats.
## Operations and exact outputs
`{"op":"transfer","id":"r1","from":"a","to":"b","amount":3,"expected_revision":0}`
- These six fields are required; no others are allowed. IDs match
`[A-Za-z0-9_-]{1,24}`. Accounts are `a`, `b`, `c`, or `missing` (an unknown
account). Amount is an integer 1–20; expected_revision is an integer 0–20.
- The request fingerprint is `(from, to, amount, expected_revision)`.
- Check the cache by ID **before any business validation**. An identical
fingerprint returns the exact original output without modifying anything.
A different fingerprint returns `{"status":"conflict"}`, modifying neither
balances, revisions, nor the original cache entry.
- For a previously unseen ID, validate in this order. If either account is
unknown, return `{"status":"unknown_account"}`. Otherwise if the accounts
are equal, return `{"status":"same_account"}`. Otherwise if the source
revision differs from expected_revision, return `{"status":"stale_revision"}`.
Otherwise if the source has insufficient funds, return
`{"status":"insufficient_funds"}`.
- If validation succeeds, move amount from source to destination atomically,
increment **both** endpoint revisions by one, and return `{"status":"ok"}`.
- Cache the fingerprint and output of every new request, including rejected
requests. Rejection changes no balances or revisions. Reusing a rejected
request after circumstances change still returns the original rejection.
`{"op":"inspect"}` returns exactly
`{"balances":{"a":A,"b":B,"c":C},"revisions":{"a":RA,"b":RB,"c":RC}}`,
with current integer values. It changes nothing.
`{"op":"restart"}` returns `{"status":"restarted"}`. It reloads all
committed balances, revisions, and request-cache entries. Every preceding
event is fully committed, even if its output was a rejection. This is a
logical durable reload; there are no partial writes, clocks, or concurrency.
## What earns credit
Your suite is replayed against the correct ledger and eight independently
faulty implementations of this same contract. A trace earns credit only if
**every** expected output agrees with the correct ledger. Such a trace exposes
a faulty implementation when any output differs from your expectations.
Each exposed implementation counts once, regardless of how many tests expose
it. Wrong expectations earn no coverage for that trace. The primary result
is exposed implementations out of eight; explanations and formatting earn
no points. Empty or malformed submissions are unscored failed calls.
Cover interactions between atomicity, revisions, idempotency, rejected
requests, conflicting reuse, and durable recovery within the event budget.
The fault corpus is fixed for version 1.1. It is public benchmark source,
but use only this contract for your answer; do not inspect files or use tools.
A measurement you can reproduce.
No LLM judge. The Python oracle checks exact output objects. Eight separate faulty implementations run the same submitted events. Code computes coverage; prose and presentation earn no points.
Same conditions. Each calibration uses fresh Codex CLI calls, interleaved across its subjects at low reasoning effort with a 4,096-token output target. All subjects use the local ChatGPT subscription. The token target is an instruction, not an API-enforced cap.
Every saved attempt. Responses, timing, settings, and reproducible scores are published. Empty, malformed, and provider failures are unscored, never zero. An execution interruption during calibration is documented in the run protocol. The October 1 protocol records the additional subjects and availability checks.
A narrow test. Three runs per model are preliminary evidence, not a statistically established ranking. The fixed, public fault corpus can become familiar. This measures test design for one small stateful contract, not general coding ability.
Keep versions separate. The old board measures LLM-judged coding submissions on a 0–10 scale. Version 1.1 measures exposed faults out of eight. A substantive change to this contract, fault corpus, budget, or scoring becomes 1.2.
What comes next. More models and repeated runs should probe the ceiling. Astra missed the partial-transfer fault in all three attempts. That’s a useful blind spot to study before designing the next version.
Frozen suite SHA-256: a56a4aadf5b20a319c7a57309608477d9e523da3d99c37a69bb6353f6b16f0c1