Evals
How the checks and repairs are doing
Every repair is logged with its error class, the prompt used, whether it fixed the number, and how many pixels it touched outside its region. The repair prompt for the next attempt is chosen from these numbers.
Problem: a repair prompt that usually fails wastes money and time. What this does: each prompt's fix rate is tracked per error class, and the next attempt samples the likeliest winner (Thompson sampling). Why it matters: the loop gets cheaper and more reliable with use, and you can see why. How to use it: watch which prompts win; a new prompt starts untested and earns its place.
Every sample organisation (Wrenfield, Alderbank, Harrowfield County, Kettleby Advisory) is invented; every run is real.
First pass by setting
every live run, finished or not| Run | Setting | Engine | Facts | First check | Unsourced | Final | Rounds | Cost |
|---|---|---|---|---|---|---|---|---|
222025_12a0e9alderbank | 2K · thinking on | 2026-10-06.1 render-v1 | 9 | 9/9 | 0 | 9/9 | 0 | $0.076 |
222114_6c4783wrenfield | 2K · thinking on | 2026-10-06.1 render-v1 | 31 | 31/31 | 0 | 31/31 | 0 | $0.085 |
222359_62eb47wrenfield | 1K · minimal thinking | 2026-10-06.1 render-v1 | 31 | 31/31 | 3 | 31/31 | 2 | $0.270 |
222702_ecf75dwrenfield | 1K · minimal thinking | 2026-10-06.1 render-v1 | 31 | 31/31 | 0 | 31/31 | 0 | $0.058 |
222738_675707wrenfield | 1K · thinking on | 2026-10-06.1 render-v1 | 31 | 31/31 | 0 | 31/31 | 0 | $0.071 |
222824_caa176wrenfield | 1K · thinking on | 2026-10-06.1 render-v1 | 31 | 31/31 | 0 | 31/31 | 0 | $0.079 |
222929_0174ecwrenfield-full | 2K · thinking on | 2026-10-06.1 render-v1 | 57 | 56/57 | 88 | interrupted | $1.102 | |
225247_03fff1wrenfield-full | 2K · thinking on | 2026-10-06.2 render-v2 | 57 | 55/57 | 2 | 56/57 | 3 | $0.417 |
225829_4e73dbharrowfield | 2K · thinking on | 2026-10-06.2 render-v2 | 10 | 9/10 | 3 | 9/10 | 3 | $0.548 |
230652_c5b051harrowfield | 2K · thinking on | 2026-10-06.3 render-v3 | 10 | 10/10 | 1 | 10/10 | 1 | $0.146 |
231852_440ed7wrenfield | 2K · thinking on | 2026-10-06.4 render-v3 | 31 | 31/31 | 0 | 31/31 | 0 | $0.087 |
231949_a3b9fawrenfield-draft | 1K · minimal thinking | 2026-10-06.4 render-v3 | 31 | 30/31 | 0 | 30/31 | 3 | $0.135 |
232749_82f43fwrenfield-draft | 1K · minimal thinking | 2026-10-06.5 render-v3 | 31 | 31/31 | 0 | 31/31 | 0 | $0.056 |
001353_926ed2wrenfield | 2K · thinking on | 2026-10-06.6 render-v3 | 31 | 31/31 | 0 | 31/31 | 0 | $0.096 |
001501_f5ccf0wrenfield-draft | 1K · minimal thinking | 2026-10-06.6 render-v3 | 31 | 29/31 | 0 | 30/31 corrected · receipt | 1 | $0.167 |
001645_7db54aalderbank | 2K · thinking on | 2026-10-06.6 render-v3 | 9 | 9/9 | 0 | 9/9 | 0 | $0.077 |
002103_932e88wrenfield-draft | 1K · minimal thinking | 2026-10-06.7 render-v3 | 31 | 30/31 | 1 | 31/31 | 2 | $0.232 |
015012_50d896wrenfield-draft | 1K · minimal thinking | 2026-10-06.8 render-v3 | 31 | 31/31 | 0 | 31/31 | 0 | $0.075 |
Cost = every API call the run made (render, reads, strict checks, edits, re-reads, any full re-render), summed from the run's own ledger lines. Controls are shared across runs and counted separately. Engine = the rules version that produced the row: the 9/10 Harrowfield row is a false flag from engine 2026-10-06.2 (its strict reader wrote “$18.6” for “$18.6 million”); the gate now confirms by digits, and the re-run is 10/10. “Corrected” = a later, stricter rule re-scored the run's own stored reads, and its receipt shows both results. “Interrupted” = stopped before a receipt; its spend still counts.
All API spend so far: $4.52: runs $3.78 (17 finished, 1 interrupted at $1.10), controls $0.50, /verify $0.11, the rest probes and experiments.
Repair prompts, ranked
— attemptsNo repairs logged yet.
Controls (latest)
—Controls run with the first live render.
Measured cost per call
median| Render · 1K · minimal thinking | $0.0485 (n=7) |
| Render · 1K · thinking on | $0.0609 (n=2) |
| Render · 2K · thinking on | $0.0729 (n=9) |
| Masked edit · 1K · minimal thinking | $0.0454 (n=8) |
| Masked edit · 2K · thinking on | $0.0596 (n=24) |
| Blind read | $0.0027 (n=132) |
.data/llm-calls.jsonl (successful calls; price from each call's own usageMetadata)