Skip to content
FactFrame

Public demo: every run here is a recorded replay of a real run, labelled with when it ran. Nothing on this site calls a model. Live runs are by request during early access.

Evals

How the checks and repairs are doing

Every repair is logged with its error class, the prompt used, whether it fixed the number, and how many pixels it touched outside its region. The repair prompt for the next attempt is chosen from these numbers.

Problem: a repair prompt that usually fails wastes money and time. What this does: each prompt's fix rate is tracked per error class, and the next attempt samples the likeliest winner (Thompson sampling). Why it matters: the loop gets cheaper and more reliable with use, and you can see why. How to use it: watch which prompts win; a new prompt starts untested and earns its place.

Every sample organisation (Wrenfield, Alderbank, Harrowfield County, Kettleby Advisory) is invented; every run is real.

First pass by setting

every live run, finished or not
RunSettingEngineFactsFirst checkUnsourcedFinalRoundsCost
222025_12a0e9
alderbank
2K · thinking on2026-10-06.1
render-v1
99/909/90$0.076
222114_6c4783
wrenfield
2K · thinking on2026-10-06.1
render-v1
3131/31031/310$0.085
222359_62eb47
wrenfield
1K · minimal thinking2026-10-06.1
render-v1
3131/31331/312$0.270
222702_ecf75d
wrenfield
1K · minimal thinking2026-10-06.1
render-v1
3131/31031/310$0.058
222738_675707
wrenfield
1K · thinking on2026-10-06.1
render-v1
3131/31031/310$0.071
222824_caa176
wrenfield
1K · thinking on2026-10-06.1
render-v1
3131/31031/310$0.079
222929_0174ec
wrenfield-full
2K · thinking on2026-10-06.1
render-v1
5756/5788interrupted$1.102
225247_03fff1
wrenfield-full
2K · thinking on2026-10-06.2
render-v2
5755/57256/573$0.417
225829_4e73db
harrowfield
2K · thinking on2026-10-06.2
render-v2
109/1039/103$0.548
230652_c5b051
harrowfield
2K · thinking on2026-10-06.3
render-v3
1010/10110/101$0.146
231852_440ed7
wrenfield
2K · thinking on2026-10-06.4
render-v3
3131/31031/310$0.087
231949_a3b9fa
wrenfield-draft
1K · minimal thinking2026-10-06.4
render-v3
3130/31030/313$0.135
232749_82f43f
wrenfield-draft
1K · minimal thinking2026-10-06.5
render-v3
3131/31031/310$0.056
001353_926ed2
wrenfield
2K · thinking on2026-10-06.6
render-v3
3131/31031/310$0.096
001501_f5ccf0
wrenfield-draft
1K · minimal thinking2026-10-06.6
render-v3
3129/31030/31
corrected · receipt
1$0.167
001645_7db54a
alderbank
2K · thinking on2026-10-06.6
render-v3
99/909/90$0.077
002103_932e88
wrenfield-draft
1K · minimal thinking2026-10-06.7
render-v3
3130/31131/312$0.232
015012_50d896
wrenfield-draft
1K · minimal thinking2026-10-06.8
render-v3
3131/31031/310$0.075

Cost = every API call the run made (render, reads, strict checks, edits, re-reads, any full re-render), summed from the run's own ledger lines. Controls are shared across runs and counted separately. Engine = the rules version that produced the row: the 9/10 Harrowfield row is a false flag from engine 2026-10-06.2 (its strict reader wrote “$18.6” for “$18.6 million”); the gate now confirms by digits, and the re-run is 10/10. “Corrected” = a later, stricter rule re-scored the run's own stored reads, and its receipt shows both results. “Interrupted” = stopped before a receipt; its spend still counts.

All API spend so far: $4.52: runs $3.78 (17 finished, 1 interrupted at $1.10), controls $0.50, /verify $0.11, the rest probes and experiments.

Repair prompts, ranked

— attempts

No repairs logged yet.

Controls (latest)

—

Controls run with the first live render.

Measured cost per call

median
Render · 1K · minimal thinking$0.0485 (n=7)
Render · 1K · thinking on$0.0609 (n=2)
Render · 2K · thinking on$0.0729 (n=9)
Masked edit · 1K · minimal thinking$0.0454 (n=8)
Masked edit · 2K · thinking on$0.0596 (n=24)
Blind read$0.0027 (n=132)

.data/llm-calls.jsonl (successful calls; price from each call's own usageMetadata)