Skip to content
FactFrame

Public demo: every run here is a recorded replay of a real run, labelled with when it ran. Nothing on this site calls a model. Live runs are by request during early access.

How it works

The method, step by step, with the evidence for each choice.

FactFrame is built around what was measured about Nano Banana 2.1 and its reader on 2026-10-06, including the ways the reader gets things wrong.

  1. 1

    Ingest the PDF

    pdfplumber extracts each page's text (pypdf is the fallback if pdfplumber fails on a file). Every number is found with its kind (amount, percent, year, date, count), its page, and the exact sentence or table row it sits in. Running headers, footers and headings are skipped.

    Why: A fact you can't point to on a page can't be defended.

  2. 2

    Pick and ground the facts

    You choose the facts to feature and their captions. Each one must appear verbatim on its page; a caption may not contain digits, because every number on the image has to be a checked fact. Caption words are scored against the source sentence.

    Why: A fact that isn't in the PDF is refused, not “fixed”.

  3. 3

    Render with Nano Banana 2.1

    One call, 2K, portrait 4:5 or 16:9, thinking left on, with the facts given as an exact, unnumbered list and an instruction to add no other number anywhere.

    Why: Google's model card scores Infographic Factuality at 0.521 with thinking and 0.328 without (scale undefined), so thinking stays on.

  4. 4

    Read it back, blind, twice

    gemini-3.8-flash (thinking low; sampling at the model's defaults, since Google says custom temperature has had no effect since Gemini 3.6 Flash; runs before engine 2026-10-06.4 sent temperature 0, which per Google changed nothing) transcribes every text element, one per line, with each figure's caption joined onto its number's line. If the two reads disagree on any number, a third read decides by majority.

    Why: On 2026-10-06, 18% of transcripts differed between two reads of the same image (that test sent temperature 0, which per Google changes nothing), and the reader split two-line text. One read is not a measurement.

  5. 5

    Match each number to its source

    One canonical rule: separators are formatting (“1,234” = “1234”), scale words expand exactly (“$4.5 million” = “$4,500,000”), nothing is ever rounded, kind and currency must match, and a comma in the wrong place is a visible typo. A number counts as on its caption only if the caption's words are on its line; a right number under the wrong caption is wrong (caption). Each number gets a status: match, wrong, missing, unsourced or damaged.

    Why: “Close enough” is how a wrong figure gets published.

  6. 6

    The strict gate

    A second prompt lists every number and says whether any character is damaged. A number is a match only when a strict read lists it as clean; otherwise a blind crop of its region is checked. The gate is trusted only after it passes its controls. If the gate can't run (its controls failed, or the reader fell back to tesseract), it clears nothing: every number is listed as unconfirmed and the run can't be all clear.

    Why: Readers forgive broken digits, and over garbled regions they can confabulate a fluent, plausible number. Checking the image, not the transcript alone, is the point.

  7. 7

    Repair by masked edits

    Each non-match gets one edit of just its region: the image, then a mask (white = may change), then the instruction (two prompt variants per kind of fix, picked by their record). A missing figure with nothing to sit next to is added without a mask, and kept only if the change is compact. The mean pixel change outside the region is measured against a JPEG re-encode floor: local at most 2× the floor + 0.25, leaky above 4× the floor + 0.5 (rejected), partial in between. An accepted edit is spliced in by region, then the image is re-read and re-matched; an edit that breaks another number is rejected. At most three rounds; too many problems, a refused edit, or a figure that can't be placed means one full re-render instead, kept only if its check is better.

    Why: In FactFrame's own measurements, every one of 32 measured 2.1 edits changed pixels outside its region beyond the re-encode floor (median 2.66× the floor). Splicing keeps every other number exactly as it was checked.

  8. 8

    Write the receipt

    JSON, printable HTML and PDF: every number, the source sentence and page, its status before and after, the rounds it took, the controls, the reader's disagreement, every edit's pixel measurement, the API cost, and what could not be verified.

    Why: The receipt is the product. The repair loop is the safety net.

Controls that can fail

The checks are checked first.

Five controls on four canvases, five reads each, cached for 24 h per image size. Every receipt says whether its controls ran fresh or were reused, and when; if a control fails, the receipt says which numbers that leaves unverified.

ControlWhat must happen
Reader floorA clean canvas of eight figures: every read must return every number exactly (5 of 5).
Spell trap“41,38,0” and “1,4O2” drawn cleanly: the reader must return them as drawn, not correct them (4 of 5).
Strict, cleanThe clean canvas through the strict prompt: nothing flagged (4 of 5).
Strict, damaged digitsTwo numbers overdrawn with stray partial digits: exactly those two flagged (4 of 5).
ConfabulationTwo figures overprinted with other digits (2,412 under 3,817 and 5,4908; 88.4% under 63.1% and 9,6.4), so no single reading is legible: the full gate must never clear them as a number (4 of 5).

When something is down

Gemini down at renderThe stored run for that scenario, badged as a replay with the time it ran; never shown as live.
Reader downtesseract reads the numbers only; badged. The strict gate can't run, so every number is unconfirmed and flagged: the run can't be all clear.
A masked edit is refused, or a missing figure can't be placedOne full re-render with the corrected list, kept only if its check is better; the receipt says which.
Spend cap reachedThe run stops where it is and the receipt lists what is unverified.

How it gets better with use

Every repair attempt is logged: error class, prompt, outcome, pixels changed outside the region, cost. The prompt for the next attempt is chosen by Thompson sampling on each prompt's fix rate per error class, so a prompt that keeps failing stops being picked. See the evals.

Measured cost per call (median)

render 1K · minimal thinking: $0.049 (n=7) · render 1K · thinking on: $0.061 (n=2) · render 2K · thinking on: $0.073 (n=9) · blind read: $0.0027 (n=132) · a fresh set of controls: $0.066–$0.248 by image size

Median time: render at 2K 29.1 s (n=9), each blind read 5.7 s.

Sources for every figure on this page

  • 0.521 / 0.328 Infographic Factuality: Google DeepMind's model card for Nano Banana 2.1 (deepmind.google/models/model-cards/nano-banana-2-1/), read on 2026-10-06. The card does not define the metric's scale.
  • 18% of transcripts differed between two reads: the Signage Test (my own test of Nano Banana 2.1's text rendering), 18 of 100 images (results/RESULTS.md, “Reader noise on real images”), measured with temperature 0 sent (results/reader_config.json), which Google's notice says has no effect.
  • Readers forgive broken letters and digits: the same test's showcase render 02 (a broken “ODELIA FLOWERS” read back exactly), and a damaged-digit control read clean in 5 of 5 reads (the text-rendering campaign's controls).
  • Filling in garbled text: the text-rendering campaign's garbled billboards (campaign/results/LESSONS.md), where the reader returned fluent wording on one read, different wording on the next, and a wrong model number. In FactFrame's own confabulation control, across 30 reads of the two garbled slots, the reader returned the string overprinted on the slot 13 times and an invented or run-together one the rest (3,817, 35,4908, 35,4908817, 35,841708, 35,8419708, …); the gate cleared none. The control applies the gate's two clearing tests in a slightly simpler form (one strict read, no digit-string confirmation, the known slot box); under the production gate's rule none is cleared either.
  • Custom temperature has no effect: Google AI Studio's deprecation notice of 2026-10-06 (“Since Gemini 3.6 Flash, sampling parameters have been set to default values”).
  • Kanji repaired by the reader: the text-rendering campaign's controls (campaign/results/LESSONS.md §3).
  • Costs: medians of the successful calls in FactFrame's own ledger, each priced from its own token counts at Google's published per-token rates (archived with the date read in the project's research notes).
  • Edits reach outside their region: FactFrame's own repair log, 32 measured edits (.data/repairs.jsonl (outside_mean / re-encode floor, per measured edit)).

Limits, stated plainly

  • FactFrame checks numbers: amounts, percentages, years, dates, clock times and plain counts. It does not judge layout, icons or decorative text.
  • Units of measure are not compared: “11.4 km” and “11.4 m” read as the same number. A scale letter after a space is a unit (“186 M” is 186 m, not 186 million) unless it follows a currency amount (“$4.5 M”); glued to the digits (“4.2M”) it is a scale.
  • Latin-script text only. The reader corrected misspelled Japanese kanji in the text-rendering campaign, so non-Latin text is reported as unverified.
  • US number formatting (“,” groups thousands, “.” is the decimal point).
  • A scanned PDF with no text layer has no facts to ground; run OCR first.