Judging AIMay 13, 2026·5 min read

How to Check a Number an AI Gave You

Figures are the highest-risk thing AI produces and the thing people check least, because a number looks like a fact. Four checks that take under a minute each, and the one that catches the most.

By Patin Team · Examples are illustrative composites

Prose gets read sceptically. Numbers don't, and that asymmetry is the single most exploitable thing about how people review AI output.

A sentence that sounds off invites a second look. £4.2m, up 19% year on year invites none, because a figure carries the shape of a fact — it looks like something that was measured rather than something that was written. It reads as retrieved even when it was generated.

Figures are also where the consequences live. A slightly wrong adjective is a style note. A slightly wrong number in a board paper is a decision made on the wrong basis.

The four checks

Where did it come from? Not "what's the source" — the model will name one. Which specific document, table, and row. A figure that can't be traced to a location isn't sourced; it's attributed, and attribution is generated the same way the number was.

Is it about the thing you asked about? The most common real failure and the least dramatic. The number is genuine, the source is genuine, and it's measuring an adjacent thing: a different period, a different region, a different definition of the same metric. Revenue and bookings. Users and active users. Financial year and calendar year.

Does the arithmetic hold? Percentages that don't total, a growth rate inconsistent with the two figures it sits beside, a subtotal that isn't the sum of its parts. Models are producing plausible text about arithmetic even when they can compute — and if you can't check it in your head, that's a reason to check it, not a reason to accept it.

Is it the right order of magnitude? The cheapest check and the one that catches the worst errors. You know roughly what your own numbers look like. A figure that's out by a factor of ten is usually obvious in one second to someone in the domain, and invisible to everyone else in the room.

The check that catches the most

Of the four, "is it about the thing you asked about?" finds the most real problems.

Order-of-magnitude errors are caught by anyone paying attention. Fabricated figures are increasingly rare when the model has the source in front of it. What survives is the definitional mismatch — a real number about a slightly different thing, which passes every other check because everything about it is true except its relevance.

So when you trace a figure back, don't just confirm it exists at the source. Read the column header. Read the footnote about what's included.

Citations are numbers too

Everything above applies to references. A citation is a factual claim with a specific format, and formats are exactly what models produce well.

The failure has shifted over the past couple of years — wholly invented sources are less common now than real sources that don't say what they're cited as saying. Which means checking a citation exists is no longer a check. You have to open it and find the claim.

If that sounds like too much for every reference, it is. Which is why the checking effort should scale with what the figure is doing: a supporting detail in an internal note gets a glance, a figure in something going to a client or a regulator gets opened.

What to ask for instead

Some of this is preventable upstream. Two requests change what you get back:

For every figure, give the source document and the specific table or section it came from.

If you can't find a figure in the sources provided, say so rather than estimating.

The second one matters more than it looks. Without it, a model asked for a number will produce a number, because that's what was asked for. With it, "not present in the provided documents" becomes an available answer — and it's the answer you want in exactly the cases you'd otherwise get burned by.

Sasha — the year that wasn't the year

Sasha is a finance business partner at a manufacturer. A market-sizing summary put a competitor's revenue at a figure she found high but not implausible.

Traced back, the number was real and correctly transcribed from the company's own filing. It was the group figure, and she'd asked about a division. The model had found the most prominent revenue number on the page.

Nothing about it was fabricated. It was simply about a different thing than the sentence around it claimed, and every check except that one would have passed it.

Tobias — the sentence that asked for permission to fail

Tobias runs research at a trade body. His reports were mostly sound, with an occasional figure nobody could locate afterwards.

The change was one line added to every request: if a figure isn't in the sources provided, say so rather than estimating.

The reports got shorter and slightly less satisfying to read — a few sentences now end in "not available in the provided sources". His view is that this is the report being honest about what it knows, and that the previous version had been filling those gaps silently.

The one thing

Numbers read as retrieved and are often generated. Check where it came from, whether it's about the thing you asked about, whether the arithmetic holds, and whether the magnitude is sane.

The definitional mismatch is the one that survives casual review — a real figure about a slightly different thing. That's the one to look for.

Reading about it only gets you so far

Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.

Just want the writing? .