Judging AIMay 24, 2026·5 min read

Reviewing AI Output: What to Actually Check

Reading it and approving it isn't a review — it's the exact filter AI output is best at passing. Interrogating is a different activity, it takes about ninety seconds, and it's most of what separates the people getting value from AI.

By Patin Team · Examples are illustrative composites

Most people review AI output by reading it. If it reads well, it goes forward.

That isn't a review. Reading well is what these systems are best at — it's the one property that's reliably present regardless of whether the content is correct, complete, or appropriate for its audience. Reviewing for plausibility filters on exactly the dimension the output is strongest.

Interrogating is a different activity. It takes about ninety seconds once you know what you're looking for, and the difference between the two is most of what separates professionals getting consistent value from AI from those getting occasional wins and a lot of rework.

Set the standard before you generate

The reason reviews take so long is that people are deciding what "good" means while looking at a specific attempt. That's slow, and it's biased — you end up assessing the output against itself.

Three criteria, written before the prompt, changes the economics. Not "make it good", but: what claim must each section support, what evidence has to be present, what's the format constraint. Then the review is a comparison rather than a judgement, and comparisons are fast.

People who do this consistently report review times dropping by more than half. The output doesn't improve. The standard gets clear enough that output either meets it or doesn't.

The five questions

Where did this number come from? Aggregates and percentages are where quiet errors live. Trace one per section.

What is this claim's mechanism? For any causal statement, can you say why A would produce B? If not, you have a pattern the model noticed, not a finding you can act on. This is the question that catches "customers who contact support churn more" — true, and backwards.

What's missing that I'd expect? The hardest one, because absence is invisible. A model asked to cover a topic covers it evenly; it has no view on what matters most. Qualifiers, caveats, and the inconvenient counter-example are what disappear first.

Is this stated with more certainty than the source warrants? Confidence flattening is the most consistent failure in summarisation. A finding replicated twenty times and a single suggestive result come back described in the same register.

Whose framing is this? Models mirror the framing of the prompt. If you asked whether to do X, you'll get a case for X. If the output agrees with everything you implied, you haven't received analysis — you've received your own assumptions with better formatting.

When it agrees with you, that's data

Push back on an AI's answer and it will frequently concede. Not because you're right — because agreement is the path of least resistance in how these systems are trained.

That means agreement carries almost no information, and a reversal under pressure carries a lot: it tells you the original position wasn't held for a reason.

The practical move is to stop asking "is this right?" and start asking it to argue the other side. What's the strongest case against this? What would have to be true for the opposite conclusion? You learn more from the answer than from any amount of agreement.

Use disagreement, not consensus

Running the same question through two models is often described as fact-checking. It isn't — they share training data and failure modes, and both can be confidently wrong in the same direction.

What it's genuinely good for is finding divergence. Where two models disagree, the question has more than one defensible answer, and that's precisely where your judgement needs to go in rather than your editing.

Nadia — the ninety-second review

Nadia is a content marketing manager at a 35-person B2B software company. She reviewed AI drafts by reading them for three months, and her review times were getting longer rather than shorter.

She now writes three criteria before any draft starts: the specific claim each section must support, one piece of evidence required per claim, and no more than two supporting sentences per point.

Review time went from 40 minutes to 12. She's clear that the drafts didn't get better — she just stopped deciding what good meant while looking at each one.

Adaeze — the framework that had no ground under it

Adaeze is a senior analyst at a consulting firm. Individual AI outputs on her team looked fine, which is why the problem took a while to surface.

A colleague presented a client framework that seemed analytically solid. It had been derived entirely from AI synthesis of secondary sources — no original analysis, nothing checked against a primary source. Every step was defensible. The whole thing rested on nothing.

She now runs important questions through two models before using the output — not to pick a winner, but to find where they diverge. Divergence marks the places her own judgement has to go in.

The one thing

Reading it isn't reviewing it. Reading is the check the output is designed to pass.

Set the standard before you generate, then ask five questions of what comes back: where's the number from, what's the mechanism, what's missing, is it overconfident, and whose framing is this? Ninety seconds, and it's the skill that stops being optional the moment your name is on the output.

Reading about it only gets you so far

Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.

Just want the writing? .