How to Verify Work You're Not Qualified to Check
AI's most attractive use is work you couldn't do yourself — which is exactly where you can't assess the result. There are real checks available, and they're structural rather than technical.
By Patin Team · Examples are illustrative composites
The most tempting use of AI is the work you couldn't do yourself. A marketer producing a data analysis. An operations manager drafting something quasi-legal. A founder writing a technical specification.
It's also the use where verification is hardest, because the thing that makes the output valuable — you lack the expertise — is the thing that makes it uncheckable. And the failure is silent: without the domain knowledge, a wrong answer and a right one look the same.
There are real checks available. None of them require you to acquire the expertise, and all of them are structural rather than technical.
What you can check without knowing the field
Whether it engages with the actual specifics. Domain-correct output refers to the particulars of your situation. Generic output refers to the category. An analysis of your data that would read identically applied to anyone else's data hasn't looked at yours.
Whether it names conditions and exceptions. Real expertise is full of "unless" — unless the contract predates the change, unless the sample is skewed, unless you're in a regulated sector. Output with no conditions in it is either a simple question or a shallow answer, and you can usually tell which by whether experts in the field argue about it.
Whether it acknowledges a trade-off. Almost every meaningful decision in every field involves giving something up. An answer that's straightforwardly good with no cost is describing something simpler than the situation.
Whether it's internally consistent. You can check whether a document contradicts itself between page one and page four without knowing anything about the subject. This catches more than you'd expect, and it's completely accessible.
Whether the confidence is even. In any real analysis some parts are firmer than others. Uniform certainty across claims of obviously different kinds means the uncertainty hasn't been tracked.
The three questions that work in any field
"What would an expert in this field disagree with here?" The most useful single question available to a non-expert. It surfaces where the contested ground is, which tells you where to be careful — and an answer that claims nothing is contested is itself informative, because something always is.
"What would you need to know to be more confident?" A good answer names specific missing information. A weak one repeats generalities. This also frequently reveals that the analysis rests on an assumption you could have corrected in one line.
"What's the standard approach here, and is this it?" Deviations from convention may be correct, but they should be deliberate. If the output has quietly departed from what's normally done and doesn't say why, that's worth someone's attention.
Match the check to the stakes
Beyond a certain level of consequence, none of the above is sufficient, and it's worth being honest about where that line sits.
Low stakes — internal, reversible, will be corrected in conversation. The structural checks are proportionate on their own.
Medium — goes to a client, informs a real decision, embarrassing if wrong. Structural checks plus one external cross-reference: a reputable published source, a professional body's guidance, anything not produced by the same model.
High — legal, financial, medical, safety, regulatory. A person qualified in the field reads it. There is no substitute available here, and the availability of a fluent draft doesn't change that. AI got you a better-prepared starting point and shortened their time, which is a genuine and sufficient benefit.
The mistake is letting the fluency of the draft move a task down a category. A well-written contract clause you can't evaluate is not lower risk than a badly written one — it's higher, because you're less likely to seek help.
The failure that isn't about accuracy
One thing worth flagging separately: the most common problem with expert-adjacent AI output isn't an error, it's an omission.
The analysis is correct as far as it goes and doesn't mention the consideration that anyone in the field would have raised immediately. You can't detect that by reading harder, because there's nothing on the page to catch. It's the strongest argument for the "what would an expert disagree with?" question and for one conversation with someone who knows, however brief.
Amara — the analysis with no exceptions
Amara runs marketing at a subscription business and used AI for a cohort retention analysis, well outside her training.
Her check wasn't statistical. She noticed the output contained no conditions at all — no "assuming", no "unless", no caveat about the data. Her instinct was that anything real about cohort analysis has caveats.
She asked what an analyst would disagree with. The response named three things, one of which was a seasonality effect that changed the conclusion materially. Her comment: she still can't do the analysis, but she now knows the shape of an answer that's been thought about.
Callum — the clause that read well
Callum is a founder at a small company. AI drafted a contract clause that read considerably better than the version he'd have written himself.
He nearly used it, and his reason for not doing so was structural: it contained no exceptions, and every clause he'd ever seen a lawyer write was full of them.
An hour of actual legal review found it enforceable but far broader than intended — it committed him to something he hadn't meant to promise. His summary: the quality of the writing had almost moved a high-stakes task into the low-stakes pile, which is exactly the wrong direction.
The one thing
You can't check the domain, but you can check the shape: does it engage with your specifics, name exceptions, acknowledge trade-offs, stay internally consistent, and vary its confidence?
Ask what an expert would disagree with. Scale the check to the consequence. And don't let a well-written draft persuade you that a high-stakes task is a low-stakes one.
Put this into practice
Reading is a start — but skill comes from doing. Try these drills now.
Reading about it only gets you so far
Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.
Just want the writing? .
Keep reading on this
OpenAI Just Warned That Its Own Agent Makes 'Finished Mistakes.' Here's How to Catch Them.
OpenAI's own safety guidance for ChatGPT Work admits the system can produce 'finished mistakes at a scale that is harder to catch.' The same week, Wharton research killed prompt tricks and Cursor data showed a 46x productivity gap. The differentiator is now verification, not prompting.
6 min readAI Makes Up URLs. Attackers Are Registering Them.
Unit 42 tested 685,339 prompts and found 2.1 million AI-generated URLs — 13,229 already live and malicious. In one case, researchers predicted which domain an AI would hallucinate. Twenty-three days later, an attacker registered it and deployed a phishing kit.
4 min readFord Spent Billions Learning What a German Court Just Made Law: Your AI's Mistakes Are Your Mistakes.
Ford rehired 350 engineers after AI quality inspection produced billions in costs. The same week, a Munich court ruled AI-generated false claims are the deploying organisation's own speech. The gap between 'the AI made an error' and 'you submitted it' is closing.
5 min read