Supervising AgentsJune 20, 2026·5 min read

What an Agent Needs to Know That You Never Wrote Down

Most agent failures aren't reasoning failures. They're handover failures — the agent didn't know a thing everyone in your team knows and nobody has ever typed. Here's how to find those things before they cost you.

By Patin Team · Examples are illustrative composites

When an agent produces something wrong, the natural diagnosis is that it reasoned badly. Usually it didn't. It reasoned perfectly well from a picture of the situation that was missing something everyone on your team has known for years and nobody has ever written down.

This is a handover problem, and it's the same one that makes a competent new starter's first month expensive. The difference is that a new starter asks. An agent doesn't — it fills the gap with something plausible and proceeds.

The four things that are never written down

Which sources are actually authoritative. Your team knows the spreadsheet is right and the dashboard lags by a day. That's not documented anywhere, and an agent given both will treat them as equally true — or pick the one that's easier to read.

What the exceptions are. Every process has them. This client gets the report early. That category is always reviewed by a person. Exceptions are the part of a process that lives entirely in people's heads, because the documented version is the version without them.

What "done" means here. Not the task definition — the local standard. Whether a summary means one page or three. Whether "check the numbers" means recalculating them or confirming they match the source. Two reasonable people disagree about this within any given team; a model has no basis to guess.

What to do when something's missing. The single highest-value line in most briefs and the one almost nobody includes. Told nothing, an agent will improvise: use the adjacent data, estimate, proceed with a caveat buried in paragraph four. Almost always you'd rather it stopped and said so.

How to find them without writing a manual

You cannot document your job. You can do something much cheaper.

Run it and read for surprise. Give the agent a small real task and watch for every point where it does something you didn't expect — including reasonable things. Each surprise is a piece of context you didn't know you were holding. Three or four short runs surface most of what matters, and they surface the parts that matter to this task, which a general write-up wouldn't.

Do the last-mile question. Before a run, ask yourself: if a capable contractor started this on Monday with the brief as written, what would they have to ask by lunchtime? Those questions are the missing context, and you can usually list four in ninety seconds.

Keep a running exceptions note. Not a document — a list, in whatever you already use. Every time you catch yourself thinking "well, obviously, except when…", that's a line. It grows to about a dozen for most tasks and then stops.

Say what to do when it's stuck

Worth stating on its own, because it changes the failure mode more than anything else on the list.

If the source you need isn't available, or the data looks inconsistent with what I've described, stop and tell me rather than working around it.

Without that, the model's default is to be helpful, and helpful under uncertainty means substituting. With it, an unrecoverable run becomes a two-line message and a five-minute fix.

The context that shouldn't be handed over

Some of what you know shouldn't go in the brief at all — because it's confidential, because it's about a person, or because it's the judgement you're paid for rather than an input to it.

That's a real limit, and the right response isn't to work around it. If a task can't be done without context you can't hand over, the task isn't a delegation candidate. That's a finding, not an obstacle.

Malik — the four questions

Malik is a project manager at a construction consultancy. His brief for a weekly status agent looked complete, and the outputs were subtly wrong for a month.

He tried the contractor test: what would a capable person have to ask by lunchtime on day one? He got four questions in under two minutes. Which system is authoritative when they disagree. Whether "at risk" means the date or the budget. Whether to include the subcontractor items. What to do when a lead hasn't updated their section.

All four had been silently answered by the agent, and three of the four answers were wrong. Adding the four lines fixed a month of drift.

Yuki — the line she added to everything

Yuki leads a data team at a retail group. Her recurring failure was substitution: agents that couldn't reach a source used an adjacent one and mentioned it in passing, somewhere nobody read.

She now adds one line to every agent brief — if the source you need isn't available or the data contradicts the brief, stop and tell me rather than working around it.

Her description of the effect: it converted a category of expensive silent failures into cheap loud ones. The agents now stop about once a fortnight, and each stop is a two-minute fix instead of a discovered mistake.

The one thing

Agents mostly fail on handover, not reasoning. What's missing is the authoritative source, the exceptions, the local standard for done, and what to do when something's absent.

You won't find those by writing documentation. You'll find them by running small tasks and reading for surprise — and by telling the agent, explicitly, to stop rather than improvise.

Reading about it only gets you so far

Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.

Just want the writing? .