Judging AIJuly 30, 2026·4 min read

AI Skills for Researchers: Speed Up Discovery Without Cutting Corners

AI can read more than you can and is confidently wrong about what it read. For researchers, that's a specific and manageable risk. Here's how to get the speed without inheriting a fabricated citation.

By Forge Team · Examples are illustrative composites

Research has a property that makes AI both unusually useful and unusually dangerous: the output looks the same whether it's right or not. A synthesis of twelve papers and a confident summary of twelve papers the model half-remembers are formatted identically, written in the same register, and equally pleasant to read.

That's not a reason to avoid it. It's a reason to be specific about which parts of research it should touch.

Where it's genuinely strong

Triage. Given forty sources, which eight are worth your time? Excellent — and low-risk, because you're going to read the eight anyway.

Extraction from a source you've supplied. "What method did this paper use, what was the sample, what did they conclude?" Reliable when the text is in front of it, because it's reading rather than recalling.

Cross-cutting. "Across these six documents, where do the authors disagree?" Genuinely hard by hand and well-suited to a machine.

Structuring. Turning a pile of notes into an outline. Mechanical.

Where it fails, and how the failure looks

Citing from memory. Ask what the literature says on a topic and you may get real papers, real authors attached to papers they didn't write, and papers that don't exist — in one list, in one voice. The fabrications aren't sloppy; they're plausible, with realistic titles and credible author names.

The rule that handles this: if the model wasn't given the source, treat every citation as a hypothesis. Not a fact to verify at leisure — a hypothesis, until you've seen it.

Currency. Models have a knowledge cutoff and are frequently vague about where it falls. For anything time-sensitive, "as of when?" is a question worth asking explicitly, and worth distrusting the answer to.

Confidence flattening. A finding replicated across twenty studies and a single suggestive result get described in similar language. The strength of evidence is exactly the thing that should survive summarisation, and it's the thing most often lost.

Skill: separate reading from recalling

The most useful distinction in AI-assisted research is whether the model has the text.

Given the text, it's a reading assistant, and it's good at that. Without the text, it's producing a reconstruction from training data, and the quality of that reconstruction is unknowable from the output.

Structure your work so it's almost always in the first mode: supply sources, ask about what you supplied. When you can't, treat the answer as a list of leads.

Tara — the citation that got through

Tara is a policy researcher at a think tank. She used AI to draft a background section, checked the sources, and found them fine — except one, which she'd skimmed because the author was someone she'd read before.

The paper didn't exist. The author was real, the topic was one they worked in, the year was plausible, and the title was exactly what such a paper would be called. It got as far as an internal review before a colleague tried to download it.

Her rule since: every citation gets opened before it goes in a document, without exception, including the ones from authors she trusts. Especially those — familiarity is what let it through.

Miguel — the disagreement he'd have missed

Miguel researches labour markets for a sector body. He was working through fourteen reports on workforce automation and used AI to map where they disagreed.

Most of the disagreements were ones he'd have found. One wasn't: two reports using apparently identical methodology reached opposite conclusions about a sector, and the difference came down to how each defined "routine task" — buried in an appendix in one and a footnote in the other.

He'd read both. He hadn't compared the definitions, because they were in different places and he'd read the reports three weeks apart.

That's the class of thing worth delegating: not judgement, but the mechanical comparison across more documents than fit in working memory.

The one thing

The distinction that matters more than any other: does the model have the source in front of it, or is it recalling?

Given the text, it's a fast and capable reader. Without it, it's generating something that looks like research. Both produce the same formatting, and only one is worth citing.

Reading about it only gets you so far

Forge turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.

Just want the writing? .