The first time an AI tool invents something at you, it is unsettling. Not because it is wrong, everything is wrong sometimes, but because of how it is wrong. There is no hedging, no "I think", no change in tone. A confidently invented statistic reads exactly like a correct one, in the same measured prose, with the same air of having checked.
This is not a bug that will be patched away, and it is not evidence the technology is useless. It is a direct consequence of how these tools work, and once you know where it tends to happen you can catch nearly all of it in about two minutes.
The short version
- The model produces plausible text. Usually plausible and true overlap. When they do not, nothing signals it.
- Errors cluster in predictable places: names, numbers, dates, citations, and anything niche.
- Work it can see is far safer than work it must recall.
- Two minutes of checking removes almost all the risk. The failures are all cases where nobody spent them.
Why it happens
A large language model answers by working out what text most plausibly comes next, over and over. It is not looking anything up. There is no internal store of facts to consult and no mechanism that distinguishes "I have seen this a thousand times" from "I am filling a gap that shape".
So when you ask about something it read constantly, the pattern is strong and the answer is right. When you ask about something rare, the pattern is thin, and it produces what the answer would look like if it existed. A case reference in exactly the right format. A statistic with a plausible figure and a plausible source. A quotation in the right voice.
That last point is the dangerous one. It fails most convincingly precisely where you are least able to check, because your question was niche in the first place. If this mechanism is new to you, What is a large language model? covers it properly.
Where it tends to happen
| Risk | Task | What to do |
|---|---|---|
| High | Statistics, research findings, market sizes | Assume invented until you find the source yourself |
| High | Legal references, case names, regulation numbers | Verify against the actual register or take advice |
| High | Quotations attributed to a named person | Find the original or cut it |
| High | Anything about your specific sector's niche rules | Thin training data; check with your trade body |
| Medium | Dates, thresholds, deadlines, tax figures | Check against GOV.UK; these change yearly |
| Medium | Which supplier or product does what | Check the vendor's own current page |
| Low | Summarising a document you pasted in | Skim for anything not in the original |
| Low | Rewriting, restructuring, tone changes | Read it. You would anyway. |
| Low | Brainstorming, first drafts, outlines | Nothing. Being wrong is fine at this stage. |
The pattern is simple: the more the answer depends on recall, the more you check. The more it depends on material you supplied, the less you need to.
The two-minute check
Run this before anything leaves the business. It is quick because it only looks at the four things that go wrong.
- Circle every number, name and date. Those are your error sites. Verify each one at source, not by asking the AI again.
- Test one citation. If it gave you three references and the first is real, the others usually are. If the first does not exist, treat the whole answer as unreliable and start again.
- Ask what it is least sure of. Genuinely useful. "Which parts of that answer are you least confident about, and which should I verify independently?" It will often flag its own weakest claim.
- Read it as the recipient. Not for accuracy, for judgement. Does it say something we would not say? Commit us to something we cannot do? This catches the tone failures a fact-check misses.
Go back over your last answer. List every factual claim you made that I would need to verify independently, and mark each one: confident, uncertain, or I am inferring this. Do not defend the answer, just mark it up.
How to make it happen less
Give it the source. The most reliable fix there is. Paste the policy, the contract, the report, and ask it to work from that. It cannot invent a clause that is not in a document it is reading.
Say what to do when it does not know. Models are trained to be helpful, and helpfulness pulls toward answering. Adding "if you are not certain, say so rather than guessing" to the end of a prompt works better than it has any right to.
Turn on web search for anything current. If the tool can look it up, it is reading rather than recalling. Then check the link it gives you actually says what it claims.
Ask for the working, not just the answer. Reasoning is easier to audit than a conclusion. You will often spot the wrong turn without knowing the right answer.
Never ask it to check itself in the same breath. "Are you sure?" frequently produces a confident apology followed by a different wrong answer. Verification has to come from outside.
The two failure modes to avoid
Businesses go wrong in two opposite directions here, and both are expensive.
The first is trusting it because it sounds right. This is the one that ends in a solicitor's letter. The prose is polished, so the thinking must be. It is a reasonable instinct built on decades of experience: in human writing, care with commas usually correlates with care with facts. That correlation does not hold here.
The second is abandoning it because it was wrong once. A colleague who is right most of the time and needs their figures checked is still a useful colleague. You would not sack a bright graduate for one bad statistic. You would check their figures.
The workable position is in between and it is not complicated: use it heavily where you can judge the output, check the four things that go wrong, and never let anything leave the building that no human has read.
What good practice looks like
In businesses that get real value from this, the checking is invisible because it is habitual. Numbers get verified as a reflex. Nobody sends a first draft. There is a name at the bottom of every document and that person read it. The one-page policy says as much in a single line, and there is a template for it in Write an AI policy for your small business in one page.
Watching someone check their own AI output in real time is more instructive than reading about it. That is a large part of what the free AI Breakfast Club webinar is: Gary works on screen, live, every other Friday morning, and leaves the mistakes in on purpose.