What is the AI actually doing? / 2. Hallucination

"That didn't happen at all." Why AI hallucinates

The answer has a date, a version number and a tidy explanation. It even sounds as though the model checked. Then you open the folder and the thing it described is not there.

The important mistake is believing the answer has already been checked

Hallucination is commonly used for generated material that is false or unsupported: an invented reference, a wrong fact, a detail added to fill a gap. In business work, the invented detail can be about the project itself. A file supposedly exists. A review supposedly happened. A draft supposedly includes the latest correction.

One of the founder's direct corrections in the GR archive was to reanchor an account of document versions in the actual folder. That is a useful instinct. The generated description and the thing being described need to be compared.

The risk increases when an unsupported answer becomes something another person relies on: a decision, an instruction, a customer statement or the assumed starting point for more work.

Why a capable model can still produce a false answer

Generating a plausible answer and establishing that each claim is true are different tasks. OpenAI's research argues that common training and evaluation practices can reward guessing over admitting uncertainty. A system can therefore improve at answering many questions while still producing confident errors.

The immediate source of a mistake can also be an ambiguous question, missing evidence, an outdated source or a faulty interpretation of retrieved material. A citation helps only if it exists, is relevant and supports the claim attached to it. Reading the cited page is part of the check.

Agreement can add another layer of false reassurance. OpenAI documented a GPT-4o update that became excessively agreeable and rolled it back. An answer that validates your preferred conclusion is not independent confirmation of that conclusion.

Make the claim checkable

For the statements that matter, ask what supports them and inspect that support. Separate what the source says from the interpretation being proposed. Leave an unresolved question visibly unresolved when the material cannot answer it.

For calculations, use a calculation and check the inputs. For a claimed file change, open the resulting file. For a claimed external action, inspect the receiving system or its record. A second paragraph from the same model saying it is confident supplies another statement, not the missing evidence.

A second reviewer can help, especially when given the source material and an explicit question. Repeating the same unsupported premise to another model can also repeat the mistake. The review needs a way to reach the evidence.

Give AI a useful job even where certainty is limited

AI can prepare a draft, organize evidence, compare versions and identify missing information while uncertainty remains. Define which of those jobs is underway, and what would justify relying on the result.

Use proportionate checks at the point where an error would matter. You do not need to verify every phrase of an exploratory note as though it were a public commitment. You do need to know when a phrase is about to become one.

Good Remedy began asking whether its own impressive-looking development was supported. That question belongs inside the work: what exists, what has been checked and what can another person reasonably use now?

Put the idea to work

For AI sales research flowing into a CRM, use the handoff guide to connect claims, evidence and the next decision. The implementation guide explains what to ask a delivery partner to demonstrate.

Published by Good Remedy. First-engagement pricing is introductory where shown; standard pricing applies after the introductory engagement. Unusual complexity is confirmed before work begins. Examples are illustrative unless explicitly identified otherwise.