AI hallucinations at work: how to catch invented details before you send
Learn what AI hallucinations are, what a real court case shows, and how to check names, numbers, sources, and promises before you send AI-assisted text.
You paste your notes from a client call into a chatbot and ask for a follow-up email. The draft reads well. It also contains a date, a figure, or a promise that you don't remember putting in your notes. Once you send it, the client reads each of those details as your statement. Nobody on their side knows which parts a tool wrote.
Here you'll learn what an AI hallucination is and what a real court case shows about who answers for one. You'll also get a list of details to check before AI-assisted text goes out, a worked example, and a way to draft so the facts come from you.
What an AI hallucination is
The US National Institute of Standards and Technology, NIST, covers this risk in its voluntary risk guide for generative AI. It calls the problem confabulation and notes that people commonly say hallucination or fabrication. The term describes a tool presenting false or wrong content with confidence. The definition also covers output that departs from your prompt or the material you gave it, and output that contradicts something the tool said earlier in the same conversation.
That second part matters most in everyday work. When you ask for a summary or a rewrite, the tool can add things that were never in your notes. A smooth sentence and a sure tone say nothing about whether a claim is correct. NIST also points out that invented reasoning or citations can make false output look more trustworthy.
The risk starts when someone trusts the output and acts on it or passes it on. Before that, a false detail is just a line in your draft that you can delete.
Fake citations in Mata v. Avianca and why the lawyers were penalized
A well-documented example comes from a 2023 decision by a US federal court. In Mata v. Avianca, lawyers filed court decisions that did not exist, complete with fake quotations and citations that ChatGPT had generated.
The opposing lawyers and then the court questioned the cases, and the lawyers kept relying on them. The court found bad faith, based on conscious avoidance and on false or misleading statements. It ordered the lawyers to pay a joint $5,000 penalty and to notify their client and each judge falsely named as the author of a fake opinion.
The fake citations started the problem. The bad-faith finding also rested on what the lawyers did after the cases were questioned. At the same time, the court said there is nothing inherently improper about using a reliable AI tool for help. Lawyers still have a gatekeeping duty to make sure their filings are accurate.
That duty applies to lawyers filing in court, and the case doesn't set a rule for every job. It still offers two practical lessons:
- A tool generated the false content, and the lawyers answered for it.
- When someone questions a detail, go back to the source and correct it if it's wrong. In Mata, standing by unchecked material made things worse.
What two studies found about AI help and missed errors
None of this means AI makes work worse. In a randomized experiment with 758 consultants at Boston Consulting Group, consultants using GPT-4 on 18 tasks designed to suit the model completed 12.2 percent more tasks. They finished them 25.1 percent faster on average and received higher quality scores. On one strategy task designed to fall outside what the model could handle, consultants using AI were 19 percentage points less likely than the control group to get the right answer.
That 19-point drop isn't an error rate for AI in general. It came from a single task, which the authors call a limit of their study. The experiment also used GPT-4 as it was at the end of April 2023, and the results come from consultants doing consulting work.
A 2025 survey of 319 knowledge workers looked at critical thinking in AI-assisted work. Everyone in it used generative AI at work at least weekly, and together they shared 936 first-hand examples from their jobs. Higher confidence that AI could handle a task went along with less reported critical thinking. The survey relied on self-reports from fluent English speakers, and the sample leaned toward younger, more technically skilled users. So the result is an association and doesn't prove that AI made anyone think less. The authors describe the work shifting toward checking information, fitting the answer into the task, and overseeing the result. In their account, responsibility for the work stays with the person.
One conclusion seems fair, though neither study tested it directly. Checks slip most easily on tasks where you feel sure the AI has it covered, so make checking a fixed step.
How to check AI-assisted text before it goes out
NIST recommends that organizations compare generative AI output with known facts, apply fact-checking methods, and review and verify sources and citations. For a single email, summary, or report, that comes down to three steps.
- Separate the wording from the facts. Keep the phrasing if you like it, and treat every factual detail as unconfirmed until you've checked it.
- Check each detail against the original material, such as your notes, the contract, the price list, or the source document. In a summary, anything you can't find in the source is an addition. Remove it or confirm it elsewhere.
- Open every link and cited source yourself. Make sure it exists and says what the draft claims. Asking the same chatbot whether its answer is right doesn't replace this step.
These details deserve the closest look, because an error in them changes what you're telling someone:
- Names of people, companies, and products
- Numbers, amounts, and units
- Dates and deadlines
- Quotations, compared word for word with the source
- Links and citations
- Prices, policies, and rules
- Commitments made in your name, such as discounts, deadlines, or next steps
- Negations, since a missing "not" turns a sentence into its opposite
Match the effort to what's at stake. A clumsy phrase in an internal draft is easy to fix later. A wrong price, policy, legal citation, medical detail, or customer promise affects someone else. If you can't judge an important claim yourself, ask a qualified colleague or check an authoritative source before you send it.
Example: an invented discount in a client email
The following example is made up. It shows an ordinary office task and doesn't come from any particular tool.
Your notes after a call:
Call with Norberg Studio. They want a revised quote for two extra workshops. Send it next week.
One sentence from the AI draft:
As agreed on our call, we'll send the revised quote by Friday, October 2, including the 10 percent discount for the two additional workshops.
Now compare the sentence with the notes, one detail at a time.
- "As agreed" claims an agreement. The notes only record a request.
- "Friday, October 2" is a specific date. The notes say "next week".
- "The 10 percent discount" appears nowhere in the notes.
All three are additions that depart from your input, which is the kind of output the NIST definition includes. The sentence sounds polite and professional, and that makes it easy to wave through. Once it's sent, the client has a discount and a date in writing that nobody agreed to.
The corrected sentence keeps to what the notes support:
We'll send you the revised quote for the two additional workshops next week.
Drafting with Yappee so the facts come from you
Checking is easier when you know where each fact in a draft came from. With Yappee, you speak your thoughts on your iPhone or Mac and choose a format such as Email, Meeting Recap, Chat Message, or Status Update. Those are only some of the formats. You can also have the text written in another of the more than 50 supported languages.
Yappee is built to work only with what you say. If you didn't say something, it isn't in the finished text, and Yappee doesn't fill gaps with invented details. Yappee handles the format and the wording, and the facts come from you.
In practice, say every fact the text needs while you record. If the email should contain a price, a date, or a promise, put it in the recording. If you leave it out, it's missing from the email too, and you can add it or record again. Your check then becomes a comparison between the draft and what you said.
This is a design rule, and it doesn't promise that every output is free of errors. Yappee also can't turn a wrong fact you said into a correct one. So before you send, check names, numbers, dates, and commitments, in translated text as well.
Let AI help with the wording and check every fact against its source. When someone questions a detail, go back to the source and fix it if it's wrong. Your name is on what you send, whichever tool wrote the first draft.