What Is an AI Hallucination, and Why Is It Dangerous in Legal Work?
A documented legal example of fabricated authority, why fluent AI output can be false, and what lawyers should verify before relying on AI-assisted research.
A fabricated citation rarely looks absurd. It usually looks exactly like a legal citation should.
In Mata v. Avianca, ChatGPT produced a fake Eleventh Circuit opinion, Varghese v. China Southern Airlines Co., Ltd., 925 F.3d 1339 (11th Cir. 2019). Counsel cited it and, when the court ordered the decisions produced, filed a sworn affidavit annexing an excerpt of it. Inside that excerpt was this citation:
Zaunbrecher v. Transocean Offshore Deepwater Drilling, Inc., 772 F.3d 1278, 1283 (11th Cir. 2014).
Case name, reporter, volume, page, pincite, court, year. Every element occupied the right slot.
The sanctions order found that no such decision existed. The reporter citation fell within a real 2014 Eleventh Circuit decision, Witt v. Metropolitan Life Insurance Co., which begins at 772 F.3d 1269.
Judge P. Kevin Castel said that the court need not describe every deficiency. The examples it did identify included:
- Decisions that did not exist at all.
- Invented case names attached to citations belonging to real, unrelated decisions.
- Citations that did not resolve as given, although a real decision with the same name appeared elsewhere.
- Quotations that did not appear in the cited decision.
- Real authorities that did not support the proposition offered.
- A Third Circuit opinion incorrectly identified as a Second Circuit opinion.
Why can an AI system produce something this convincing?
The National Institute of Standards and Technology calls this confabulation and describes it as a natural result of how generative models are designed. These systems produce output that approximates the statistical distribution of their training data. Predicting the next token is one example. That process can produce factually accurate text, but it can also produce text that is factually inaccurate or internally inconsistent.
NIST defines confabulation as confidently presented erroneous or false content. It also notes that an output may contain confabulated logic or citations purporting to justify the answer, which can further mislead people into trusting it.
That general mechanism does not tell us exactly why ChatGPT produced the Zaunbrecher citation. It does explain why fluent legal language should never be mistaken for verified legal authority.
What the court actually sanctioned
The sanctions were not based on the use of ChatGPT alone. Judge Castel wrote that there is nothing inherently improper about using a reliable AI tool for assistance.
The initial submission of nonexistent authorities was nevertheless part of the conduct before the court. The order also emphasized what followed: failure to review the cited authorities, continued reliance after opposing counsel and the court questioned the cases, a sworn affidavit annexing excerpts of fake opinions, and false or misleading statements to the court. One lawyer knowingly claimed that he was away on vacation to obtain an extension when another lawyer was the person who was away.
The distinction matters. The risk was not simply that a machine produced false text. The lawyers presented it as legal authority without performing the verification their roles required.
Verification is a professional responsibility
ABA Formal Opinion 512, applying the ABA Model Rules of Professional Conduct, states that duties to the tribunal require lawyers to review AI-generated analysis and citations before submitting them to a court and to correct errors. Lawyers should also consult the binding rules and guidance in their own jurisdictions.
For legal work, verification should include:
- Does the authority exist?
- Does the cited text appear in the authority?
- Does the authority support the proposition?
- Are the court and jurisdiction correctly identified?
- Is the authority current and still good law?
Legal databases and citators can help confirm a case’s existence, court, citation, and subsequent treatment. They do not decide for the lawyer whether a quotation is accurate in context or whether the holding supports the proposition for which it is offered. Those questions require reading the authority.
Retrieval reduces risk, but does not eliminate it
The number worth remembering is not limited to the sanctions in Mata.
LexisNexis had advertised “100% hallucination-free linked legal citations.” Thomson Reuters said it avoided hallucinations by grounding answers in trusted legal content. A peer-reviewed study by Stanford and Yale researchers reported that, in tests conducted from March through May 2024, the evaluated legal AI products produced hallucinated responses between 17% and 33% of the time.
The researchers used a broader legal definition of hallucination than a fabricated citation. They classified an answer as hallucinated if it contained incorrect information or falsely asserted that a source supported a proposition. The legal research products performed better than the general-purpose GPT-4 system tested in the study, suggesting that retrieval helped, but it did not eliminate error.
The vendors challenged aspects of the study’s definitions and methodology, and the products have continued to change. The figures should therefore be read as results for the systems and test conditions examined in 2024, not as a current product ranking. The durable lesson is the difference between retrieving a real source and verifying that the source supports a correct legal conclusion.
A citation that looks professional is not necessarily a citation that exists. A citation that exists is not necessarily one that supports the answer.
What verification does your firm require before AI-assisted research enters a filing?
References
- Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023), Opinion and Order on Sanctions
- NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, section 2.2
- ABA Standing Committee on Ethics and Professional Responsibility, Formal Opinion 512
- Varun Magesh et al., Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, 22 Journal of Empirical Legal Studies 216 (2025)
- LexisNexis response to the Stanford study
- Thomson Reuters response on error definitions and evaluation
This article is general information for legal professionals, not legal advice or an ethics opinion. Rules of professional conduct vary by jurisdiction—consult yours.