Hallucinated Justice
Large Language Models, Fabricated Legal Citations, and the Future of Deterministic AI in Courtroom Advocacy
Abstract
The rapid integration of Large Language Models (LLMs) into legal practice has introduced significant opportunities for efficiency in legal research, drafting, and litigation support. However, these systems also present severe risks to judicial integrity because of their tendency to hallucinate—confidently generating fabricated legal citations, nonexistent cases, inaccurate quotations, and false legal reasoning. Recent court sanctions against attorneys submitting AI-generated fictitious authorities demonstrate that hallucinations are no longer theoretical concerns but operational threats to the administration of justice. This paper examines the structural reasons why modern transformer-based LLMs hallucinate in legal contexts and analyzes how such hallucinations can undermine the logical architecture of legal briefs. Drawing on documented judicial incidents, legal ethics principles, e-discovery scholarship, and AI governance literature, the paper argues that unconstrained general-purpose LLMs are structurally incompatible with authoritative legal citation generation. The paper further proposes that the future of reliable legal AI lies in specialized Small Language Models (SLMs) integrated with deterministic legal retrieval and citation-verification engines capable of producing auditable, explainable, and authenticated legal outputs. Systems such as Lexi by SigmaCore.AI represent an emerging model for AI architectures aligned with the epistemic and procedural requirements of modern justice systems.
1. Introduction
The legal profession has historically evolved through cautious integration of technologies that enhance research, advocacy, and judicial administration. From printed reporters to digital legal databases such as Westlaw and LexisNexis, legal innovation has traditionally focused on improving access to authoritative sources while preserving the integrity of legal reasoning. The emergence of Large Language Models (LLMs), however, represents a fundamentally different technological transition. Unlike traditional legal research systems, LLMs are generative probabilistic systems capable of producing original text, arguments, summaries, and legal analyses through statistical prediction rather than deterministic legal retrieval.
The adoption of these systems by attorneys has accelerated rapidly. Lawyers increasingly employ generative AI to draft motions, summarize depositions, prepare appellate briefs, generate discovery responses, analyze contracts, and assist in legal research. The efficiency gains are substantial. Tasks that previously required many hours of manual legal drafting can now be completed within minutes. Yet this productivity comes with significant structural risks.
One of the most serious dangers is hallucination—the generation of fabricated or inaccurate information presented with persuasive confidence. In legal practice, hallucinations may include nonexistent cases, invented quotations, false procedural histories, incorrect legal standards, fabricated statutory references, or mischaracterized precedents. Because legal reasoning depends heavily on precise citation structures and authoritative precedent, hallucinations threaten not merely drafting accuracy but the reliability of judicial process itself.
The most widely publicized example occurred in Mata v. Avianca, Inc. (S.D.N.Y. 2023), where attorneys submitted a federal court brief containing multiple fictitious cases generated by ChatGPT. The cited authorities included fabricated judicial opinions, nonexistent quotations, and invented legal reasoning. Judge P. Kevin Castel imposed sanctions on the attorneys, emphasizing that legal professionals retain an independent duty to verify all cited authorities before filing documents with the court. Since that decision, additional courts across the United States have reported similar incidents involving fabricated citations and AI-generated hallucinations in legal submissions.
These events demonstrate that hallucinations are not isolated technical anomalies but recurring structural characteristics of probabilistic language-generation systems. The legal system is especially vulnerable because judicial decision-making depends upon the authenticity and integrity of cited authority. A legal brief functions as a logically interconnected argumentative structure in which every citation performs a doctrinal role. Once hallucinated authorities enter this structure, the integrity of the entire argument may become compromised.
This paper examines the legal, technical, and jurisprudential implications of AI hallucinations in legal advocacy. It analyzes why LLMs hallucinate, how hallucinations can disrupt the logical structure of legal argumentation, and why deterministic verification architectures are necessary for the future of trustworthy legal AI systems.
2. The Rise of Generative AI in Legal Practice
2.1 Expansion of AI-Assisted Legal Drafting
Generative AI tools are increasingly used throughout the legal profession for drafting motions and pleadings, summarizing judicial opinions, reviewing contracts, preparing discovery requests, generating deposition outlines, conducting preliminary legal research, drafting client communications, and assisting with appellate briefing. The attraction is understandable because LLMs can synthesize enormous volumes of legal text and produce highly polished prose rapidly and inexpensively.
Legal technology vendors have aggressively integrated generative AI into legal workflows, often marketing these tools as transformative productivity platforms. However, many such systems remain fundamentally dependent upon general-purpose transformer architectures that prioritize linguistic fluency rather than validated legal accuracy.
2.2 Judicial Responses to AI Hallucinations
The judiciary has responded with increasing concern regarding hallucinated legal authorities. In Mata v. Avianca, Inc., attorneys submitted a brief citing several nonexistent cases generated by ChatGPT. When opposing counsel and the court failed to locate the cited authorities, the attorneys initially insisted the cases were real because the AI system had represented them as authentic. The court ultimately sanctioned counsel, emphasizing that attorneys remain professionally responsible for all filings submitted to the court. This case became a landmark warning regarding the dangers of unverified AI-generated legal research.
Subsequent reports documented additional incidents involving fabricated AI-generated citations. Legal publications, including LawFuel, reported cases in which courts criticized or sanctioned attorneys for submitting hallucinated authorities generated through AI-assisted drafting systems. These examples collectively demonstrate that AI hallucinations have become operational legal risks rather than hypothetical technological concerns.
3. Why Large Language Models Hallucinate
3.1 Probabilistic Token Prediction
Modern LLMs are primarily based on transformer architectures introduced by Vaswani et al. (2017). These systems generate language by predicting statistically likely token sequences based on patterns learned from massive text corpora. Importantly, LLMs do not “understand” law in a doctrinal or jurisprudential sense. They do not retrieve authorities through deterministic legal validation. Instead, they generate plausible language patterns probabilistically. Consequently, when an LLM lacks certainty regarding a citation, quotation, or precedent, it may generate a syntactically plausible but entirely fictitious legal authority.
3.2 Hallucination as a Structural Phenomenon
Hallucinations arise from several structural characteristics of transformer-based systems. First, LLMs are optimized for linguistic continuation rather than factual verification. Their objective function rewards coherent sequence generation rather than truth validation. Second, legal citations involve highly precise references involving jurisdiction, reporter volume numbers, procedural posture, and doctrinal context. Even small deviations may produce entirely nonexistent authorities. Third, when prompts are vague or incomplete, models may interpolate missing information by generating statistically probable but inaccurate legal references. Finally, LLMs frequently produce hallucinations with high rhetorical confidence. This characteristic is especially dangerous in legal practice because persuasive presentation may conceal factual unreliability.
3.3 The Difference Between Search and Generation
Traditional legal research platforms operate through deterministic retrieval systems in which cases exist independently within validated legal databases. Search systems retrieve authentic authorities that already exist in the database. By contrast, unconstrained LLMs generate text dynamically through statistical prediction. The distinction between retrieval and generation is foundational. Retrieval systems locate existing authority, whereas generative systems create probabilistic textual outputs. This architectural distinction explains why hallucinations occur.
4. The Logical Structure of Legal Briefs
4.1 Legal Briefs as Structured Logical Systems
Legal briefs are far more than persuasive narratives. They are carefully constructed logical frameworks built upon interconnected authorities, procedural standards, factual records, evidentiary references, and doctrinal hierarchies. Every section of a brief performs a specific analytical function within the broader legal argument. A properly developed submission typically includes a statement of jurisdiction, procedural history, statement of facts, issues presented, applicable standards of review, governing legal authorities, application of law to fact patterns, and a request for relief. These elements are not isolated components but interdependent layers of legal reasoning that collectively support the attorney’s theory of the case. The integrity of the entire structure therefore depends upon the authenticity, precision, and reliability of the authorities cited throughout the document.
4.2 Citation as Epistemic Infrastructure
Within legal reasoning, citations function as a form of epistemic infrastructure. They do far more than merely reference supporting materials. Citations authenticate legal propositions, establish precedential authority, demonstrate doctrinal continuity, support procedural legitimacy, and provide the foundation upon which judicial reasoning may proceed. Courts rely upon citations as signals that an attorney’s assertions are grounded in verified legal authority rather than unsupported advocacy. In this sense, citations operate as the connective tissue between legal argument and institutional legitimacy. Consequently, when an LLM hallucinates a citation, invents a quotation, or misstates a precedent, the damage extends well beyond factual inaccuracy. Such hallucinations compromise the epistemic reliability of the legal submission itself and potentially undermine the court’s confidence in the integrity of the broader argument.
4.3 How Hallucinations Corrupt Legal Logic
Hallucinated authorities can corrupt legal reasoning at multiple levels simultaneously. A fabricated case may falsely imply that binding or persuasive precedent exists for a proposition unsupported by actual law. Incorrect or hallucinated standards of review may materially distort appellate analysis and alter how a court evaluates legal questions. In other instances, LLMs may improperly merge doctrines from separate jurisdictions, confuse federal and state precedents, or inaccurately represent procedural histories, thereby distorting the applicability of cited authorities. Particularly dangerous are situations in which authentic cases are cited but their holdings are inaccurately characterized, since such errors can be more difficult to detect during ordinary review. Because legal briefs operate as interconnected logical systems, even a single hallucinated authority may contaminate multiple portions of the argument, producing cascading analytical failures throughout the submission. What initially appears to be an isolated citation error may therefore compromise the doctrinal coherence and credibility of the entire brief.
5. Illustrative Examples of Hallucination Damage
5.1 Mata v. Avianca
The fabricated authorities submitted in Mata v. Avianca provide one of the clearest illustrations of how AI hallucinations can destabilize legal reasoning. The nonexistent cases generated by ChatGPT were not peripheral references but foundational authorities supporting the plaintiffs’ argument. Because the cited precedents did not exist, the doctrinal structure supporting the legal analysis effectively collapsed once the hallucinations were discovered. The court emphasized that attorneys cannot delegate legal verification responsibilities to artificial intelligence systems and reaffirmed that counsel remain professionally accountable for all authorities submitted to the court.
5.2 Fabricated Quotations
Courts and legal commentators have also documented incidents in which AI systems generated quotations inaccurately attributed to authentic judicial opinions. These hallucinations are particularly dangerous because they often appear superficially credible during cursory review. A fabricated quotation inserted into an otherwise authentic case citation may materially distort the meaning of a judicial holding while remaining difficult to detect without direct source verification. Such distortions threaten the accuracy of judicial interpretation and may improperly influence legal reasoning if left undiscovered.
5.3 Mischaracterized Holdings
Large Language Models may also cite authentic cases while incorrectly describing their holdings, procedural posture, or doctrinal significance. This form of hallucination can be especially problematic because the underlying authority exists, creating the appearance of legitimacy. However, the legal characterization itself may be materially inaccurate. Misstated holdings can alter the perceived relevance of precedent, distort doctrinal boundaries, and improperly support legal conclusions unsupported by the actual case law.
5.4 Erosion of Judicial Trust
Repeated submission of hallucinated authorities risks producing broader institutional consequences beyond individual cases. As courts encounter increasing numbers of fabricated citations and inaccurate AI-generated submissions, judicial confidence in AI-assisted legal practice may erode generally. Courts may respond with heightened scrutiny, stricter certification requirements, expanded verification obligations, or broader skepticism toward AI-assisted submissions altogether. Such reactions could impose additional procedural burdens even upon attorneys who employ AI responsibly and cautiously.
6. Legal Ethics and Professional Responsibility
6.1 Attorney Duties of Competence and Candor
Rules of professional responsibility impose affirmative duties of competence, diligence, and candor toward tribunals. Under ABA Model Rule 3.3, attorneys may not knowingly make false statements of law or fact to a court. Even negligent submission of hallucinated authorities may create ethical exposure because attorneys are expected to independently verify the accuracy of all cited legal authority before filing. Courts have increasingly emphasized that reliance upon AI-generated research does not diminish these professional obligations.
6.2 Malpractice and Fiduciary Exposure
AI-generated hallucinations may expose attorneys and law firms to significant professional and fiduciary liability. Potential consequences include malpractice claims, sanctions, fee disputes, reputational harm, judicial disciplinary actions, and direct prejudice to client interests. These risks extend beyond procedural embarrassment because inaccurate authorities may materially affect litigation outcomes, appellate review, settlement negotiations, or judicial decision-making.
6.3 Nondelegable Verification Responsibility
Courts consistently maintain that verification responsibilities remain nondelegable regardless of the technological tools employed during drafting. Artificial intelligence systems may assist legal work, but they cannot replace professional judgment, ethical accountability, or attorney oversight. The ultimate responsibility for ensuring the accuracy and authenticity of legal submissions remains with counsel.
7. Deterministic Legal AI: Toward Reliable Judicial Systems
7.1 Limitations of Monolithic LLMs
General-purpose LLMs are structurally unsuitable for authoritative legal citation generation because they lack deterministic verification mechanisms capable of guaranteeing citation authenticity. Their probabilistic architecture means that no amount of prompting alone can fully eliminate hallucination risk. As long as such systems generate legal text through statistical prediction rather than validated retrieval, the possibility of fabricated authority remains inherent within the architecture itself.
7.2 Specialized Small Language Models (SLMs)
A more promising approach involves specialized Small Language Models integrated with deterministic legal retrieval systems. Such systems may incorporate verified legal databases, deterministic citation engines, Shepardizing verification, jurisdictional validation, procedural-rule verification, explainable reasoning chains, and comprehensive audit logs capable of tracing how authorities were retrieved and applied. Unlike unconstrained generative systems, these architectures seek to embed legal validation directly into the AI workflow itself.
7.3 Deterministic Retrieval Architectures
Under deterministic legal architectures, authorities are retrieved exclusively from authenticated databases rather than probabilistically generated. The language model functions only as a summarization layer, drafting assistant, or narrative synthesizer operating on verified underlying data. Critically, the system lacks the ability to invent citations outside validated repositories. This architectural separation between retrieval and generation dramatically reduces hallucination risk while preserving many of the efficiency benefits associated with AI-assisted legal drafting.
7.4 Lexi by SigmaCore.AI as an Emerging Model
Emerging systems such as Lexi by SigmaCore.AI illustrate how legal AI may evolve toward deterministic governance architectures aligned with the reliability requirements of judicial systems. By integrating constrained language generation with authenticated legal retrieval mechanisms, such platforms seek to preserve citation integrity while still delivering AI-assisted drafting efficiency. The future of trustworthy legal AI will likely depend upon such hybrid governance models rather than unconstrained generative architectures that prioritize linguistic fluency over verified legal accuracy.
8. Broader Implications for the Justice System
The justice system depends fundamentally upon informational reliability. Courts assume that attorneys submit verified authorities, accurate quotations, and authentic precedents. If hallucinated authorities become widespread, the adversarial system itself incurs systemic epistemic contamination. Judges, opposing counsel, clients, and litigants become exposed to hidden informational risk generated by probabilistic systems masquerading as authoritative legal reasoning engines.
The issue therefore extends beyond technology adoption into the preservation of procedural legitimacy and public trust in judicial institutions. The future integration of AI into legal systems will require deterministic verification, explainable reasoning, human oversight, citation authentication, auditable retrieval systems, and transparent governance architectures. Without such safeguards, generative AI risks undermining the very reliability upon which the rule of law depends.
9. Conclusion
Large Language Models present substantial opportunities for efficiency within legal practice, but they also introduce severe structural risks to judicial integrity because of their propensity to hallucinate legal authorities, quotations, and reasoning.
The problem is not merely occasional factual inaccuracy. Hallucinations strike at the epistemic foundation of legal advocacy because legal briefs function as logically interconnected systems dependent upon authentic precedent and accurate doctrinal support.
Documented sanction cases such as Mata v. Avianca demonstrate that hallucinations are operational realities capable of damaging litigants, misleading courts, and exposing attorneys to professional liability.
The core problem is architectural. General-purpose LLMs are probabilistic text-generation systems rather than deterministic legal verification engines. As a result, unconstrained use of such systems within legal briefing creates unavoidable risks.
The future of trustworthy legal AI therefore lies not in unrestricted generative systems but in specialized deterministic architectures integrating verified legal databases, authenticated retrieval systems, citation validation, explainable reasoning, and constrained language generation.
Systems such as Lexi by SigmaCore.AI point toward an emerging governance model capable of preserving both technological innovation and judicial reliability. Ultimately, the survival of trustworthy AI-assisted advocacy will depend upon whether legal AI systems can align with the evidentiary precision, procedural accountability, and epistemic integrity required by modern justice systems.
References
-
American Bar Association. Model Rules of Professional Conduct.
-
Bommasani, R., Hudson, D. A., Adeli, E., et al. (2021). On the opportunities and risks of foundation models. Stanford Center for Research on Foundation Models.
-
Electronic Discovery Reference Model (EDRM). (2026). EDRM Model. https://edrm.net/edrm-model/current/
-
Electronic Discovery Reference Model (EDRM). (2026). Cite-checking to find hallucinated cases deemed insufficient. https://edrm.net/2026/04/cite-checking-to-find-hallucinated-cases-deemed-insufficient/
-
LawFuel. (2024). Lawyers sanctioned again for relying on bogus legal AI citations. https://www.lawfuel.com/lawyers-sanctioned-again-for-relying-on-bogus-legal-ai-citations/
-
Lipton, Z. C. (2018). The mythos of model interpretability. Queue, 16(3), 31–57.
-
Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023).
-
OpenAI. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774.
-
Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
-
Clearbrief. (2026). https://clearbrief.com/