Who Checks the Citations?

LePhantomCite is a benchmark for detecting hallucinated legal citations: 1,300 real brief excerpts with injected fabrications across five error types drawn from actual court filings.

Patty Liu, Dominik Stammbach, Peter Henderson  ·  Princeton University

Leaderboard

390-example test set · segment-level detection · values in %

Rank by
# Model Harness F1 Precision Recall Avg tool calls

Agentic runs use our BOED agent (max 30 steps) with CourtListener search, local opinion retrieval, and open web search; temperature 0.8, max tokens 8,192 (10,000 for GPT-5). † Claude Code runs in its own native harness with Opus 4.8 (default effort, max 30 turns) and WebSearch, WebFetch, and CourtListener tools via MCP; it has no non-agentic result. ±95% CIs in the paper.

Overview

Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictions that newer models would hallucinate less or that court sanctions would deter negligent filers, we found over 1,000 filings containing fabricated citations — with this number growing year-over-year. This study evaluates whether AI-based systems can mitigate these errors by automatically detecting hallucinations. We propose a taxonomy of legal citation hallucinations grounded in actual court filings and introduce a dataset of 1,300 brief excerpts containing injected errors. Benchmarking five models in agentic and non-agentic settings reveals that while the latest iterations perform better — GPT-5 achieves 84.4% recall and a 55.0% F1 score in an agentic framework — all models struggle with subtle error categories. Agentic verification remains resource-intensive, with GPT-5 averaging 15.3 steps per excerpt. Furthermore, restricted information access limits the efficacy of even the best agents. This gap creates policy concerns, as it disadvantages both AI systems and litigants who lack subscriptions to commercial legal databases.

The Problem

Legal citations play a prominent role in U.S. legal practice. Attorneys must point to past judicial decisions and laws to make their case. Fabricating citations, or misrepresenting the content of those citations, is the same as pointing to made-up law to win the case — one judge called it:

"an abuse of the adversary system" that puts the integrity of the judicial process at risk. — Noland v. Land of the Free, 114 Cal.App.5th 426, 445 (2025)

The rapid adoption of large language models (LLMs) in the legal system has turned rare individual instances of fabrication into a systemic problem. Pro se litigants, trained attorneys, and even judges are using LLMs to generate briefs, motions, and other court filings. A prevailing argument has been that this problem is temporary — that hallucination rates would diminish as models improve and that high-profile sanctions would induce greater caution. Our findings challenge these assumptions.

Real-world impact
1,000+ Hallucinated Filings
Court filings containing fabricated citations identified across diverse U.S. jurisdictions, growing steadily year-over-year.
→
Trend reversal
Rates Not Decreasing
GPT-5.1 produces hallucinated citations at 6.57%, significantly higher than the best mid-2024 GPT-4o release at 1.23% (p = 0.001).
→
Growing burden
The Jevons Paradox
Newer models cite more cases per document from broader, less canonical sets — harder to verify even if individual rates improve.
Hallucination rate over time across GPT models and cumulative hallucinated court filings

Legal hallucination rates are not consistently decreasing across GPT model generations, while hallucinated citations in court filings grow steadily. Hallucination rate is the percentage of fabricated citations in all model-generated citations.

Judges repeatedly describe the resulting burden as an "enormous waste of judicial resources," and increasingly impose sanctions because "lesser sanctions have been insufficient to deter the conduct." Mid Cent. Operating Eng’rs Health & Welfare Fund v. HoosierVac LLC, 2025 WL 574234, at *3 (S.D.Ind. Feb. 21, 2025); Powhatan County Sch. Bd. v. Skinger, 2025 WL 1559593, at *10 (E.D.Va. June 2, 2025). Courts impose these sanctions while pro se litigants — who stand to benefit most from AI's ability to improve access to justice — are least equipped to detect hallucinated citations and lack access to commercial legal databases.

Through a controlled experiment querying eight generations of ChatGPT models on 92 legal drafting prompts, we find that hallucination rates are no longer consistently decreasing across model generations. Early GPT-4o models released in mid-2024 exhibit the lowest rates at 1.23%, substantially improving over GPT-3.5's ~25%. GPT-5.1 reverses this trend at 6.57%. Newer models also generate more citations per document, drawn from a broader and less canonical set of cases that are individually harder to verify.

The LePhantomCite Dataset

To assess the promise of AI in automatically checking legal filings, we introduce LePhantomCite (Legal Phantom Citation), a benchmarking dataset of legal brief excerpts augmented with injected hallucinations. It is accompanied by a taxonomy of citation hallucination types derived from failure modes observed in real court filings.

Sources
245 Appellate Briefs
From 13 U.S. Courts of Appeals (2012–2021). Pre-2022 filings minimize risk of AI-generated content. Converted via olmOCR and segmented into 5,648 coherent passages.
→
Dataset
1,300 Excerpts
1,000 from appellate briefs + 300 from Dahl et al. (2024). 4,499 total citations; 1,107 contain injected hallucinations across five types. Balanced 50/50 clean/hallucinated split.
→
Evaluation
390 Test Examples
70-30 train/test split. Segment-level evaluation with relaxed matching (match if one segment is a substring of another).

Hallucination Taxonomy

We propose five categories of legal citation hallucinations grounded in failure modes documented in actual court filings. Categories are not mutually exclusive, but only one hallucination type is introduced per citation in the dataset.

Benchmark Results

We evaluate five models in agentic and non-agentic settings using a custom harness with access to CourtListener search, local opinion retrieval, and open web search. We adapt BOED agent from Zheng et al., 2026 as our agentic framework. BOED agent maintains an explicit, language-based belief state that is updated after each action. We adopt this framework because case citation verification is inherently sequential and information-dependent: the agent must extract citations from the brief excerpts, decide on how to gather information and when it has gathered sufficient evidence to make a hallucination determination. We set a maximum of 30 steps per episode. To compare against a production agent harness, we additionally evaluate Claude Code with Opus 4.8, using the same prompt and a matching 30-turn budget.

Full precision / recall / F1 and per-type recall are in the leaderboard above.

See agent behavior analysis

GPT-5 devotes 37.3% of its actions to local opinion search (highest of any model) and 8.3% to open web search as a fallback for citations absent from CourtListener. Among BOED agents it has the lowest false positive rate on citations not found in CourtListener; Claude Code is lowest overall, attempting to look up the cited opinion’s content directly. Weaker models often treat absence from CourtListener as evidence of hallucination.

Model Avg Tool Calls False Positive Rate
(citations not on CourtListener)
GPT-OSS 120B16.840.9%
Claude Code (Opus 4.8)15.410.9%
GPT-515.325.0%
Gemini 2.5 Flash9.765.9%
Qwen3.6-27B9.362.1%
Qwen3-8B7.546.2%

A notable side effect: the agent surfaced over 10 citation typos in pre-LLM briefs that human authors missed, demonstrating practical value beyond the benchmark.

Citation

@misc{liu2026checks,
  title         = {Who Checks the Citations? Benchmarking Legal Hallucination Detection},
  author        = {Liu, Patty and Stammbach, Dominik and Henderson, Peter},
  year          = {2026},
  eprint        = {2606.21155},
  archivePrefix = {arXiv},
  url           = {https://arxiv.org/abs/2606.21155}
}