Our results reveal that some failures are structural rather than purely technical. The most consequential barrier is access to legal data.
Recommendations
Improve public access to legal data. Although judicial opinions are public domain, they are not freely accessible through a centralized service. PACER charges per-page fees; commercial providers bill up to $100 per query; public repositories like CourtListener have incomplete coverage and contain no pincite information. Contrast with Canada: CanLII provides free, near-complete access to all court judgments — one Canadian court noted that verifying AI-cited cases requires only "a simple search on CanLII," an assumption US courts cannot reasonably make. Improving automated verification will require either broader public access to legal databases or verification systems built atop commercial platforms.
Support pro se litigants with targeted AI literacy guidance. The share of AI hallucinations in courts attributable to pro se filers has risen every year. Courts already provide self-help resources for unrepresented litigants — these should include guidance on AI citation risks and concrete verification best practices. Raising sanctions without addressing structural limitations makes AI a false promise for the litigants who stand to benefit most.
Develop and deploy automated citation verification tools. Currently, verifying citations is largely a manual process performed by attorneys, law clerks, or judges — especially costly when a cited case does not exist. Our benchmark, dataset, and agent harness provide a foundation for building and auditing such tools. With sufficient data access, automated tools could ease court verification burdens and assist non-lawyers directly in verifying AI-generated citations.