Three years ago, a federal judge in Manhattan sanctioned two attorneys for filing a brief built on cases that did not exist. Their research tool was ChatGPT, and it had invented the opinions: plausible captions, plausible citations, plausible internal quotes. When opposing counsel could not locate the cases, one of the lawyers went back to the same tool and asked whether they were real. It assured him they were.
Mata v. Avianca made headlines because it was first. It is no longer close to alone. A researcher who tracks these decisions now catalogs more than 1,800 of them worldwide, and the pattern has climbed the professional ladder: in early 2025 a federal court in Minnesota threw out an expert declaration on AI-generated misinformation, from a Stanford professor who studies exactly that subject, because the declaration itself contained AI-fabricated citations. An expert on machine deception, undone by unverified machine output, in a case about machine deception.
The reflex is to laugh at the lawyers. That is the wrong lesson. Nobody in these cases failed at law. They failed at verification, which is our job, and they failed at it in a way that is coming for investigators faster than it came for attorneys. Because the same tools that will hallucinate a court case will just as fluently hallucinate a person, a business address, a criminal history, or the "confirmed" location of a photograph. If your finding ends up in a report, a declaration, or testimony, the question waiting for you is the one this newsletter exists to answer: how do you know?
So this issue is not an argument for or against AI in investigations. I use it, and the productivity gain is real. This is about the discipline that lets you take the gain without inheriting the failure.
The machine fails in a specific, predictable way
A language model is not a database, and querying it is not retrieval. It generates the most plausible continuation of your prompt, one token at a time, based on patterns in what it was trained on. Most of the time the most plausible answer is also true. When it is not, nothing about the output changes. Same tone, same formatting, same confidence. The machine was built to be fluent, and it is accurate only as a byproduct of fluency usually landing on the truth.
I spent several years as a detective before I moved into forensics, and every human source I handled in those years had some tell when they were inventing something. Models have none. Worse, they do not know when they do not know, so the tell you are unconsciously waiting for never comes.
The scale is measurable. A Stanford study of legal queries found that general-purpose models of the 2023 generation invented or misstated the law in more than half of test questions, with rates varying by model and task. The purpose-built, retrieval-grounded legal research tools that followed are better, and still produced false or misgrounded answers in roughly one in six to one in three responses in a follow-up benchmark. Grounding a model in real documents reduces hallucination. It does not eliminate it, because the model can still misread, overextend, or misattribute the very source it is quoting.
That last failure is the dangerous one for us. A fabricated case dies on the first docket search, but a real source cited for something it does not quite say survives the first check, and sometimes every check until cross-examination. The most damaging AI errors are distortions with a genuine citation attached, not outright inventions.
What the machine is actually good for
Used correctly, AI sits downstream of collection and upstream of judgment. Inside that band, it is the best force multiplier to hit this field in a decade:
- Working your own corpus. Summarizing, translating, and extracting entities from material you collected and preserved yourself. The facts are yours and verifiable; the model is doing labor, not sourcing.
- Hypothesis generation. Asking what else could explain a pattern, what a boring explanation would look like, what you might be missing. An earlier issue of this newsletter argued that if you have only one hypothesis you are not analyzing, you are confirming. A model is a cheap, tireless generator of rival hypotheses.
- Query and pivot drafting. Search strings, dork variations, alternate spellings and transliterations, places to look next. Every lead it suggests gets checked by you, in the sources, on the record.
- Drudgery. Cleaning transcripts, structuring timelines from your own logs, first-pass triage of bulk material you already hold.
The band has a hard edge, and it is this: the model is never the source of a fact. Not a name, an address, a date, a quote, or a "this account belongs to this person." The moment an unverified model statement enters your case file as a fact, you have broken the chain between your findings and reality, and you will not necessarily find the break before someone else does.
Leads, not evidence
Issue 1 of this newsletter, Search Is Cheap. Judgment Is Not., walked through the Admiralty system: grade the source A through F, grade the information 1 through 6. Run a language model through it honestly and you get F6. Reliability: cannot be judged, because the "source" is a statistical average of the internet with no stake in being right. Credibility: cannot be judged, until you have done the work. F6 sounds insulting until you remember it is also the correct starting grade for any tipster you know nothing about, and a tipster with perfect grammar is what a model is.
You already know how to handle a tipster. The same rules transfer whole:
- Every model-supplied fact is a lead until independently sourced. Primary source or it does not enter the file. If you cannot find independent support, the fact does not exist, no matter how specific the model was.
- Demand the sourcing, then check it yourself. Asking a model for its sources is a useful pressure test, but the citations it produces are subject to the same failure as everything else it produces. A citation is a lead to a document, not proof the document says what the model claims. Open it. Read it.
- Never show it the conclusion you want. A model is agreeable by construction. Frame a prompt as "confirm that X" and it will oblige, which makes it a confirmation-bias amplifier strapped to a jet engine. Ask it to attack your hypothesis instead. It is genuinely good at that, and that version of the tool makes your analysis stronger rather than softer.
- Do not let it write your conclusions. Findings, confidence levels, and opinions are the parts of the work that are yours to defend under oath. Drafting help on structure is one thing. Reasoning outsourced is testimony you did not author and cannot fully explain.
Protect the case from your own tools
Two more exposures, less discussed and more dangerous than hallucination.
Confidentiality. Pasting case material into a consumer chatbot is disclosure to a third party. Depending on what the material is, that can mean waived privilege, a data-protection violation, a breached protective order, or a violation of your agency's policy, and the platform's terms may allow the content to be retained or used for training. The rule is simple to state and has to be decided before the case, not during it: case data goes only into deployments approved to hold it. If no approved deployment exists, the model works from your generic, sanitized questions only, and the case material never leaves your environment.
The record. If AI touched the workflow, your contemporaneous log should say so: what tool, what date, what was asked, what came back, and what you did to verify what survived. Federal judges began issuing standing orders on generative AI within weeks of Avianca, and disclosure requirements keep spreading. But the deeper reason is older than any standing order. "Did you use artificial intelligence in preparing this report?" is already being asked in depositions, and there are only two honest answers. "Yes, and here is the verification trail for every fact it touched" is a strong answer. "Yes" followed by silence is the whole case.
Treat the model's role exactly like any other tool in your workflow: named, versioned, logged, and subordinate to a documented verification step. Nobody gets hurt in that deposition for having used AI. The damage lands on the examiner who cannot show what happened after it answered.
The protocol, in one box
Before you prompt:
- Decide what the model is allowed to see. Case data only in approved deployments.
- Frame the ask neutrally. No preferred conclusions in the prompt.
While you work:
- Treat every factual statement as an F6 lead, not a finding.
- Ask for sources, then open and read each one yourself.
- Ask the model to argue against your working hypothesis.
- Keep the prompts and outputs. They are part of your record.
Before anything enters the file:
- Independently source every fact that survived. Primary source or it does not exist.
- Check every quote against the original document, word for word.
- Log the tool, date, use, and verification in your case notes.
- Be ready to describe all of it, out loud, under oath, without notes.
Bottom line
AI belongs in your workflow for the same reason every prior force multiplier did: it buys back hours. It stays safe in your workflow only under the discipline this trade already invented, because a language model is a source like any other, except faster, more confident, and completely indifferent to being wrong. Grade it F6. Work its output like tips from a stranger. Source everything independently, keep the case data out of unapproved tools, log the machine's role like you would log any instrument, and never let it author a conclusion you have to defend. The investigators who thrive with these tools will not be the best prompters. They will be the best verifiers, which is what they were before the tools arrived.
Take it further
This issue is one module of a larger method. In the Certified OSINT Investigator, Court-Ready Practitioner (COI-CRP) course, this exact discipline is Module 9, AI-Assisted OSINT and Hallucination Control, and it sits inside the full arc: mission framing, operational security, legal authority across borders, search and identity resolution, SOCMINT, image and geolocation analysis, dark web and cryptocurrency tracing, and evidence handling through reports, testimony, and a mock cross-examination. Self-paced and online: 10 modules, 80 lessons, 8 downloadable templates, module quizzes, and a 100-question final exam.
Tuition is $2,497, with government purchase orders, net-30 invoicing, and agency licensing available.
Syllabus and enrollment: thewaldrepcompany.com/courses/coi-crp/
If this issue was useful, subscribe to Digital Forensics Today below and pass it to an investigator, analyst, or attorney who is being told to "use AI" without being told how. Every issue is about doing OSINT and digital forensics that stands up when someone asks how you know.
Sources and further reading
- Mata v. Avianca, Inc., No. 22-cv-1461 (PKC), Opinion and Order on Sanctions (S.D.N.Y. June 22, 2023). courtlistener.com
- Damien Charlotin, AI Hallucination Cases database (tracking court decisions involving hallucinated AI content). damiencharlotin.com/hallucinations
- Kohls v. Ellison, No. 24-cv-3754 (D. Minn. Jan. 10, 2025) (order excluding expert declaration containing AI-fabricated citations).
- Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E. Ho, "Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models," Journal of Legal Analysis 16, no. 1 (2024). doi.org/10.1093/jla/laae003
- Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning, and Daniel E. Ho, "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools," Journal of Empirical Legal Studies (2025). doi.org/10.1111/jels.12413
- Judge Brantley Starr, Mandatory Certification Regarding Generative Artificial Intelligence, N.D. Tex. standing order (2023). txnd.uscourts.gov
- Ziwei Ji et al., "Survey of Hallucination in Natural Language Generation," ACM Computing Surveys 55, no. 12 (2023). doi.org/10.1145/3571730
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), 2023. nist.gov
- Office of the Director of National Intelligence, Intelligence Community Directive (ICD) 203: Analytic Standards. dni.gov (PDF)
- Headquarters, Department of the Army, FM 2-22.3, Human Intelligence Collector Operations, Appendix B (source reliability and information credibility). irp.fas.org (PDF)