# Understanding AI in a Month 19: Why Fluent Systems Are Unreliable

Understanding AI in a Month — Day 19 · 2026-09-30

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

On the twenty-fifth of September, a federal judge in Massachusetts, Angel Kelley, sanctioned a lawyer over briefs he had filed in an insurance lawsuit. The briefs cited cases for things they did not say, quoted cases for words they do not contain, and cited cases that do not exist. For one of the briefs, the lawyer admitted using AI. According to the order, he told the court he had believed that the enterprise version of the AI software he was using did not hallucinate cases. The judge ordered his firm to pay the other side's costs, up to ten thousand dollars, and revoked his permission to appear in the case.

Three days later, the researcher Damien Charlotin updated his public database of court decisions dealing with AI-invented material. It now lists two thousand and ninety-five of them. Today: why a system can sound entirely competent and still be wrong.

Two readings of stories like this mislead.

The first is that a system good enough to pass the exams can be trusted with the work. In March twenty twenty-three, OpenAI reported that GPT-4 scored around the top ten per cent of test takers on a simulated bar exam. The lawyer's belief was a version of the same thing. In a paper posted in May twenty twenty-four, researchers at Stanford and Yale reported testing legal research tools whose makers had described them as avoiding hallucinations, one as offering hallucination-free citations. On their test questions, each tool hallucinated between seventeen and thirty-three per cent of the time. A score, or a label, earned under one set of conditions does not travel automatically to another.

The second reading is the opposite: that stories like this prove the tools are useless. Charlotin's count is a count of court decisions, not a rate of error. It could rise with more use, more checking or more reporting, and there's no count of filings to divide it by. Careful studies find large gains on some tasks alongside failures on others. Neither blanket trust nor blanket distrust fits the evidence. The question is where. Three ideas help.

The first idea is distribution shift, and a classic cautionary tale is about flu.

In early two thousand and nine, Google researchers reported in the journal Nature that they could estimate flu activity from what people typed into its search engine, a week or two ahead of the government's own reports. In February twenty thirteen, Nature reported that Google Flu Trends was predicting more than double the share of doctor visits for flu-like illness that the Centers for Disease Control were recording. A year later, in the journal Science, David Lazer and three colleagues dissected what went wrong. Part of it was that the original system had latched on to search terms that rise every winter, flu or no flu, and it completely missed the out-of-season pandemic of two thousand and nine.

> In short, the initial version of G F T was part flu detector, part winter detector.

> — *David Lazer and colleagues, 'The Parable of Google Flu: Traps in Big Data Analysis', Science 343, 14 March 2014, page 1203, section 'Big Data Hubris'; author copy at https://gking.harvard.edu/files/gking/files/0314policyforumff.pdf*

> As printed in the source: “In short, the initial version of GFT was part flu detector, part winter detector.”

Another likely culprit, they argued, was that Google's search engine, and the way people used it, kept changing underneath the model. That's distribution shift: a system is built and checked on one slice of the world, then used on another. The model needn't change for its accuracy to change. A bar exam question and a brief in a live lawsuit are different slices of the world.

The second idea is calibration, and it comes from weather forecasting. A forecaster is well calibrated if, on the days they say seventy per cent chance of rain, it rains on about seventy per cent of them. Even then, it stays dry on three of those days in ten. In nineteen fifty, Glenn Brier of the U.S. Weather Bureau worried that the way forecasts were scored could push forecasters to game the score.

> This may lead the forecaster to forecast something other than what he thinks will occur

> — *Glenn W. Brier, U.S. Weather Bureau, 'Verification of Forecasts Expressed in Terms of Probability', Monthly Weather Review 78(1), January 1950 (issued 15 April 1950), page 1, 'Introduction'*

His answer was probability forecasts, scored in a way he argued could not push the forecaster in any undesirable way.

Here's the catch. Underneath, a language model does assign probabilities to the words it might write next. OpenAI's GPT-4 report found that the model, before the extra training that turns it into an assistant, was highly calibrated on part of a multiple-choice test: its confidence generally matched how often it was right. After that training, OpenAI found, calibration was reduced. And in an ordinary answer, none of those probabilities is shown to you. A sentence built around an invented case can read exactly like one built around a real case. Sounding certain is a property of the writing, not a measure of how often answers like this are right.

The third idea is jagged capability. In July twenty twenty-four, the AI researcher Andrej Karpathy gave it a name: jagged intelligence. Some things these systems do extremely well by human standards, others they fail badly, and

> it's not always obvious which is which

> — *Andrej Karpathy, 'Jagged Intelligence', post on X, 25 July 2024, 17:50 UTC, in the paragraph that restates the heading 'Jagged Intelligence.' and continues 'Some things work extremely well'; https://x.com/karpathy/status/1816531576228053133*

In people, as Karpathy noted, abilities tend to move together. Someone who can draft a strong legal argument can usually check that a case exists. In these systems, doing the hard thing well tells you less than you'd expect about the easy thing next to it.

Jaggedness first. In a study published in September twenty twenty-three, business-school researchers gave seven hundred and fifty-eight consultants at Boston Consulting Group realistic tasks, some with GPT-4 and some without. On tasks the model could handle, those using it completed about twelve per cent more tasks, about a quarter faster, and at markedly higher quality. On one task chosen to fall outside what it could do, those using AI were nineteen percentage points less likely to reach the correct answer. The authors called the tasks seemingly similar in difficulty. The results weren't.

The legal tools fit the same pattern. Their makers pointed to a design that looks up real case law before answering. The Stanford and Yale test found this reduced hallucinations compared with GPT-4, but didn't remove them.

And the courts have been keeping score. Charlotin's published file has fewer than eighty decisions dated through twenty twenty-four, about nine hundred and thirty through last year, and over two thousand now. By my count of his published data, the pace has held at roughly a hundred to a hundred and eighty decisions a month over the twelve complete months to August. More than half involve people representing themselves; over eight hundred involve lawyers.

There are two postures towards all this, and they put the fix in different places.

The first says: change what the machine is rewarded for. In September twenty twenty-five, OpenAI published a paper arguing that hallucinations persist partly because of how models are graded, like students on a multiple-choice exam, where a blank scores zero and a guess sometimes scores a point. Its blog was careful to say that evaluations don't directly cause hallucinations, but that

> most evaluations measure model performance in a way that encourages guessing rather than honesty about uncertainty.

> — *OpenAI, 'Why language models hallucinate', blog post, 5 September 2025, section 'Teaching to the test', first paragraph; https://openai.com/index/why-language-models-hallucinate/*

That's Brier's worry, seventy-five years on. In OpenAI's own example, a newer model that declined about half of a quiz's questions got a quarter of them wrong. An older one that almost never declined got slightly more right, and three quarters wrong. Those are OpenAI's numbers about its own models.

The second says: test the machine and the person together, because that's what gets used. In February this year, researchers at Oxford published a trial in Nature Medicine with almost thirteen hundred members of the British public, each given a medical scenario. It ran in late twenty twenty-four, with GPT-4o and two other models. Tested alone, the chatbots named a relevant condition in about ninety-five per cent of cases. People using those same chatbots did so in fewer than thirty-five per cent, which was worse than people left to use whatever they'd normally use.

> Standard benchmarks for medical knowledge and simulated patient interactions do not predict the failures we find with human participants.

> — *Andrew M. Bean and colleagues, 'Reliability of LLMs as medical assistants for the general public: a randomized preregistered study', Nature Medicine, published 9 February 2026, abstract; https://www.nature.com/articles/s41591-025-04074-y*

The court's order adds a complementary duty: whatever the tool, the person using it stays responsible.

> There is no rule against the use of AI in researching and drafting legal papers, but it must be utilized responsibly.

> — *United States District Court for the District of Massachusetts, Memorandum and Order on Motion for Sanctions, Civil Action No. 25-CV-12395-AK, Document 155, Judge Angel Kelley, 25 September 2026, page 12, section II 'Discussion'*

The difficulty is old. In nineteen eighty-three the psychologist Lisanne Bainbridge wrote about the ironies of automation: the more advanced a system, the more crucial the person watching it may become, and the harder that person's job gets. I think the two postures need each other. A system that says when it's unsure makes checking cheaper. A person who knows where the system was tested knows where to check.

One. The Massachusetts case. The judge gave the parties until the twenty-third of October to report whether they have agreed the fees to be paid, and set a status conference for the ninth of November.

Two. Charlotin's count, which stood at two thousand and ninety-five on the twenty-eighth of September. Look again at the end of October and compare months, not totals, since the newest weeks are likely to be incomplete. Does the monthly pace fall as courts' warnings pile up, or hold above a hundred?

Three. The International AI Safety Report. Last year its first interim update came on the fifteenth of October. If the timing holds, another could arrive within weeks. Its February report named an evaluation gap: existing evaluation methods do not reliably reflect how systems perform in real-world settings.

The idea to keep is that sounding right is not evidence of being right. A fluent answer shows that a system is good at producing fluent answers. The lesson isn't to trust it less on every task. It's to ask what evidence supports this answer, in this setting, and how a mistake would be caught. So: is this task like the ones the system was tested on, or has the world shifted? Does it tell you when it's unsure, and has anyone checked that its confidence means something? And does being good at the thing next to this one tell you anything at all?

To read more: The Parable of Google Flu, by David Lazer and colleagues, in Science, March twenty fourteen. It's three pages, about search data rather than chatbots, which is exactly why it's worth reading.

---

## Sources (21)

- U.S. District Court for the District of Massachusetts, Memorandum and Order on Motion for Sanctions, Civil Action No. 25-CV-12395-AK (Judge Angel Kelley) — 25 September 2026
- Damien Charlotin, *AI Hallucination Cases Database* (page and CSV), damiencharlotin.com/hallucinations — last updated 28 September 2026; read 30 September 2026
- OpenAI, *GPT-4 Technical Report*, arXiv 2303.08774 — March 2023
- Magesh, Surani, Dahl, Suzgun, Manning and Ho, *Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools*, arXiv 2405.20362 — 30 May 2024
- Ginsberg et al., *Detecting influenza epidemics using search engine query data*, Nature 457 — 19 February 2009
- Lazer, Kennedy, King and Vespignani, *The Parable of Google Flu: Traps in Big Data Analysis*, Science 343 — 14 March 2014
- Zech et al., *Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs*, PLOS Medicine — 6 November 2018
- Gwern Branwen, *The Neural Net Tank Urban Legend*, gwern.net/tank — 2011, revised 2023
- Glenn W. Brier, *Verification of Forecasts Expressed in Terms of Probability*, Monthly Weather Review 78(1) — January 1950
- Andrej Karpathy, *Jagged Intelligence*, post on X — 25 July 2024
- Dell'Acqua et al., *Navigating the Jagged Technological Frontier*, Harvard Business School Working Paper 24-013 — 22 September 2023
- Artificial Analysis, *Benchmarking GPT-6 Astra* — 9 September 2026
- Vectara, hallucination leaderboard — 22 September 2026
- Lisanne Bainbridge, *Ironies of Automation*, Automatica 19(6) — 1983
- Parasuraman and Manzey, *Complacency and bias in human use of automation*, Human Factors 52(3) — June 2010
- Becker, Rush, Barnes and Rein (METR), *Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity*, arXiv 2507.09089 — July 2025
- METR, *We are Changing our Developer Productivity Experiment Design* — 24 February 2026
- Kalai, Nachum, Vempala and Zhang, *Why Language Models Hallucinate*, arXiv 2509.04664 — 4 September 2025
- OpenAI, *Why language models hallucinate* (blog) — 5 September 2025
- Bean et al., *Reliability of LLMs as medical assistants for the general public: a randomized preregistered study*, Nature Medicine — 9 February 2026
- International AI Safety Report 2026, arXiv 2602.21012; publications page — 3 February 2026; read 30 September 2026

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
