Better language models won't fix hallucination. They'll make it quieter.
Every upgrade this year made our models smarter, and after every upgrade we gave them less to do. The better they got, the harder their mistakes were to see.

Imagine a photocopier that doesn't copy. It reads your page, memorizes it, and retypes it from memory. It's a spectacular typist, and it gets page after page right. Then one day it types a seven where you had a one, and hands the page back with the same confident whirr.
That machine is a large language model. Every number you've seen one write, it retyped.
Do hallucinations still matter? The case for no is easy: better models, lower rates, citations that look real. I think the opposite. They matter more, because they've gone quiet.
A number's digits don't contain its fact. A language model outputs a string, and a string can't say which fact it meant, however good the model is. So every number in an answer needs provenance a machine can check: a route back to the fact it came from. Hallucination is only part of the problem, and it's the part better models shrink. The rest needs a different design. If approximately right is good enough for your product, stop here. In the industries Kepler works in, one wrong figure in 1,000 means someone checks all 1,000.
What makes a fact a fact
A number from LinkedIn's last 10-Q before Microsoft agreed to buy it: 1,322,500. Everything the digits leave out:
- Currency: US dollars.
- Scale: thousands, from the table header, so $1,322,500,000.
- Sign: a credit balance, money LinkedIn owes.
- Precision: to the nearest thousand, by the filing's own precision tag.
- Concept: the principal of LinkedIn's convertible senior notes. The label says principal; the tag underneath is named carrying amount.
- Date: as of March 31, 2016.
- Entity: LinkedIn Corporation.
- Location: one tagged fact with its own ID, in the debt footnote.
Change any one and you have a different fact with the same digits. The same filing tags $1,322,500,000 three times: the principal at March 31, 2016; the principal at December 31, 2015; and the face amount on the day the notes were issued, November 12, 2014.
Three facts, one string. Search the filing for 1,322,500 and you get a match whichever of the three you meant. The match tells you nothing about which.
A fact is a value plus its context, and the context is most of it. That gives three ways to get a number wrong:
- Wrong digits. The value changes on its way into the answer: 1,322,500 comes out as 1,332,500.
- Wrong fact. The digits are right and the context is wrong: thousands read as dollars, December's column reported as March's.
- Wrong choice. Every fact is right, but the question needed a different one: the balance-sheet debt, when a buyer takes on the principal.
The first two are what people call hallucination. Better language models make them rarer, and harder to see. The third costs the most, and no upgrade fixes it, because the model was never the one who should decide.
A language model has no copy button
Start with hallucination in its simplest form, wrong digits. Pretend the digits are the whole fact, and the model only has to carry 1,322,500 from the filing into its answer. Even that isn't guaranteed, because a language model doesn't copy. It learned something that usually looks like copying.
The networks I learned on couldn't represent a number at all. In Andrew Ng's deep learning course at Stanford, I built a recurrent neural network that scored news articles. Each word entered as a GloVe vector, placed near words used alike: "billion" beside "million." Good for meaning, useless for an exact value: 1,322,500 wasn't in the vocabulary, so it became UNK, the token for every unseen word. The network knew the sentence was about money, not how much.
Summarizers built on those networks wrote UNK wherever a name or a number belonged, so in 2017 the field added a copy button: at each word, a switch between generating and copying from the source. Then subword tokenizers made every string spellable, pretrained Transformers beat the pointer without one, and no language model anyone ships has had a copy button since.
Nothing took its place. What today's models have instead is learned: copying from the context is done by a few attention heads, the parts that look back at the input: under one in twenty in the models studied, and not the same ones from token to token. Switch off the token-level ones and the model paraphrases where it used to copy. There has never been a guaranteed copy inside the model.
And a number rarely leaves the way it arrived. The filing prints 1,322,500 under "(in thousands)"; the model writes "$1.3225 billion." OpenAI's tokenizer splits them into 1 | , | 322 | , | 500 and $ | 1 | . | 322 | 5 | billion. Getting from one to the other means moving the decimal and swapping the unit: arithmetic, not copying, done in the same pass that writes the sentence. And the arithmetic is learned the same way. Anthropic traced its own model adding 36 and 59 along two parallel pathways, one for rough magnitude and one for the last digit. Asked how, the model described carrying the one, which it hadn't done.
So a language model has no copy-paste, only retyping that is usually right. At one slip in 1,000 figures, a 400-figure workbook carries a wrong number every two or three runs, and nothing says which. Training doesn't drive that to zero, because copying was never a primitive.
Even a perfect copier couldn't carry the fact
Now suppose the copy is perfect. That fixes wrong digits and leaves wrong fact untouched.
A language model's only output is a sequence of tokens, each sampled from a distribution conditioned on the tokens before it. Provenance isn't in that channel: no token records the cell it came from. All three of LinkedIn's facts decode to the same tokens, and "as of March 31, 2016" is more tokens, sampled the same way and bound to nothing.
A better copier can't fix that, because the fact isn't a span. The scale sits in the table header, the date in the column heading, the concept in the row label. Copying any one of them is copying. Assembling them into a sentence is generation, and generation is where December's column arrives beside March's digits. Correctness can't live in the digits. It has to live in a pointer to the fact.
Better language models fail quieter
A weak model fails loudly: an UNK where a name belongs is obvious to anyone. A strong model fails with a different right number: December's column instead of March's, the right digits at the wrong scale. Each is a real number from the real filing, so each passes a search.
An August preprint on numerical claims from 10-Ks, with GPT-5.5 among the models tested, found exactly this: models "often produce plausible numerical claims from financial filings while using the wrong reporting period, unit, line item, or formula." A better model doesn't remove this failure. It writes it more fluently.
Some numbers the document can't settle
The third failure, the wrong choice, survives a perfect model and costs the most. Price Microsoft's offer of $196 a share for LinkedIn and you need enterprise value: every diluted share at the price, plus the debt you inherit, minus the cash you keep. Ask the filing what LinkedIn owed on its notes and it gives three numbers, all true:
- $1.138 billion, the carrying amount on the balance sheet, net of discount
- $1.197 billion, what the notes were worth on the market
- $1.3225 billion, the principal LinkedIn had to repay, tagged on the day the notes were issued and again as the same principal at March 31, under a tag literally named "carrying amount"
A buyer takes on the principal, because in a takeover the holders can make LinkedIn buy the notes back at 100% of it. Use the balance-sheet figure instead and enterprise value is $184 million off, with a perfect citation. Deal teams know the rule, and a capable model may know it too. Nothing in the answer says which figure it chose. A choice like that belongs written down, where the next reader can see it.
But what about…
"Put the document in context and make it cite." Retrieval puts the right page in front of the model; it doesn't make the answer come from it. The citation features the major labs now ship return a pointer to the passage the model chose. The number in the sentence is still generated beside it. And nothing forces the citation to be chosen for the number: plant a few words of a model's own answer into a document it hadn't cited, and a production retrieval model moves its citation there up to 57% of the time, answer unchanged.
"Hallucination rates are falling. Just wait." They are. OpenAI's own GPT-5.5 card shows the catch: each claim got 23% more likely to be right, but only 3% fewer responses had an error in them, because the model now makes more claims per response. Anthropic's Opus 5 card reports accuracy up 11% and hallucinations up 6% in the same release. A falling rate is good news for a chatbot. An auditor doesn't need the rate; she needs to know which figure is the wrong one, and no rate says.
"Use structured outputs, or a model that can't hallucinate." One launched last week. TypeSafe's Jev gives up strings: every output is a typed value in a schema fixed in advance, which is useful for routing and classification. TypeSafe says it "can't hallucinate," and its post explains the 0% in its charts: "Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots." Schema matching guarantees shape, not truth. A schema can require a number in the debt field and a date in the date field; it can't require the principal, or March 31. Fill the fields with 1,138,264 and December 31, 2015, and the output passes every type check and is still wrong. A citation field changes nothing: it's one more value the model fills in, with nothing tying it to the number beside it.
Each of these stops a step short of the sentence, makes the model a better typist, or constrains the shape of what it types. None proves the right digits landed under the right fact. What proves it is a record, kept outside the model, of where every number came from and which choice put it there. That record is the provenance layer, and the rest of this piece is about building one.
Language models will name numbers, not type them
Software has fixed this shape of problem before. SQL injection came from pasting user input into the query text. Data and code traveled in one string, so a quote mark and a few keywords typed into a form ran as a command. The fix wasn't pasting more carefully. It was parameterized queries, where the value travels separately from the query and is never parsed as code. A language model retyping a number into a sentence is the same single string. References are the parameters.
At Kepler, every tagged figure in a filing is resolved into a fact, keyed to the ID the company itself assigned. The model doesn't type a figure. It names one, [@linkedin-10q-q1-2016-notes-principal], and the system renders the filed value, linked to its line. In the words of Kepler's codebase, that "makes display/value drift unrepresentable": no number can disagree with its source, because the number is never written.
Calculations take references and return references, so a total three steps deep still opens onto the lines it came from. An unreferenced number is treated like a compile error. That is the work we took away from the model: typing the value, doing the arithmetic, deciding the convention.
By hand, LinkedIn's enterprise value at $196 is an afternoon of pulling share counts, options and debt from two documents and checking each against its page. Ask Kepler and in about 60 seconds you get an Excel workbook: every cell either cited to the fact in the filing or computed by formula, and the debt basis a named setting, principal, written down rather than typed.
Every system that has to defend a number to an auditor will end up here: a store of facts the model can point at but not write. Nobody will choose it for its elegance. They'll choose it after the third time a right number lands under the wrong fact, when retyping stops looking like a shortcut.
What would change my mind
Show me a system that lets the model type its numbers and, on filings full of same-digit traps, lands 400 figures under the right fact a hundred runs in a row, or flags the one it missed before a reader does. I don't expect to see it. A typed number can say where it came from, but then that statement needs checking too. Provenance you take the model's word for is just more output.
Nobody fixed the photocopier by hiring a better typist. They stopped retyping the page.
If this is the kind of problem you want to work on, we're hiring.
Sources
- LinkedIn Corporation, Form 10-Q for the quarter ended March 31, 2016, and its XBRL instance. The three tagged facts are
DebtInstrumentCarryingAmountat March 31, 2016 and December 31, 2015, andDebtInstrumentFaceAmountat November 12, 2014. - See, Liu and Manning, "Get To The Point: Summarization with Pointer-Generator Networks", ACL 2017.
- Wu et al., "Retrieval Head Mechanistically Explains Long-Context Factuality", ICLR 2025.
- Lin et al., "Retrieval Heads are Dynamic", ACL 2026.
- Feucht, Todd, Wallace and Bau, "The Dual-Route Model of Induction", COLM 2025.
- Lindsey et al., "On the Biology of a Large Language Model", Anthropic, March 2025.
- Hall, Shome and Eiers, "VeriFin: A Neurosymbolic Framework for Verifying LLM-Generated Financial Claims", preprint, August 2026.
- Wallat, Heuss, de Rijke and Anand, "Correctness is not Faithfulness in Retrieval Augmented Generation Attributions", ICTIR 2025.
- OpenAI, GPT-5.5 System Card, April 2026; Anthropic, Claude Opus 5 System Card, July 2026.
- TypeSafe, "Introducing System One Models & Jev", September 2026.
- rain.forest.puppy, "NT Web Technology Vulnerabilities", Phrack 54, December 25, 1998.
