NewFeatured by AnthropicRead more →
Products
Kepler for FinanceKepler for Industry Soon
Sections
The ProblemApplicationsCareers
Research
Our BlogOur Research
About
Request a DemoLog in
Engineering Research · Research note

Better language models won't fix hallucination. They'll make it quieter.

Every upgrade this year made our models smarter, and after every upgrade we gave them less to do. The better they got, the harder their mistakes were to see.

One of these 1,000 figures is wrong. Kepler shows you the fact behind each one: what it is, when, and where it came from. A spreadsheet of figures, with a magnifying glass over 1,322,500.

Imagine a photocopier that doesn't copy. It reads your page, memorizes it, and retypes it from memory. It's a spectacular typist, and it gets page after page right. Then one day it types a seven where you had a one, and hands the page back with the same confident whirr.

That machine is a large language model. Every number you've seen one write, it retyped.

Do hallucinations still matter? The case for no is easy: better models, lower rates, citations that look real. I think the opposite. They matter more, because they've gone quiet.

A number's digits don't contain its fact. A language model outputs a string, and a string can't say which fact it meant, however good the model is. So every number in an answer needs provenance a machine can check: a route back to the fact it came from. Hallucination is only part of the problem, and it's the part better models shrink. The rest needs a different design. If approximately right is good enough for your product, stop here. In the industries Kepler works in, one wrong figure in 1,000 means someone checks all 1,000.

What makes a fact a fact

A number from LinkedIn's last 10-Q before Microsoft agreed to buy it: 1,322,500. Everything the digits leave out:

  • Currency: US dollars.
  • Scale: thousands, from the table header, so $1,322,500,000.
  • Sign: a credit balance, money LinkedIn owes.
  • Precision: to the nearest thousand, by the filing's own precision tag.
  • Concept: the principal of LinkedIn's convertible senior notes. The label says principal; the tag underneath is named carrying amount.
  • Date: as of March 31, 2016.
  • Entity: LinkedIn Corporation.
  • Location: one tagged fact with its own ID, in the debt footnote.

Change any one and you have a different fact with the same digits. The same filing tags $1,322,500,000 three times: the principal at March 31, 2016; the principal at December 31, 2015; and the face amount on the day the notes were issued, November 12, 2014.

Three facts, one string. Search the filing for 1,322,500 and you get a match whichever of the three you meant. The match tells you nothing about which.

Figure 1
One string, three facts
$1,322,500,000PrincipalAs of March 31, 2016Rounded to the thousandDebtInstrumentCarryingAmountPrincipalAs of December 31, 2015Prior year-end columnDebtInstrumentCarryingAmountFace amountIssued November 12, 2014Tagged as exactDebtInstrumentFaceAmount
Three tagged facts in LinkedIn's Q1 2016 10-Q share one value: March 2016, December 2015, and the day the notes were issued.

A fact is a value plus its context, and the context is most of it. That gives three ways to get a number wrong:

  • Wrong digits. The value changes on its way into the answer: 1,322,500 comes out as 1,332,500.
  • Wrong fact. The digits are right and the context is wrong: thousands read as dollars, December's column reported as March's.
  • Wrong choice. Every fact is right, but the question needed a different one: the balance-sheet debt, when a buyer takes on the principal.

The first two are what people call hallucination. Better language models make them rarer, and harder to see. The third costs the most, and no upgrade fixes it, because the model was never the one who should decide.

A language model has no copy button

Start with hallucination in its simplest form, wrong digits. Pretend the digits are the whole fact, and the model only has to carry 1,322,500 from the filing into its answer. Even that isn't guaranteed, because a language model doesn't copy. It learned something that usually looks like copying.

The networks I learned on couldn't represent a number at all. In Andrew Ng's deep learning course at Stanford, I built a recurrent neural network that scored news articles. Each word entered as a GloVe vector, placed near words used alike: "billion" beside "million." Good for meaning, useless for an exact value: 1,322,500 wasn't in the vocabulary, so it became UNK, the token for every unseen word. The network knew the sentence was about money, not how much.

Summarizers built on those networks wrote UNK wherever a name or a number belonged, so in 2017 the field added a copy button: at each word, a switch between generating and copying from the source. Then subword tokenizers made every string spellable, pretrained Transformers beat the pointer without one, and no language model anyone ships has had a copy button since.

Nothing took its place. What today's models have instead is learned: copying from the context is done by a few attention heads, the parts that look back at the input: under one in twenty in the models studied, and not the same ones from token to token. Switch off the token-level ones and the model paraphrases where it used to copy. There has never been a guaranteed copy inside the model.

Figure 2
Where the copy button went
2014Word vectorsunseen wordsbecome UNK2017Pointer-generatora learned gatecopies from the source2019BARTwins with nocopy mechanism2024Retrieval headsa few headslearn to copyNowReferencesthe copy movesoutside the model
Neural networks that write text have never had a guaranteed copy. The fix moved outside the model.

And a number rarely leaves the way it arrived. The filing prints 1,322,500 under "(in thousands)"; the model writes "$1.3225 billion." OpenAI's tokenizer splits them into 1 | , | 322 | , | 500 and $ | 1 | . | 322 | 5 | billion. Getting from one to the other means moving the decimal and swapping the unit: arithmetic, not copying, done in the same pass that writes the sentence. And the arithmetic is learned the same way. Anthropic traced its own model adding 36 and 59 along two parallel pathways, one for rough magnitude and one for the last digit. Asked how, the model described carrying the one, which it hadn't done.

Figure 3
Same value, different tokens
IN THE FILING, UNDER "(IN THOUSANDS)"1,322,500WHAT THE MODEL WRITES$1.3225billionmove the decimal, swap the unit:arithmetic, not copying
The same principal as OpenAI's tokenizer splits it. Two pieces in common.

So a language model has no copy-paste, only retyping that is usually right. At one slip in 1,000 figures, a 400-figure workbook carries a wrong number every two or three runs, and nothing says which. Training doesn't drive that to zero, because copying was never a primitive.

Even a perfect copier couldn't carry the fact

Now suppose the copy is perfect. That fixes wrong digits and leaves wrong fact untouched.

A language model's only output is a sequence of tokens, each sampled from a distribution conditioned on the tokens before it. Provenance isn't in that channel: no token records the cell it came from. All three of LinkedIn's facts decode to the same tokens, and "as of March 31, 2016" is more tokens, sampled the same way and bound to nothing.

A better copier can't fix that, because the fact isn't a span. The scale sits in the table header, the date in the column heading, the concept in the row label. Copying any one of them is copying. Assembling them into a sentence is generation, and generation is where December's column arrives beside March's digits. Correctness can't live in the digits. It has to live in a pointer to the fact.

Better language models fail quieter

A weak model fails loudly: an UNK where a name belongs is obvious to anyone. A strong model fails with a different right number: December's column instead of March's, the right digits at the wrong scale. Each is a real number from the real filing, so each passes a search.

An August preprint on numerical claims from 10-Ks, with GPT-5.5 among the models tested, found exactly this: models "often produce plausible numerical claims from financial filings while using the wrong reporting period, unit, line item, or formula." A better model doesn't remove this failure. It writes it more fluently.

Some numbers the document can't settle

The third failure, the wrong choice, survives a perfect model and costs the most. Price Microsoft's offer of $196 a share for LinkedIn and you need enterprise value: every diluted share at the price, plus the debt you inherit, minus the cash you keep. Ask the filing what LinkedIn owed on its notes and it gives three numbers, all true:

  • $1.138 billion, the carrying amount on the balance sheet, net of discount
  • $1.197 billion, what the notes were worth on the market
  • $1.3225 billion, the principal LinkedIn had to repay, tagged on the day the notes were issued and again as the same principal at March 31, under a tag literally named "carrying amount"
Figure 4
What did LinkedIn owe?
$1.10B$1.15B$1.20B$1.25B$1.30B$1.35B$1.138BCarrying amountbalance sheet$1.197BFair valuedisclosed in the notes$1.3225BPrincipalwhat a buyer takes on$184 million, with a perfect citation
Three true answers to one question, $184 million apart.

A buyer takes on the principal, because in a takeover the holders can make LinkedIn buy the notes back at 100% of it. Use the balance-sheet figure instead and enterprise value is $184 million off, with a perfect citation. Deal teams know the rule, and a capable model may know it too. Nothing in the answer says which figure it chose. A choice like that belongs written down, where the next reader can see it.

But what about…

"Put the document in context and make it cite." Retrieval puts the right page in front of the model; it doesn't make the answer come from it. The citation features the major labs now ship return a pointer to the passage the model chose. The number in the sentence is still generated beside it. And nothing forces the citation to be chosen for the number: plant a few words of a model's own answer into a document it hadn't cited, and a production retrieval model moves its citation there up to 57% of the time, answer unchanged.

"Hallucination rates are falling. Just wait." They are. OpenAI's own GPT-5.5 card shows the catch: each claim got 23% more likely to be right, but only 3% fewer responses had an error in them, because the model now makes more claims per response. Anthropic's Opus 5 card reports accuracy up 11% and hallucinations up 6% in the same release. A falling rate is good news for a chatbot. An auditor doesn't need the rate; she needs to know which figure is the wrong one, and no rate says.

"Use structured outputs, or a model that can't hallucinate." One launched last week. TypeSafe's Jev gives up strings: every output is a typed value in a schema fixed in advance, which is useful for routing and classification. TypeSafe says it "can't hallucinate," and its post explains the 0% in its charts: "Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots." Schema matching guarantees shape, not truth. A schema can require a number in the debt field and a date in the date field; it can't require the principal, or March 31. Fill the fields with 1,138,264 and December 31, 2015, and the output passes every type check and is still wrong. A citation field changes nothing: it's one more value the model fills in, with nothing tying it to the number beside it.

Each of these stops a step short of the sentence, makes the model a better typist, or constrains the shape of what it types. None proves the right digits landed under the right fact. What proves it is a record, kept outside the model, of where every number came from and which choice put it there. That record is the provenance layer, and the rest of this piece is about building one.

Language models will name numbers, not type them

Software has fixed this shape of problem before. SQL injection came from pasting user input into the query text. Data and code traveled in one string, so a quote mark and a few keywords typed into a form ran as a command. The fix wasn't pasting more carefully. It was parameterized queries, where the value travels separately from the query and is never parsed as code. A language model retyping a number into a sentence is the same single string. References are the parameters.

At Kepler, every tagged figure in a filing is resolved into a fact, keyed to the ID the company itself assigned. The model doesn't type a figure. It names one, [@linkedin-10q-q1-2016-notes-principal], and the system renders the filed value, linked to its line. In the words of Kepler's codebase, that "makes display/value drift unrepresentable": no number can disagree with its source, because the number is never written.

Figure 5
Named, not typed
KEPLERWhat did LinkedIn owe on its convertible notes?LinkedIn owed$1.3225 billionin principal as of March 31, 2016.THE MODEL WRITES[@linkedin-10q-q1-2016-notes-principal]THE READER SEESthe filed value, linked to its lineLINKEDIN 10-Q / CONVERTIBLE SENIOR NOTES(in thousands)Mar 31, 2016Dec 31, 2015Principal1,322,5001,322,500Net carrying amount1,138,2641,126,534THE FACT IT NAMESPrincipal, convertible senior notesAs of March 31, 2016, in thousands of US dollarsDebtInstrumentCarryingAmount
The model names the fact; the system fills in the value and links it to its line.

Calculations take references and return references, so a total three steps deep still opens onto the lines it came from. An unreferenced number is treated like a compile error. That is the work we took away from the model: typing the value, doing the arithmetic, deciding the convention.

By hand, LinkedIn's enterprise value at $196 is an afternoon of pulling share counts, options and debt from two documents and checking each against its page. Ask Kepler and in about 60 seconds you get an Excel workbook: every cell either cited to the fact in the filing or computed by formula, and the debt basis a named setting, principal, written down rather than typed.

Figure 6
The choice, written down
EXCEL WORKBOOK / LINKEDIN ENTERPRISE VALUE AT $196$ in thousands, except per-share and share dataOffer price per share$196.00Merger agreementFully diluted shares143,469,093Agreement, 10-QEquity value$28,119,942= formula+ Debt, at principal$1,322,50010-Q, note+ Redeemable noncontrolling interest$27,32110-Q, balance sheet− Cash and equivalents$759,45110-Q, balance sheet− Short-term investments$2,400,18710-Q, balance sheetEnterprise value$26,310,125= formulaSETTINGSDebt basisPrincipalRedeemable NCIIncludedThe choice is a setting,not a sentence.Switch to carrying amount and theworkbook moves by $184 million,with every input still linked.
Every cell is cited to the filing or computed by formula, and the debt basis is a setting.

Every system that has to defend a number to an auditor will end up here: a store of facts the model can point at but not write. Nobody will choose it for its elegance. They'll choose it after the third time a right number lands under the wrong fact, when retyping stops looking like a shortcut.

What would change my mind

Show me a system that lets the model type its numbers and, on filings full of same-digit traps, lands 400 figures under the right fact a hundred runs in a row, or flags the one it missed before a reader does. I don't expect to see it. A typed number can say where it came from, but then that statement needs checking too. Provenance you take the model's word for is just more output.

Nobody fixed the photocopier by hiring a better typist. They stopped retyping the page.

If this is the kind of problem you want to work on, we're hiring.

Sources

Run this on your own data

Our engineers deploy on each firm's own data, and code pulls, computes, and cites every number.

Request a demo

Work with us

We're hiring engineers to work on problems like this one. Every open role is listed on our jobs page.

See open roles →
Susannah Meyer

Susannah Meyer

Susannah is a Founding Engineer at Kepler, where she built the citation infrastructure that keeps the model from retyping numbers.

LinkedIn →