Stop Paying to Forget
Most document AI rereads the same pages for every question and keeps almost nothing from the last read. We read each document once, keep every number, name and relationship we can recover from it, and answer the next question with ordinary software.
Part 1 The $20,000 Question
Retrieval reads too little. Brute force reads everything and then forgets it. Either way, every new question pays to read the documents again.
Most document AI makes reading a runtime expense.
Ask a question and the system goes looking for the pages it thinks matter, then asks a model to read them. Ask a second question and it does the same thing again. The first answer leaves almost nothing behind except prose.
We think that is backwards. We read every filing once, keep every number, name and relationship we can recover from it, and pay about a seventh of a cent per median filing to do it. After that, a question over what has already been read is a lookup, not another reading bill.
We read a 28 MB fund report end to end for about 15 cents. Extracting the same numbers, entities and relationships with a frontier model would cost an estimated $185.1 The fund report is an unusually large document. On a full 10-K, the same comparison is about 5 cents against roughly $34, or around 700 times cheaper.2
The price matters because the argument depends on it. Reading everything has to be cheap enough to do once, up front, and keep.
At that price, you stop choosing which documents deserve reading.
And cheap is not sloppy: on our held-out evaluation, scale is right on 99.4% of tagged facts across 40 filings.
The map
Picture every fact your firm holds as a point on a map, placed by meaning. Everything about one fund sits near the other things about that fund; every covenant, counterparty and quarter has its own neighborhood. A question lights up the points it needs.
A complete answer needs all of them, not just the ten nearest. Ask what one LP pays in fees: the LPA sets the standard rate, and a side letter may give that LP a discount. Miss the side letter and the answer is wrong.
There are two popular ways around this map.
Walking is folders, themes and tags. Someone organized the corpus, which encodes real knowledge, and walking a folder can be exhaustive inside the folder. But every structure has seams. The bubble was drawn by somebody who did not know what you were going to ask.
Teleporting is search, embeddings, RAG, routing and agents. Type a phrase and land near it, anywhere. For discovery it is unbeatable. But search is excellent at finding something and has no idea whether it found everything. The missing result looks exactly like no result.
Query routing is the same failure one level up: a search over searches. An agent is forty teleports instead of one. It still cannot tell you the size of what it skipped.
Context engineering does not help here either. It decides what to do with the pages search found, and has nothing to say about the ones it missed.
More context is not coverage
This matters because the documents are not clean databases waiting to be retrieved.
Apple's FY2025 10-K calls its author "the Company" 789 times, "Apple" 105 times, and "we" or "our" 43 times. The filer tagged 934 numbers as facts, and 909 of them print a string that is not the value: the table says 416,161 and means $416,161,000,000. Net sales are printed six times. Airbnb's S-1 prints 2019 gross booking value as "$38.0 billion" in prose and "37,962.6" in a table headed in millions.
Even the filer's own tags, the nearest thing to an answer key, stop early. Disney tags segment information in the notes while the Domestic, International, Consumer Products and Direct-to-Consumer lines analysts actually model live in the untagged MD&A. Step outside public filings and there is no answer key at all. The deck, the LP report, the credit agreement and the side letter follow their own house styles because nobody wrote them for a parser.
Retrieval finds some of that and misses the rest. A smarter model makes the miss harder to see: an answer built from the pages retrieval happened to find comes back fluent and fully cited, with nothing in it to show which pages were never read.
Brute force has the right instinct and the wrong bill
There is an obvious answer to the coverage problem: stop choosing. Give the model everything.
For a one-off question over a handful of documents, that is often the correct tool. We use frontier models that way ourselves. The trouble starts when the question is not one-off and the set is not small.
Harvey co-founder Gabe Pereyra said in June that a review over 100,000 contracts can cost $20,000.3 That can be a fair price for an exhaustive review. The problem is that the next question starts from zero, because in most systems nothing from the first read survives as structure you can query. The coverage was bought by the token and then thrown away.
A context window is where the model thinks. It is a very expensive place to store a database.
Read once. Keep what you read.
Exhaustive lookup over structured facts is a database problem. Databases solved it before most of us were born. Reading one messy page, with its scale header three paragraphs away and its forty names for the issuer, is a model problem. The mistake is using the model as the database too.
Instead, read every document once (ahead of time if you want search, lazily if you do not) and keep the result. Keep the entities, the relations, the non-numeric facts, and the numbers with their scale, unit, period and position. "The Company" and "Apple" become one thing. A figure printed six ways becomes one value with six addresses and six representations.
That is what we built our reader to do, with no per-form rules, no schema supplied with the document, and one GPU. The seventh of a cent is the median across a test corpus of about five thousand public filings. On the 28 MB fund report, the run took 136 seconds and produced 83,204 facts, 45,199 entities and 9,042 relations.
Those economics matter because cheap changes the architecture. If reading is cheap enough, you stop asking which twelve documents deserve to be read. You read all four hundred and decide which twelve matter afterwards.
We are not alone in betting that more capable models should become better components inside software, rather than a reason to make every operation a chat completion. TypeSafe's Jev, launched this month, is built to make fast, typed decisions that software can use directly.4 It solves a different problem with the same instinct: let the model make the judgment, then give ordinary software something durable to work with.
The public filings are the proving ground because they are documents we can show and, sometimes, documents with a partial answer key. The private documents are the reason the work exists. The most valuable deck or credit agreement is usually the one you are not allowed to paste into somebody else's chatbot.
Give database problems to a database and model problems to a model. Read every document once, keep what you read, and stop paying to forget.
Part 2 looks at what "read it once" means in practice, using a filing that contains an org chart nobody ever drew.
Part 2 The Org Chart Nobody Drew
A 28 MB fund report names a trust, twelve funds, one adviser and more than forty sub-advisers. It never draws the chart. Reading the filing means recovering the structure it assumed you already knew.
The filing names holdings, managers, boards and contracts. It uses full legal names in one section, short names in another, and "the Fund" everywhere else. What it never does is draw the org chart.
The people who run the complex do not need the chart. They already carry it in their heads. The filing assumes the reader will do the same.
That assumption is most of what makes reading hard.
A document is not a bag of numbers
A number on a page is maybe half a fact. Revenue is revenue of a segment. Net assets are net assets of a fund. A holding is held by one fund and managed by one firm. The value matters, but so does the thing it is about and the relationship that puts it there.
Public financial statements give you a shortcut: Inline XBRL tags many numbers with concept, value, scale, unit, period and sometimes segment. We use that answer key where it exists. But it covers only part of the document and none of the private data we ultimately care about. A deck does not tag its numbers. Neither does an LP report, a credit agreement or a data room.
So the reader gets the document with no schema, templates or per-form rules and returns numbers with scale and unit, entities with types, relations between them, non-numeric facts, positions, and a confidence on each. It reads the whole thing once.
One fund, four links
Start with Bridge Builder International Equity Fund. The filing also calls it "International Equity Fund". Olive Street Investment Advisers advises it. Marathon Asset Management sub-advises it. It sits inside the Bridge Builder Mutual Funds family and holds shares in Fortescue. Those statements live in different places in a report hundreds of pages long.
The picture only holds together because "Bridge Builder International Equity Fund" and "International Equity Fund" were recognized as the same thing. The reader made hundreds of same-entity decisions like that in this document. Those decisions, plus some simple legal-suffix cleanup, are what turn the raw adviser names into a list of firms. If the names do not merge, the graph does not exist.
Then the whole shape appears
Across the report, the relationships recover most of the shape. The filing says Olive Street advises all twelve funds; today's extraction links it to nine, still far more than any other firm. T. Rowe Price appears on four. Parametric is the odd one. The adviser table in the notes lists it as one more sub-adviser. About twenty pages later, the board's review describes it as the tax-overlay and direct-indexing manager for exactly the three Tax Managed funds.
No one sentence says, "here is the hierarchy." The hierarchy is the shape left by many individually boring sentences.
One shareholder, three names
Rivian's S-1 gives the same test a different shape. Amazon appears through Amazon.com NV Investment Holdings LLC, Amazon Logistics, Inc., and Amazon Web Services. Those names live in different sections and carry different economic roles. Within that filing, the reader resolves them back to the same parent and recovers Amazon as shareholder, customer and cloud supplier. The filing never says that in one place.
That is why the structure matters. "Which counterparties have more than one role?" sounds like a research request if your only representation of the filing is pages. Once the entities and relations have been read and kept, it becomes a query over those relations.
The hard part was reading the sentences hundreds of pages apart and deciding what they meant. Once that judgment has been made and kept, rerunning the judgment on every question is waste.
Public filings are the demo we can show
The fund report is useful because you can inspect our work. It is also an unusually friendly document compared with the reason we built this.
The same shape sits in documents that never reach a regulator: a fund of funds and its underlying managers; a family office's holdings across custodians; a lender's exposure across a borrower group and its guarantors. Every one arrives as a PDF or deck, describes its hierarchy in prose and signature blocks, and has no answer key.
At a seventh of a cent per median filing, you can afford to read those documents before you know which question will matter. Nobody manually opens four hundred documents in a data room. They open the twelve they think matter and hope.
What it cannot do yet
There is a less flattering diagram we could draw too.
In the current system, the fact and entity layers are still separate. The net assets, expense ratios and NAVs in the fact layer do not yet attach directly to the fund entities in the graph.

Identity also currently stops at the document boundary. "Amazon" resolved across three names inside Rivian's S-1 is one entity. The Amazon in another filing is not yet automatically the same entity. Cross-document identity is a separate layer still in progress.
A pretty graph is not an ontology. The useful version is where "net assets of the Core Bond Fund at June 30, 2026" is one thing you can ask for. Today the numbers are on one side, the names are on the other.
That gap matters because the point of reading once is to leave behind a representation that makes the second question easier than the first.
Part 3 asks what work the model should still be doing once the document has been read. Our answer is less than most systems ask of it.
Part 3 The LLM Is Not the Database
The first question may need a model. The second one usually should not. The whole point of reading a document once is that the judgment survives the read.
Eight milliseconds
We asked the stored output of Rivian's 2021 IPO filing a question we had not extracted for in advance: which shareholders are also commercial counterparties?
The query scanned 4,101 relationships and answered in about eight milliseconds.5 No document was retrieved. No context window was assembled. No model was called.
| Shareholder | What else the filing says |
|---|---|
| Amazon | Customer through Amazon Logistics; cloud supplier through AWS. |
| Ford | Its wholly owned subsidiary Troy Design and Manufacturing built Rivian's prototype and pre-production vehicle bodies and later supplied components. |
| Global Oryx | Its affiliate guaranteed Rivian's $200 million term loan from Standard Chartered Bank. |
| Cox | A Cox Automotive entity had a master services agreement with Rivian. |
Amazon is the answer you would expect. Ford and Global Oryx are the useful ones: nobody named either company in the query. They fell out of a filter over what the filing had already told us. The expensive judgment had happened during the read; the second question was ordinary software.
The role matters, not just the link
The same thing shows up in the fund report. We asked a stranger question: which manager is described in terms of what it does with the other managers' portfolios, rather than simply being named as a manager of a fund?
One firm stood out: Parametric, the tax-overlay manager from Part 2. The board's review says it combines the other sub-advisers' model portfolios, implements their recommendations, and can vary from them for tax management. It does that for exactly the three Tax Managed funds.
The fund questions took about 30 milliseconds in total, again with no model call. What matters more than the speed is that the answer kept enough of the original meaning to distinguish what kind of manager Parametric was, even though the document never gives you one neat field called "role".
Use the LLM where the ambiguity lives. Store the result where the ambiguity does not.
Our first query was wrong
This gets more interesting when the filing changes over time. Our first pass at the fund relationships put Artisan Partners on the Large Cap Value Fund's current roster. The June 30 adviser table said otherwise.
At first that looked like an extraction error. Then we read the source. The board's June review still discussed Artisan's performance on the fund, while the fund's own report says Artisan was removed during the year and the June 30 table no longer lists it. The document was describing two different points in time. Our first query had flattened them together.
Once the question respected the timing in the filing, Artisan moved from current to former. The same query can reconstruct the announced changes to the two Small/Mid Cap funds: the filing says both will be renamed, managers will leave and others will join, with the terminations expected in late September or early October and the broader changes expected by October 28, 2026. The dated query runs in about a tenth of a second with no model call.
| Fund | Leaving | Joining |
|---|---|---|
| Small/Mid Cap Growth → Mid Cap Fund | Eagle; Federated MDTA | Boston Partners; LSV; Vaughan Nelson |
| Small/Mid Cap Value → Small Cap Fund | Boston Partners; LSV; Vaughan Nelson; Diamond Hill | Eagle; Federated MDTA; PIMCO; Stephens |
Both resulting rosters match the filing's subsequent-events note. One part is less clean. The filing names all the changes explicitly, but our reader did not preserve the planned status cleanly for the three Mid Cap additions, so the query has to infer that they are joining. We treat that part as provisional. The useful property is that the remaining uncertainty is visible and debuggable without asking the model to reread 28 MB of text.
The expensive part should happen once
There are things models are unusually good at. "The Fund" can mean the same entity as a long legal name introduced hundreds of pages earlier. One paragraph can describe a former manager while another describes the current roster. A sentence can explain that a firm operates on the other managers' portfolios even though a table elsewhere lists it as just another adviser.
Those are judgment calls over messy text. That is where a model earns its keep.
But once the judgment has been made, keep it. A fact does not become more true because a frontier model rereads the paragraph tomorrow. A company does not stop being the same company because the user phrases the next question differently.
If the second question still requires rereading the same document with the same class of model, the system did not really read it the first time.
Within one filing today
This is also where the current boundary matters. Rivian works because the names can be resolved inside one S-1. The fund queries work because the funds, managers and changes live inside one report. We do not yet have durable identity across every filing in a corpus.
So "ask another question about this already-read filing" is a capability we can demonstrate now. "Run this every morning across every portfolio company" still depends on the cross-document identity layer we have not finished yet.
Building that layer is the next database problem, and it is in progress.
Put the model at the edge
The architecture we want is deliberately asymmetric. At the edge, where the page is messy and the meaning is implicit, use the model. Let it decide what the sentence means.
Then cross the boundary. Keep the result and let ordinary software do what ordinary software is good at: filter it, join it, count it, compare it and ask another question without paying to understand the same sentence again.
The first question can be expensive because somebody has to understand the document. The second question should be cheap because the system already did.
An LLM is a very good reader. It is a terrible database.
Notes and sources
- Using list per-token prices for a frontier model reading the report's 5.3 million characters in 8,000-character windows and writing out every number, entity and relation it finds. Most of that is output; merely reading the text would cost a few dollars, and would leave you with nothing kept. Our figure is the measured model time for this document. ↩
- Estimated for a frontier model reading and extracting a full 10-K at list prices, against about 5 cents for our reader on the same filing. Across filings the ratio ranges from roughly 450x to 1,100x. ↩
- Harvey token economics: Sourcery, "BREAKING: Harvey Co-Founder & Head of Applied Research on the Token Reckoning", June 2026. ↩
- Purpose-built decision models: TypeSafe, "Introducing System One Models & Jev". ↩
- Timings are medians of repeated warm runs on a laptop, excluding file loading and process startup: Rivian 8.1 ms; fund questions, including Parametric, 30.1 ms; dated fund roster 112.5 ms. None made a model call at query time. ↩
Filings
- Apple Inc., Form 10-K (fiscal 2025).
- The Walt Disney Company, Form 10-K (fiscal 2025).
- Airbnb, Inc., Form S-1 (2020).
- Bridge Builder Trust, Form N-CSR (year ended June 30, 2026): adviser table p. 568; pp. 577, 579 and 584; board review of Parametric pp. 586-588; plus the Large Cap Value, Small/Mid Cap Growth and Small/Mid Cap Value annual shareholder reports.
- Rivian Automotive, Inc., Form S-1 (2021), pp. 171-174.
