Your Mistakes Are a Dataset
When a system gets something right for the 10,000th time, you don't learn much. When it gets something wrong for the first time, you might have found something nobody knew to test for.

When a system gets something right for the 10,000th time, you don’t learn much. When it gets something wrong for the first time, you might have found something nobody knew to test for.
I’ve been thinking about this a lot while building at Kepler. Working with messy, real-world information forces you to get pretty specific about what “wrong” actually means.
There are a lot of ways these systems can fail without obviously looking broken. Sometimes a system retrieves exactly the right number, just from the wrong period. Or it identifies every entity correctly but gets the relationship between them wrong. Or it finds the right source, gives the right answer, and somehow the citation still doesn’t support the claim. Public-market data has a million versions of this: an actual, management guidance and an analyst estimate can all refer to “revenue,” all be legitimate numbers, and all mean completely different things. Retrieval isn’t really the problem there. You need to know what the number is.
Relationships are even easier to get subtly wrong. A researcher, a company and a university can show up repeatedly across the same set of documents, and a system can identify all three perfectly while incorrectly inferring that the researcher works for the company. Nothing has been made up. Every entity exists, every source is real, and you can still end up with a false picture of what happened.
I find these failures much more useful than obvious hallucinations because they expose assumptions you didn’t realize you were making. And once you find one, why throw it away?
We’ve started treating these cases as things worth keeping. Save what went in, what was retrieved, what came out, what should have happened and why it was wrong. That failure can become an eval, a hard negative, a regression case, maybe training data. Sometimes it just tells you that you’ve designed something badly. Not every bad output means the model needs fixing either. Plenty of “model problems” turn out to be retrieval problems, stale context, entity resolution, a bad assumption upstream, or perfectly good evidence being combined badly.
Agents make this messier. Once a system is doing more than returning an answer, you have the whole path to care about. It can retrieve the right things, make sensible tool calls and still end up somewhere you didn’t want it to go. It can also take a terrible path and somehow arrive at the correct answer. Looking only at the final output hides both.
So lately I’ve cared less about whether a system can pass the same test again and more about what happens when it surprises us for the first time. Can we turn that surprise into a case it never gets to surprise us with again? Do that for long enough and your eval set starts looking less like a benchmark and more like a history of everything you’ve learned the hard way: the weird query someone spent half a day debugging, the relationship nobody thought would be ambiguous, the tool call that looked completely reasonable until you saw what happened three steps later.
It’s basically organizational scar tissue you can run.
I also think these eval sets are becoming a competitive advantage. Models change, infrastructure changes, and a lot of the stack is available to everyone. A good eval set isn’t. It accumulates slowly from production: edge cases, user corrections, failures nobody anticipated and distinctions that only become obvious after enough time in a domain. Two teams can use roughly the same models and tooling and end up with very different systems because one has spent years turning those surprises into tests. In that sense, an eval set starts to look a lot like proprietary data.
Playing with Jev has made me think about the other side of this: if these failure sets are valuable, what should actually be open-source? Open-source AI shares plenty of models, frameworks, datasets and infrastructure, but not many of the embarrassing bits. Most failures get fixed, turned into an internal test if we’re disciplined, and disappear into a repo or Slack thread. There are good reasons for that. A collection of failures is accumulated domain knowledge, and giving all of that away isn’t obviously smart. But the opposite is also kind of absurd: how many teams have independently discovered the same retrieval failure or built the same regression test?
I’d genuinely like to see an open dataset that’s just a graveyard of things that broke good systems.
Give me the weird query that fooled every retriever you tried, the relationship every model confidently got wrong, the citation that looked perfect until someone actually read it, the agent run where every step looked fine and the result was a disaster. For each one: here’s what happened, here’s why it looked reasonable, and here’s why it was wrong.
Every tombstone is a test somebody else doesn’t have to discover the expensive way.
I’m not convinced all of this should be open. Failures specific to your domain may be some of your most useful proprietary data. But there are plenty of general failures where keeping them private mostly guarantees that somebody else wastes a Tuesday discovering the same thing.
That’s also why static benchmarks only tell me so much. By definition, somebody already knew which questions to put on the test. The production failures I care about are the ones that weren’t on the test because nobody knew to ask the question yet. As systems get better, those should get rarer and weirder, which is exactly why I want to keep them.
Your mistakes are a dataset. Treat them like one.