Your Context Is Filthy

Data quality used to be judged by what a human could tolerate. Context quality is judged by what a machine will act on without hesitation. Here's why your context is probably dirty.

Published:
Updated:
Your Context Is Filthy

Clean Context

Clean data was about rows. Clean context is about what your agents read before they act.

Everyone agreed on clean data a decade ago. Nobody argues for duplicate records, stale fields, or six versions of the same customer anymore. The argument was won so completely that we forgot it was ever an argument.

It was, though. For years, data quality was the thing everyone nodded at and nobody funded. It took a generation of bad dashboards, failed migrations, and board meetings where two executives showed up with two different revenue numbers before "clean data" stopped being a hygiene chore and became a discipline with owners, budgets, and tooling.

We are about to have the same argument again, one layer up. I'm calling it clean context.

From rows to reading

Clean data was a storage problem. Is this record right? Is it the only one? Is it current? The consumer was a report or a person, and both were forgiving. A human reading a dashboard notices when a number looks wrong. They ask someone. They check.

Context is a different thing. Context is whatever an agent reads in the moments before it acts: your pricing, your positioning, your refund policy, who your customer is, what you decided last quarter and why. It's the company as the agent understands it at the point of action.

And the agent does not ask someone. It does not check. It reads what it's given, treats it as true, and gets on with the job. An agent handed three conflicting versions of your pricing won't flag the conflict. It will pick one, or blend all three, and write the email.

That is the shift. Data quality used to be judged by what a human could tolerate. Context quality is judged by what a machine will act on without hesitation.

Four things make context clean

1. No duplication

One place a thing is true. Not the same policy restated in four decks, a Notion page, and somebody's draft.

Duplication is how companies have always worked, and it was survivable because people carried the reconciliation in their heads. Everyone knew the pricing deck from January was dead and the real numbers lived in a spreadsheet Sarah owned. That knowledge was never written down. It didn't need to be.

Agents don't have it. To a retrieval system, the dead deck and the live spreadsheet are two equally plausible sources. Every duplicate is a coin flip you've handed to a machine.

2. A fact layer

Facts that are curated, verified, versioned, and owned by a human, sitting in the orchestration path rather than scraped back out of documents after the fact.

Each of those words is doing a job:

  • Curated means someone decided this belongs. Not everything your company ever wrote down is a fact.
  • Verified means someone confirmed it is true now.
  • Versioned means you can see what it used to say, when it changed, and who changed it.
  • Owned means a named human is accountable for it staying true.

The last clause matters most. The fact layer has to sit in the path, where agents read from it by default. Most of the industry is trying to do the reverse: leave the mess where it is and reconstruct the truth from it at query time. I've written before that retrofitting a context layer is like reconstructing traffic patterns from dashcam videos. You can't extract discipline from documents that never had any.

3. Clean recall via API

Every agent and every person reads the same facts from the same place, so changing a fact once changes what everything knows.

This is the part that makes the first two pay off. A single source of truth that nothing can reach is just another document. Put it behind an API and it becomes infrastructure: the sales agent, the support agent, the website, the board deck, and the new hire on day one all make the same call and get the same answer.

It also changes what an update is. Today, changing your positioning means a rewrite project across dozens of assets, most of which you'll miss. With clean recall, it's one edit. That's the idea behind building the company itself as an API, and it's why I can change our strategy by updating a single fact.

4. Auditable

For any action an agent takes, you can show which facts it read, which version of each, and who owned them at the time.

This one is coming whether you plan for it or not. Agent decisions now touch every area of a business: pricing, hiring, credit, claims, customer commitments. It is inevitable that regulators, industry bodies, and government authorities will ask for the paper trail behind them. "The model decided" will not be an acceptable answer, and neither will "it was somewhere in the documents we indexed."

You can't produce that trail from a pile of retrieved chunks. You can produce it from a fact layer, because the first three properties make it almost free: one place each fact lives, a version history with a named owner, and a single API where every read can be logged. Auditability isn't a feature you add later. It's what clean context looks like from the outside.

More memory does not equal more truth

Most teams are still trying to solve this with memory. Bigger windows, more retrieval, more documents thrown at the model.

I understand the appeal. It asks nothing of you. You don't have to decide what's true, assign an owner, or delete anything. You connect the drive, index the lot, and let the model sort it out.

But the model can't sort it out, because the problem was never capacity. A bigger window holds more of your contradictions. Better retrieval finds the stale version faster. More memory just gives confusion more room.

This is the same mistake the industry is making at a larger scale, and it's the core of our thesis on how the AI bubble ends: the bottleneck isn't the raw intelligence of the model, it's the absence of a system that reliably delivers the right facts. A probabilistic model fed ambiguous inputs gives you ambiguity with good grammar. Smarter models won't fix that. Cleaner inputs will.

What dirty context costs

Dirty data gave you a bad report. Someone caught it, or they didn't, and a decision got made a little worse than it should have been. The damage was bounded by the speed of the humans in the loop.

Dirty context gives you a thousand agents confidently acting on something that stopped being true in March.

Three things make it worse than the data problem ever was:

  • It executes. A wrong number in a report waits for someone to read it. A wrong fact in an agent's context is already in the outbound email, the quote, the support reply.
  • It multiplies. Every agent you add reads the same mess and draws its own conclusion. Ten agents on fragmented sources don't give you ten times the output. They give you ten versions of your company.
  • It's quiet. Agents don't sound unsure. The output is fluent, well-formatted, and wrong, and nothing in it tells you which fact it was built on.

The more autonomy you hand over, the more your context is your company. Whatever the agents read is, for practical purposes, what the business believes.

Clean context is the work

The uncomfortable part is that none of this is a model problem, so no model release will solve it for you. It's the unglamorous job of deciding what your company knows to be true, writing each thing down once, and putting a name next to it.

A useful test: pick one fact that matters, say your current pricing or your ideal customer. Then ask three questions.

  1. How many places does it live?
  2. Who owns it?
  3. If it changed this afternoon, what would your agents say tomorrow morning?

If the answers are "lots", "nobody", and "the old version", you have a context problem, and adding agents will make it bigger.

The companies that do this work get something more than accurate agents. They get a business that humans and machines can read the same way, and that updates as fast as its facts do. That's what I mean by the living company: an organization that runs on facts, not fiction.

Clean data took a decade to go from nagging to non-negotiable. Clean context won't get that long. The agents are already reading.

At NOAN we're building the fact layer for agentic business. Clean context, every time.