Latest News & Blogs - cubesys

The corpus eats itself

Written by Paul Heaton | Sep 7, 2026, 3:11:22 AM

 

Last week my own compliance lead told me, in front of the leadership team, that I had contaminated our management system. She was right, and it took her about ninety seconds to prove it.

Her words: "on the 16th of August, you created a bunch of documents. And because they're quite recent documents, it has picked those up and used those as a source of truth over some of the compliance documents. And those ones that you've created are not compliant with our procedure documents. And they're not controlled documents for us."

Then the part that should worry anyone deploying this technology at scale: "even though I specifically said I only look at these folders and I narrowed it down, it picked up on them."

Nothing had been hacked. No permission was wrong. I had written some drafts with AI, they were recent, they were authoritative in tone, and they carried my name. The retrieval layer did exactly what retrieval layers do and promoted them over the documents we are actually audited against.

 

Why this matters now

The whole industry, mine included, has been telling mid-market organisations to get their data ready before deploying AI. Clean the permissions, retire the duplicates, fix the archive. That advice treats your corpus as a fixed inheritance: twenty years of mess, sitting still, waiting to be tidied.

It stopped being fixed the day you deployed. Every AI-assisted document your people produce lands in the same estate the AI reads from. It is well written, confidently phrased and freshly dated, and the two signals retrieval leans on hardest are recency and apparent authority. So the newest, most convincing, least reviewed material in your business is the material most likely to be served back as the answer.

That is not a housekeeping problem. It is a feedback loop, and it runs at machine speed while your document review process runs at human speed.

 

The seniority multiplier

Here is the uncomfortable part, and the reason I am writing this as the offender rather than the adviser. The worst polluter in an organisation is its most senior AI user.

One of my leadership team named it exactly: those documents "are written very authoritative and they say approved by the CEO, he takes that very seriously." My drafts were not wrong. My compliance lead was precise about that too — the material "sounds very authoritative and it's not incorrect. It's just not aligned to our procedures." Which is the harder failure, because nothing in the text tells you anything is off.

Seniority produces volume, and volume produces contamination. The executives who adopt AI fastest generate the most confident, most authoritative, least controlled content in the business, and every piece of it is indexed alongside the documents you are actually held to.

 

It is already showing up outside the tenant

A COO I spoke to this week described the same phenomenon arriving at his desk from the outside world: "like many people, I see a lot of, I won't call it AI slop, but a lot of stuff delivered to me" that has "clearly been written by AI, that the person sitting behind it... hasn't bothered to do anything other than give it a prompt, take the output it delivers, and send it." Job applications, cover letters, proposals. His response is "the delete button or the reject."

Externally, a human reads it and bins it in four seconds. Internally, nobody reads it at all. It is simply indexed, retrieved and quoted back with total confidence to the next person who asks a question.

And we have done this to ourselves before, without AI. An infrastructure manager at a 3,000-person business described a knowledge-base mandate his team ran for a couple of years: "you can't close any ticket off without having a knowledge base article linked to it. So if it doesn't exist, you need to create it then and there. Fantastic, but geez, we ended up with a mess." Duplications, outdated articles, no owner. That was the manual version, at human speed. AI is the same policy with the handbrake off.

 

Where I was wrong

My instinct in that meeting was to push back. My position was that this is a user problem: "that's the person using AI. It's not the data problem. The person using AI has to make a decision about what it's doing. You need to read it." I said we could not audit every single document we create, and I warned against building process around an edge case.

My compliance lead did not accept it, and she was right not to: "it's not about common sense. It's because I know how an audit works." Her point was that a document being technically correct is not the standard. Alignment to controlled terminology is the standard, and misalignment is a non-conformance regardless of who wrote it or how sensible they were being.

Common sense is not a control. It does not scale, it cannot be evidenced, and it fails precisely when the author is senior enough that nobody checks their work. That is the whole argument, and I lost it.

 

What to do instead of more review

The reflex answer is a review gate on AI-generated content. In a small to medium sized organisation that is unbounded work with no owner, and it will lose every budget argument it enters. The volume is the problem; adding human inspection to a machine-speed pipeline is not a fix, it is a queue.

The workable answer is a boundary rather than an inspection. Separate the workshop from the record. Draft space that AI can write into freely and cannot read back from, a single governed location that is the only thing retrieval is pointed at, and one explicit promotion step where a human decides a document has become official. Everything before that step is thinking out loud. Everything after it is evidence.

The question I could not answer in that meeting is the one worth taking away: when does a document become official? Most organisations have never had to define it, because drafts used to sit harmlessly in someone's folder. They do not sit harmlessly any more. They get quoted.

 

The honest test, and it takes ten minutes

Ask your assistant a question you already know the audited answer to, and look at what it cites. If it cites something nobody approved, you have the problem in this article and you found it in one query.

How many documents created in your tenant in the last ninety days are AI-assisted, uncontrolled, and sitting inside a location your AI is indexed against? If nobody can produce that number, the honest position is that you do not know what your assistant is quoting.

Where does a draft physically live in your business, and can retrieval reach it? If the answer is the same place as everything else, you do not have a draft state. You have an unreviewed publication.

 

And if you are a provider, as I am

We sell data readiness as a project with an end date. It is not one, and this is the clearest evidence yet that it never was. Readiness is now a rate: content is entering the estate faster than governance is being applied to it, and the gap widens every week adoption improves. A provider that cleans an archive and declares the client ready has solved the smaller half of the problem and left the growing half untouched.

That is the real content of the shift from managed services to managed intelligence. Not agents, not licences. Being accountable for the quality of what an organisation now produces at scale, including when the person producing it signs your invoices.