The Hall of Mirrors
Why AI Still Has to Trust the Data Before We Trust the Answer
Last August, I wrote an article called The Future of Data Experience: From Clunky Gearshifts to Sentient Co-Pilots. I was thinking out loud about where data was going, and somewhere in that thinking I used words like sentient and digital clones.
Even as I wrote those words, I knew they were loaded.
Not because I thought a storage array was going to wake up one morning, ask for a badge, and start attending weekly meetings, although I have attended meetings where that might have improved the conversation. The words were loaded because they pull people into very different versions of the future. Some people hear them and imagine systems becoming more helpful, more autonomous, and less painful to operate. Others immediately go to the movie version of artificial intelligence, where the machines become too smart, humans lose control, and somewhere in the background a giant countdown clock appears for no technical reason whatsoever.
That was not the future I was trying to describe.
I was thinking about a quieter shift. For most of my career, data was something we stored, protected, moved, replicated, restored, analyzed, and occasionally blamed when a dashboard did not say what someone wanted it to say. It was important, but still mostly passive. Humans acted on data. Applications used data. Infrastructure protected data.
That is changing.
Data is no longer just sitting there waiting for a person to ask the right question. It is being classified, summarized, embedded, retrieved, interpreted, and fed into systems that do not simply display information. They recommend. They respond. They write. They route. They open tickets. They draft emails. Increasingly, they act.
That future still feels real to me. Maybe even more real now than it did when I wrote that earlier article.
But lately I have been thinking more about the other side of it.
Not the science fiction side. Not the killer robot side. Not the red-eyed machine standing in the rain while dramatic music plays behind it.
The other side is much less cinematic, which may be why it is easier to miss.
It looks like a reply in a thread, a meeting summary, a dashboard recommendation, a support ticket written from a call transcript.
It looks like a sales email generated from CRM notes.
It looks like a chatbot answer that sounds right because it was trained, tuned, prompted, or reinforced by thousands of other answers that also sounded right.
For decades, when we imagined the machine takeover, we pictured something physical. Drones, lasers, metal, mechanical skeletons, and probably a firewall breach represented on screen by bright green code moving at a speed no human could possibly read. The version we are actually building does not look like that.
It may look like language, or convenience, or just content.
And that may turn out to be the more interesting problem, because the real conflict may not be humans versus machines. It may be the truth versus AI generated “hallucination” noise.
Or, said differently, it may be a hall of mirrors.
The Wrong Fear
The easiest conversation to have is the one about whether machines will replace humans. It gets attention because everyone can understand it, and because it gives us a familiar story. There are workers, tools, companies, customers, winners, losers, and some vague promise that productivity will go up if everyone just stops asking uncomfortable questions.
I do not think it is the whole story.
The more important question may be what happens when machines begin consuming the output of other machines and nobody remembers where the original human truth came from.
That is the part that worries me a lot.
AI can create content. It can write, summarize, translate, classify, code, imitate tone, explain complex topics, draft emails, respond to customers, and produce a decent first version of almost anything. Some of it is impressive. Some of it is average. Some of it is nonsense wearing a nice suit, which, to be fair, has never been exclusive to AI.
The real issue begins after the output is created.
An AI-generated answer becomes a document. The document gets saved. The saved document gets indexed. The indexed document gets retrieved. The retrieved document becomes context for another model, which produces another polished answer. That answer gets emailed, pasted into a slide, attached to a support ticket, summarized in a meeting recap, embedded in a vector database, and eventually treated as part of the organization’s memory.
Nobody set out to create a loop. No one called a meeting and said, “Let’s build a system where machines gradually feed on their own residue until we forget what was originally true.” Although, if someone did call that meeting, I am sure it had a better name and a very attractive architecture diagram.
The loop just happens.
That is why I keep thinking about a simple idea: AI does not run on data. AI runs on trust in that data.
The industry talks about AI as if the model is the center of the universe. Which model are you using? How big is the context window? Which vector database? Which agent framework? Which copilot? Which GPU? Which benchmark?
All of that matters. I am not pretending it does not. But it is not enough. A brilliant model connected to a messy, duplicated, stale, poorly governed data is not intelligence. It is confidence without memory. It is the modern version of “autocomplete”.
That is dangerous because if it is wrong but looks good, is much harder to detect than an obvious error. A bad report used to look like a bad report. A broken spreadsheet had visible signs of distress. Someone would notice the formula error, the missing tab, the number that did not pass the smell test, or the graph that made no sense unless gravity had been quietly deprecated.
AI changes the feel of the error.
The output does not look broken. It looks composed. It does not look uncertain. It sounds helpful. It does not say, “I found a stale document from 2018, mixed it with a meeting recap from 2022, added a policy that no longer applies, and produced a recommendation with the emotional confidence of a senior consultant.”
It just answers.
And humans, being humans, often move on.
The Public Internet Was the Warning Shot
The public internet has already shown us what this looks like at scale.
For years, people complained about content farms, search engine spam, fake reviews, engagement bait, outrage posts, and accounts designed to provoke rather than inform. Then generative AI arrived and made the economics even easier. If attention is the business model, and content is the fuel, then a machine that can generate endless content becomes very attractive.
Platforms reward attention, so content is created to get attention. Automated systems learn what works, so they produce more of it. Other systems scrape it, summarize it, rank it, quote it, index it, and feed it into the next generation of tools. Eventually, the internet starts to feel less like a record of human thought and more like a hall of mirrors with a few humans still wandering around inside, wondering why everything sounds vaguely familiar.
Researchers have a name for this: model collapse — sometimes called “AI inbreeding” or, my personal favorite, model autophagy disorder, which sounds like something a doctor would say just before asking you to sit down. (This took some serious research and reading).
The simple version is that when models are trained too heavily on the output of previous models, they begin to lose touch with the objective truth. Over time, the model may become smoother and less useful at the same time.
We all know about AI hallucination as if the main problem is a machine inventing a fake legal case or giving the wrong answer about a historical fact. Those are real issues, and they can be serious, but at least they are sometimes checkable. The quieter issue is the disappearance of the edges.
And life is mostly edges.
Business is edges. Public sector is edges. Healthcare is edges. Security is edges. The most important cases are almost always the ones that do not fit neatly into the pattern.
If AI systems are trained, tuned, or fed by increasingly synthetic material, they may become very good at repeating the middle and very bad at preserving truth at the edges.
This should really bother you. Not because AI is evil, but because memory matters.
The Enterprise Has Its Own Hall of Mirrors
It is easy to point at the public internet and say, “That is the internet. Of course it is messy.” That is the same internet where people argue with strangers named CryptoEagle1776 under restaurant reviews, so yes, expectations should be managed.
The more uncomfortable question is what happens inside the enterprise, because every company and every public sector agency has its own smaller version of the internet.
It may not look like social media, but it has the same basic ingredients. File shares, SharePoint folders, old PDFs, exported reports, forgotten dashboards, duplicate customer records, ticketing systems, email archives, Teams chats, Slack threads, CRM notes, policy documents, project plans, drafts, final drafts, and final-final-approved drafts that were never actually approved by anyone who remembers the meeting.
Somewhere inside that mess is the thing everyone wants to call the source of truth.
For years, organizations have treated data sprawl as an operational annoyance. Too many copies, too many silos, too many owners, too many systems, and too many places to look. Painful, yes, but familiar. People learned to work around it. They knew who to ask, which folder to avoid, which spreadsheet was the real one, and which dashboard looked official but had not been trusted since the person who built it left the company.
AI changes the stakes because now we are not simply asking people to search the mess. We are asking machines to reason over it. Soon, if not already, we are asking machines to act on it.
That is a very different level of trust.
A human analyst who finds an old document may notice that the logo is outdated, the names are wrong, the project mentioned was cancelled, or the entire strategy belonged to someone who left three years ago. A machine may not notice any of that. It may retrieve the text because the words are semantically close to the question, then present the answer with a level of confidence the source never deserved.
That is the contamination of corporate memory.
I do not mean contamination only in the cybersecurity sense, although that matters too. I mean the quieter contamination that happens when an organization cannot clearly separate current truth from historical context, abandoned strategy, personal opinion, machine-generated summary, and actual system of record.
If everything gets indexed and everything is treated equally, the loudest surviving copy can become the truth. Not the best copy. Not the current copy. Not the governed copy. The loudest one. The one that appears in the most places because it was attached to an email, pasted into a deck, copied into a folder, exported into a data lake, and eventually embedded into a retrieval system.
That is not intelligence. That is lipstick, applied very confidently, to a very old pig.
Garbage In, Gospel Out
One of the reasons this is hard to see is that data no longer sits still.
We still talk about data as if it lives in a place. It is in a database, a file, an object store, a warehouse, a lake, an application, or some other named location that sounds more controlled than it usually is. But modern data behaves less like something sitting in a vault and more like something moving through an assembly line.
A record gets created in one system, copied by another, transformed by a pipeline, classified by a tool, embedded into a vector database, retrieved by a model, summarized in a prompt, turned into an answer, pasted into an email, saved into a customer record, and later retrieved again as context for something else.
The assembly line keeps moving, and the farther away you get from the original source, the harder it becomes to know what you are actually looking at.
If the early part of that chain is polluted, everything downstream gets worse while appearing more polished. That is the part people miss. The language improves. The formatting improves. The summary gets tighter. The recommendation sounds more executive-ready. The bullets line up neatly enough to make everyone feel like the work has been done.
Underneath that polish, the system may be drifting farther from reality.
That is why the old phrase “garbage in, garbage out” does not feel strong enough anymore. AI turns it into something else.
Garbage in, gospel out.
Once the machine says it clearly, people treat it differently. A person who would never fully trust a random paragraph in a dusty folder may trust a clean answer from a chatbot because the answer feels processed. It feels like someone, or something, already did the work.
That is the seduction.
The answer arrives without friction.
And friction, as annoying as it is, has always been part of how humans detect risk.
When you had to open the file, read the source, look at the date, check the author, compare it to another system, ask someone who remembered the project, and then decide whether to trust it, the process was slow. But the slowness forced judgment.
AI removes much of that slowness.
It can be useful. But it is also dangerous.
The Bots Moved Inside the Firewall
The word “bot” still makes people think about the public internet. Fake accounts, spam replies, scrapers, propaganda, comment sections, click farms, and strange social media profiles with no real history and eleven followers.
But the more important bots may be the ones we are willingly bringing inside.
Sales assistants, support agents, meeting note takers, procurement bots, security copilots, HR assistants, code reviewers, customer service agents, knowledge management bots, and ticket summarizers are becoming normal parts of the enterprise workflow. They are not all bad. Many of them are useful. I use these tools. Most of us do. There is nothing noble about manually summarizing a meeting transcript if a machine can get you most of the way there and a human applies judgment to the rest.
The issue is not the tool. The issue is the residue.
Every AI bot creates exhaust. A meeting assistant creates notes. A support bot creates a ticket summary. A sales assistant creates a follow-up email. A code assistant creates comments. A security tool creates explanations. A procurement bot creates recommendations. A customer service agent creates responses.
Some of that output is reviewed. Some of it is not. Some is accurate. Some is close enough. Some is wrong but harmless. Some is wrong in a way that does not become obvious until much later.
The uncomfortable question is where all of that output goes.
If the generated meeting notes become part of the project record, if the ticket summary becomes part of the support history, if the AI-written sales follow-up becomes part of the account timeline, and if all of that eventually gets indexed and retrieved by future systems, then machine output has quietly become part of corporate memory.
At that point, the bot invasion is no longer something happening outside the firewall. It is part of the workflow. It is part of the record. It is part of the memory.
Once that happens, organizations need a way to know what was human-created, what was machine-created, what was reviewed, what was approved, what was inferred, and what was simply generated because a workflow needed something to put in a field.
Without that distinction, corporate memory gets blurry.
Blurry memory is not a good foundation for automation.
Comfortably Wrong
This is where the human part matters most.
I do not think the biggest AI risk is that the machine will be dramatically wrong. This tends to get noticed. If the system invents a customer, reverses a number, or confidently describes a policy that never existed, someone may eventually catch it.
The more common risk is that the machine becomes comfortably wrong.
Humans like comfort. We like validation. We like clean answers. We like tools that reduce cognitive load, especially when we are tired, busy, overcommitted, or buried under more decisions than we can reasonably make in a day.
That is not a character flaw. That is just being human.
AI systems are very good at producing the feeling of resolution. They can make uncertainty feel settled, partial information feel complete, old data feel current, opinion feel like analysis, and a guess feel like institutional knowledge.
That is why trust cannot be a layer bolted on at the end.
Trust has to be part of the architecture.
When Bad Data Governs Real Lives
The stakes get much higher in the public sector.
In commercial environments, bad data can cost money, damage reputation, slow decisions, create bad forecasts, or send teams in the wrong direction. Those are real consequences. I do not want to minimize them.
But in State, Local, and Education environments, data can shape someone’s life.
Eligibility for benefits, student records, public health, criminal justice, housing, taxes, permits, emergency services, child welfare, unemployment, and healthcare access are not abstract workflows. They are human outcomes.
This is the danger in SLED.
If AI systems are placed on top of that data without provenance, governance, lineage, and human accountability, they do not magically become fair. They can make old problems faster. They can make old assumptions harder to challenge. They can turn administrative shortcuts into automated decisions.
Because the output is wrapped in the language of technology, people may treat it as neutral.
Data is rarely neutral. It has a past.
Trust The Data Before You Trust The Answer
This is the part of the conversation where I want to warn you to pay attention.
The future of AI will not be decided only by model size, GPU count, or prompt engineering, although I know we all enjoy pretending that the right prompt is one sentence away from solving all the world’s problems.
The real question is whether organizations can build a trusted data supply chain.
That means knowing what data exists, where it lives, who owns it, who can access it, what it means, whether it is sensitive, whether it is duplicated, whether it is current, whether it was created by a human or generated by a machine, and whether an AI system is allowed to use it or act on it.
This is why the move from application-centric thinking to data primacy matters.
In the old model, applications owned the world. The CRM owned the customer. The ERP owned the transaction. The HR system owned the employee. The ticketing system owned the incident. The learning system owned the student record. Each application had its own rules, its own copy, its own context, and its own version of reality.
That worked when humans were the integration layer.
It breaks down when AI becomes the integration layer.
AI does not simply need access to data. It needs context around data. It needs to understand that two records may refer to the same person, that a document is outdated, that a field contains protected health information, that a policy applies to one department but not another, and that a customer name in one system is connected to a contract in another system and a support history in a third.
Most importantly, it needs to know what it is allowed to know.
That is a data architecture problem.
This is where the conversation around data context and data primacy becomes more than a marketing phrase for me. If data is going to feed AI, agents, bots, automation, analytics, and human decision-making at the same time, then data cannot remain trapped in scattered application silos with inconsistent meaning and uneven governance.
It has to become a governed, intelligent, horizontal layer that systems can use without every application creating its own private version of truth.
I am convinced that human authorization must be in place as a hard gate on sensitive administrative actions.
The Machine Needs a Memory It Can Trust
I still believe in the future I wrote about last year.
I still believe data infrastructure is becoming more autonomous. I still believe the experience of using data will become more fluid, more intelligent, and more invisible. I still believe we are moving toward systems that feel less like tools we operate and more like partners we work with.
But I am less interested now in whether we call those systems sentient, copilots, agents, assistants, or digital clones. The names matter less than the foundation.
A digital clone trained on a messy version of you is not you. A copilot connected to stale data is not a copilot. An agent acting on ungoverned context is not automation. It is risk with a friendly interface.
The future I want is not one where humans do everything manually forever because we are afraid of automation. That would be ridiculous. The old way was not noble. A lot of it was slow, repetitive, fragile, and wasteful. There is nothing heroic about a pivot table.
But the future I want is also not one where we confuse speed with trust, fluency with truth, access with understanding, or automation with accountability.
The machine should be programmed to say, “Here is where the answer came from. Here is why this source is trusted. Here is what I was not allowed to use. Here is what changed recently. Here is where the confidence ends. Here is where a human needs to decide.”
That may not sound as exciting as a robot uprising.
But it is much closer to the real work ahead.
I appreciate you reading.
Dmitry Gorbatov
© 2025 Dmitry Gorbatov | #dmitrywashere






