Visualização de leitura

AI is not ready to answer questions about your data

Every other week, there’s a new story about AI solving a math problem that took mathematicians decades to touch. Most recently, it was a conjecture that had sat unsolved since 1939, cracked with an assist from an AI model. Meanwhile, the typical team still can’t ask ChatGPT what we sold last week without pulling three dashboards and arguing over whose number is right. That begs the question: why is AI intelligent enough to outsmart MIT professors, but it can’t outsmart the sales intern?

Here’s part of the answer. AI doesn’t get to guide a real business decision until it can answer with real accuracy, not 95%, not 99%, all the way. Getting there means clearing three hurdles: context, determinism and cost. Solve those three and you get accuracy. Right now, none of them are solved, which is why accuracy, not intelligence, is the actual blocker nobody wants to admit.

Grounding is the first wall

Ask an AI model a question about your company and it doesn’t actually know your company. It doesn’t know who owns which decision. It doesn’t know why your warehouse manager overrides the forecast every October, or which of two conflicting reports your team trusts. A new hire picks this up by working somewhere long enough. AI needs it handed over, deliberately, in layers.

The first layer is your teams, roles and workflows: who does what, and who is allowed to change it. The second is your industry and business context; the reason a service outage means something different for a bank than it does for a media company. The third is your data itself, the tables, definitions and history that make an answer true rather than plausible. All three have to be handed over deliberately. None of them show up on their own.

What I have observed is that companies skip straight to buying an agent and skip this grounding work entirely, which doesn’t end well. The agent still sounds confident. It’s also wrong in ways nobody catches until a decision has already been made on top of it.

Businesses do not want a coin flip

A client once told me something I have not stopped thinking about. We proposed solving their facility allocation problem with a classic optimization algorithm, deterministic and auditable, the kind where the same input always produces the same output. They were disappointed, stating, “I hired you because you are AI experts. Classic optimization gives us deterministic results. We want AI that gives us indeterminate results.”

I understood what they meant. They wanted something that felt like AI. But indeterminate is not a feature you want in a system telling you how much inventory to order. In the pursuit of using AI for the sake of using AI, it is easy to lose sight of the outcome the business actually needs.

That conversation still shapes how I scope every new engagement. When a client asks for AI without naming the decision it needs to support, I ask what happens the day it’s wrong. If the answer is a bad number in a board deck, we’re not talking about the same kind of AI they think they’re asking for. Sometimes that conversation ends the engagement before it starts. More often, it reshapes it into something smaller and more useful, an agent that handles the easy 80% of questions and flags the rest for a human, instead of one system trying to do everything at once.

Here is a number to think about. Anthropic recently published how it automated its own internal business analytics queries using Claude, and even with a purpose-built system, aggregate accuracy landed around 95%. That is one of the best AI labs in the world, building for its own internal use, still getting roughly one answer in twenty wrong. Tell any CFO that number and watch how fast they walk back to their deterministic BI dashboard.

I don’t think 95% is close enough. Not when the number ends up in a board deck. Not when a wrong answer becomes a real decision with real money behind it. In my experience, executives will forgive AI for being slow to learn. They will not forgive it for being confidently wrong. AI doesn’t earn a seat at the decision table by being right most of the time. It earns it by being right every time, the same way the deterministic system it’s replacing was. That’s not a popular thing to say in a market excited about what AI can do, but popularity was never the bar a business decision needed to clear.

Cost swings before the ROI math holds still

The variance in the cost of running an AI agent is enormous. I have watched the same agent perform the same task use up to 30 times as many tokens in one run as in another. Try defending a Well Architected Framework, or any ROI model, against a cost that swings that hard.

The trend lines pull in opposite directions at the same time. The price per token has generally been falling, which should make agents cheaper over time. At the same time, multi-agent architectures are burning through more tokens to do the same job. And an agent genuinely grounded in your teams, industry context and data — the grounding I described above — will use even more tokens than a shallow one. Full accuracy costs more, not less. The good news, if the trend holds, is that this improves rather than worsens over time. But “if the trend holds” is doing a lot of work in that sentence.

Not every question your business asks needs the same level of AI maturity. How many customers are in the system is easy, and I’ll trust an agent’s answer. How revenue has changed over time is medium. Forecasting future demand by product, week and store gets harder. Figuring out how to allocate demand across factories and manufacturing lines is super hard. Crafting a strategy that takes into consideration the three layers of context plus macroeconomic trends is, honestly, impossible for AI or a human to answer with certainty. The higher you climb that ladder, the more a wrong answer costs you, and the less I’m willing to accept anything short of fully right.

The mistake I see most often is a company jumping from easy straight to hard, expecting an agent that can answer “how many customers do we have” to also handle causal questions like “why did we stop selling a product.” Context requirements, accuracy requirements and token cost all climb together as you move up that ladder. That’s also the difference between handing an agent a task and handing it a role; the judgment a planner, buyer or analyst brings to a job every day sits at the top of the ladder, not the bottom.

The mathematician who cracked that 87-year-old conjecture and the CFO who wants last week’s sales are asking for different things. The mathematician wanted a collaborator: someone to try a thousand wrong paths and surface one interesting idea, where a 5% hit rate is a triumph. The CFO wants a number she can put in front of a board, where a 5% miss rate is a liability. AI has earned the first seat. It hasn’t earned the second.

It will, but not by getting smarter. It gets there when someone does the unglamorous work of grounding it in the company’s context, constraining it to be right the same way twice, and paying for that at a cost the ROI math can survive. Intelligence was never the blocker. Accuracy is. So, until an agent can clear all three hurdles, give it the easy 80% of the ladder, keep a human on the rest and don’t let the fact that it disproved a conjecture convince you it’s ready to order your inventory.

The crisis of synthetic culture

For most of the IT era, technology leaders have treated information as an asset. It is something to be stored, secured, processed and monetized. Organizations generated vast amounts of data, and technology processed that wealth and made it useful.

Then AI arrived, and the framing broke.

AI changes the relationship with information. It interprets it, compresses patterns within it, and generates new language from those patterns. It does so with such fluency that it converts accumulated human expression into outputs that are coherent, meaningful, and often persuasive in ways that are difficult to examine or trace. The arrival of AI is a deeper shift in how organizations produce language, remember knowledge, establish authenticity and decide what deserves trust.

Generally, the discussion about AI’s impact on language, memory and meaning gets swept into the social or philosophical bucket too quickly. But these are not just philosophical questions. They are enterprise questions, because they directly affect knowledge management, brand trust, customer engagement, regulatory exposure, institutional memory, decision-making and employee learning. CIOs who brush this dimension aside will govern AI infrastructure competently and miss its deeper institutional consequences entirely.

Language is humanity’s most powerful mechanism for collective learning. We speak not only about what is present, but about what is absent, imagined, remembered, feared and hoped for. That capacity is what allowed us to preserve experience, transmit it across generations and convert it into culture, knowledge, law, philosophy, science and enterprise memory.

CIOs must govern meaning, not just data

This is why large language models are consequential in a way that earlier software was not. LLMs train on enormous volumes of text. They learn patterns, structures, associations, idioms and contextual relationships within language. By doing this, they create a dynamic simulation of human language itself. Unlike traditional software that executes defined instructions, LLMs operate within encoded human expression. They observe documents, reports, messages and public knowledge, and they generate responses that sound almost like understanding.

Almost. And that almost is where the situation gets hairy.

The power is real. AI can widen access to sophisticated knowledge. It can digest complexity, translate across domains and reveal patterns invisible to a single analyst. Inside an organization, the impact can be significant: a junior employee can consume decades of internal documents in minutes; a manager can synthesise thousands of customer interactions before a single meeting; a compliance team can identify patterns across audit data that would take months by hand. Knowledge work has a new definition.

But it is precisely this power that creates the crisis of synthetic culture.

While writing “The AI Codex: Power, Ethics and the Human Future in the Age of Intelligent Machines,” I chose to leave the engineering largely aside and follow the subtler human consequences that are easier to ignore.

Synthetic culture is one of those consequences, and it should concern CIOs immensely.

Culture is not office decoration, value posters or elaborate town halls. Culture is how an organization thinks, decides, explains, rewards and justifies its actions. It lives not only in documents but in stories, habits, unwritten rules, leadership behaviours, institutional scars and the memory of battles won or lost. It is what people have lived through and passed on. Your document management system and knowledge management system are not the custodian of your culture. The people around you are.

AI preserves the digital residue of culture. It cannot preserve the lived meaning behind it. That distinction matters enormously.

Consider what happens in practice. If AI generates a memo in the style of a CEO, does it carry the judgment of that leader, or does it merely carry the pattern of her language? When generative AI summarises a complex customer dispute, does it carry the flavour of the relational history, or does it just compress text? When it drafts a new policy, does it reflect institutional accountability, or does it mimic the average structure of similar policies? When AI produces a cultural narrative for employees, is it transmitting memory or manufacturing a convincing imitation of the same?

These are not rhetorical questions. They describe a genuine ambiguity that is already embedded in enterprise operations.

AI takes what is human, learns from it, and produces something that comes close to human expression. But proximity is not identity. AI can sound human without being human. It can produce language without emotion. It can generate meaningful output without moral agency. It can create efficiency without bearing responsibility for the consequences.

And this is where authenticity begins to fracture.

There was, until recently, something like a one-to-one relationship between content and its source. A document had an author. A speech had a speaker. A photograph recorded an event. A policy had an accountable authority behind it. These relationships were not perfect, but they were there, and they allowed organizations and societies to trace meaning back to a human being who could be questioned, challenged or held responsible.

AI can weaken or obscure that link. It can produce synthetic truths that spread rapidly, appear credible and influence decisions without being anchored to any lived reality. Deepfakes are the obvious example. But the real danger is subtler and already present inside organizations: the synthetic summary, the automated narrative, the AI-generated recommendation that nobody can fully trace. A synthetic truth is not a lie. But it can be more dangerous than one, because it poses as a plausible construction with no real thing behind it.

The new enterprise risk: synthetic truth

For CIOs, this is a new category of risk for which many organizations have no mature controls. The question is no longer only whether organizational data is secure. The question is also whether organizational meaning is secure. Can employees distinguish between original knowledge and synthetic synthesis? Can customers trust that they are interacting with accountable institutional communications, or with an automated approximation? Can leadership trace the lineage of a recommendation or a decision?

These are operational questions. The failure mode is not just a data breach but more of a memory breach.

When AI systems generate from enterprise knowledge, they shape what the organization remembers and how it remembers it.

AI is changing what organizations remember

The consequences follow from the quality of what goes in. If the underlying data reflects poor documentation, AI amplifies that poverty. If institutional knowledge records only dominant voices, dissenting experience is subdued and eventually forgotten. If past mistakes have been quietly buried, AI may reproduce the organization’s confidence without preserving its caution. The organization becomes more efficient at forgetting what it should have remembered.

CIOs also need to rethink what knowledge management means. For years, KM was treated as a repository problem: store, tag, search, retrieve. AI retrieves knowledge and generates new formulations from it. Each time a model runs, it can produce a slightly different answer. Knowledge becomes fluid and unstable. A document may be old, but it is static and accountable. An AI-generated answer may be elegant but untraceable. Both can exist in the same organization, and most people cannot tell them apart.

Those who complained about information overload in the internet age have no idea about the blizzard the AI age is about to bring. AI can identify patterns invisible to humans, but it can also manufacture alternative truths that are difficult to challenge. It can reduce noise, but it can also generate noise at industrial scale.

This is why AI governance cannot be a downstream compliance exercise. It must be a first-principle commitment. Not just checking whether the model works but asking what kind of institutional memory the model is helping to create. Where did the information come from? Who approved its use? What has been included, and more importantly, what has been excluded? Where does the audit trail begin? Where does human judgment remain mandatory and non-negotiable? These questions belong on the Post-it notes sitting on every CIO’s desk as the AI agenda gathers speed.

The search for truth cannot be a human pursuit alone in this environment, but it cannot be outsourced to machines either. It must be a governed exercise, with explicit architecture and explicit accountability.

The crisis of synthetic culture will not announce itself dramatically. It will arrive quietly, in the form of convenience: automated memos, summarised knowledge, AI-generated reports that nobody has the time or the inclination to question. Machines will not become human. But organizations will gradually grow comfortable accepting machine-generated statistical approximations as human judgment, institutional memory or cultural truth. That comfort is the real risk.

The CIO now carries an institutional mandate: scale intelligence without surrendering trust. Productivity matters and must be pursued. But the real challenge is building trust while scaling intelligence. This means ensuring AI output is traceable, sources are visible, human authority is explicit and institutional memory is always protected from synthetic distortion.

CIOs have to become not just custodians of organizational systems, but of organizational memory. In the new Badlands of the AI age, the CIO is the morally upright gunslinger. Organizational memory is the line she must defend.

Never mind clean data. Annotate as you collect it.

Generative AI is notoriously eager to help, to the point that if it can’t find something matching what you ask for, it’ll create it. So the problem with relying on guardrails is that all too often, a model will be wrong, showing a high confidence score for an incorrect answer because it’s relying on stale or non-canonical data.

Not only do you need to be able to track the lineage of data your model uses from source to token, something the EU AI Act requires, you also need to be able to take into account where the data came from, whether it’s out of date, if it changed in a way that affects the result, or if it was never really relevant or authoritative in the first place.

Gartner expects organizations will abandon 60% of AI projects because they don’t have the right metadata management, data quality, and data observability. IBM’s acquisition of Confluent also highlights the importance of real-time data with lineage, governance, and policy for AI agents, and one of IBM’s 2026 predictions was the importance of smarter data.

The usual approach is adding metadata and validation later in the data pipeline. That’s similar to the way the bronze, silver, and gold tiers of typical lakehouse architecture are supposed to represent how filtering, cleaning, and augmenting data improves structure and quality until it’s ready to use. That can mean an enormous amount of work since nearly three quarters of the CPU work in training a frontier model is data cleansing and validation.

But that can also remove a lot of the context crucial for gen AI. Rather than cleaning data and losing the original context, it’s often more effective to keep as much information about the original state of the data, says David Aronchick, open-source platform Kubeflow founder, and CEO of distributed data pipeline vendor Expanso. “You can’t pursue exactly purely clean data; that’s just not possible,” he says. “As you pull data into your ML model, every line should have some mechanism saying where it came from. Otherwise, you’re never really going to know because you can’t mix them together and tease them apart later. You can search your raw content, your raw logs, but it’s just not going to be there.”

IoT digital twin systems often tag data all the way back to the device capturing it so you can see whether a temperature spike is a critical failure, which you want to react to, or a routine calibration, which you don’t. But that information may well be relevant down the line when you want to use that data more broadly. So unless you capture at least some elements about the source of data before you move it, you’re not going to be able to easily reconstruct the context later, or at all sometimes.

Ulrik Hansen, co-CEO of Encord, a platform for managing and annotating data, calls this in-stream labelling and cautions it’s not an alternative to cleansing data. “Dirty conflates two things: actual corruption you should fix, and context dependence, where a reading only looks anomalous because you threw away the frame that explained it,” he says. “Cleansing kills both. The point isn’t to stop cleaning, it’s to stop normalizing away context you can never recover.”

Context can be cheap to capture at the source and nearly impossible to recover after, he adds. “The question isn’t whether to keep it,” he says, “it’s about curating what actually helps.”

Raw but not rancid

Aronchick characterizes the state of most bronze tiers as toxic waste because raw data doesn’t get validated before ingestion, or have a metadata wrapper on each data point. “You’ve taken raw data and stripped it of context,” he says.

Take a wind farm operator, for instance. When sensor data about the turbines is generated, it comes from a particular turbine at a particular position in a specific wind farm at a known location, running at a specific speed in specific weather conditions, at a particular time. “If you have other turbines also working in the field, the performance of your turbine will go down, but the field performance will go up,” says Aronchick. “The performance of your turbine going down isn’t a negative, but unless you have the context at the point of data collection, you’re going to make your life much harder later on, when someone asks about the efficiency.”

Metadata needs to be much richer, and it needs to be added as early in your data pipeline as possible when you have the most detail available to make sense of the structure and complexity of the data, Aronchick adds. “You want to capture as much about the data you’re collecting as possible, where it doesn’t require insane activity to do so.”

But not all the metadata you need will be generated with the data, he says. You almost certainly need to augment and annotate your data, and provide extra structure, especially for something like a point of sale system with very light metadata. “Data comes off these things in poor structure,” he says. “It’s not OpenLineage, it’s often a CSV or a text record, and you have to reconstruct them into a full structured log. So do smart things where you’re creating data. That might be compressing, sampling, converting, appending metadata to it, and enforcing schema and lineage all before you start moving anything.”

That doesn’t have to mean bloating your data, Hansen points out. He suggests capturing what’s free and unrecoverable. “The system of origin is the label,” he says. “You don’t tag HR policy, you capture that it came from the HR system. Anything a model can derive later, you can skip.”

Structure isn’t static

Routine changes to APIs, schemas, and how data is collected or stored happen in every organization, and need to be reflected in metadata that lives alongside the data or added as data is collected, not reconstructed later in a fragile process that depends on knowing about all those changes. Google’s research into these data cascades shows how easily context gets lost and how badly it affects data quality.

Shifting schema enforcement further left in your data pipeline so you deal with it as soon as possible allows you to make more effective downstream decisions. For a sensor recording temperature and humidity, you need to know the temperature scale it uses, readings, and how the timestamp is recorded. Checking that against the schema before ingesting the data lets you route it differently depending on whether it validates or triggers alerts about data quality.

“Maybe I’ll delete it, or send it off to some place where a human being or other tooling can reconstruct it into something valuable,” says Aronchick. “But what it doesn’t do is allow the polluted or bad data into my pipeline. Saying whether or not something passed your schema makes your downstream systems much more reliable.”

Sensing structure

Unstructured and semistructured data needs more augmentation. A PDF or Word document has an author and a creation date, but doesn’t necessarily include any context about the job title and department of the author, whether it’s up to date, only applies to a particular group of customers, or is based on accounting regulations that can change. If that information is available, it needs to travel with the document, not be left in a compliance spreadsheet.

Data platforms like DataHub and SurrealDB both capture and create context. The latter can analyze a photo, for instance, using vision AI to understand what’s in the image. “From completely unstructured data, we get as much structure as possible,” says the company’s CEO Tobie Morgan Hitchcock.

That’s paired with other data potentially useful for an AI agent down the line. “Understanding what happened around an event becomes a lot easier if you’re tracking the conversation, telemetry, tool and model usage, geospatial data, and the vector search and relationships,” he says. “You’re going to have a far better chance of getting an accurate understanding of that data, which started off completely unstructured, than if you weren’t capturing anything.”

Metadata about document authors, which might come from the company directory, can show how much authority a document has. He describes that as building an understanding of what trust and provenance is over time by the weight and authority of who’s updating the information. After all, he says, company-generated information has more trust or can have traced provenance compared to conversational inputs from a user.

Incentives for annotating

DataHub CTO Shirshanka Das saw how much of a mess data can be even with strong guidelines as former architect of LinkedIn’s GDPR strategy. “The data was a swamp, despite us having had pretty good data-first and schema-first practices,” he says. As well as cleaning up the data governance, they added in the first nuggets of the DevOps’ ‘shift left’ approach.

LinkedIn already required data checked in to its Kafka ecosystem to have a schema, and ran CI/CD pipelines to check backward compatibility. “I attached metadata attribution and collection around compliance metadata into that pipeline, where developers weren’t able to check in a schema until they had declared what every column meant.”

The extra work was unpopular until teams who didn’t participate saw the flood of tickets that came their way, which allowed him to extend that same proactive governance and annotation at source approach to pretty much every data set being produced.

“The starting point of data at most companies is a lot more swampy,” he says. “Many people are using Kafka, which is a very schema forward system, and yet they’re just shoving in JSON and unstructured stuff.”

That’s common, agrees Megha Kumar, research VP for analytics and AI at IDC, because while collecting more metadata provides better context and cleaner data lineage, it’s hard in practice. “Most organizations batch process data, so real-time context capture rarely happens,” she says. “Even the ones that process in real-time tend to have pre-defined schemas, so adding context requires changes to the data, which unfortunately happens later.”

People don’t know how to start, says Das, so DataHub Cloud tries to add back context by collecting operational metadata from multiple systems, including queries and BI tools to extrapolate a semantic model. “We confront the mess by giving them something they can react to,” he says. “They can quickly validate, and then it starts becoming a governance layer on top where humans annotate at source.”

Online whiteboard provider Miro, for example, dramatically improved AI agent query accuracy from about 50% to 90% using DataHub. Then they applied GitOps principles on top of what was inferred with a human in the loop for approvals.

So getting people to do the work happened the same way at LinkedIn, says Das. “When a data scientist gets 10 times more requests because they didn’t document their work well, resulting in the AI making lots of mistakes and stakeholders constantly pinging them for answers, they have the incentive to add the annotation when they produce an analysis, because then they get out of the critical path.”

DBOMs and data contracts

Provenance and lineage of data is critical, Aronchick says, so you can preserve details like who collected the data, when, from where, if the source was authoritative or canonical, what transformations were run, and exactly what the model saw.

“It’s not just about the version and the metadata,” he says. “Where things really start to change is when you can say along the way this data has gone through these steps, this is the root source, and these were the other elements.” You want to be able to find out if there were any experimental flags, like a new customer campaign running when it was collected, as well as what claims the data contributes to.

Aronchick advocates for a SLSA-style data bill of materials using a tool like Makoto, which can add signed provenance and attestation to simplify applying central concepts of governance and structure to upstream data.

The notion of a data contract or a data product spec is starting to become common in the financial sector says Das, defining it as a data set, or a group of data sets, bound together by a contract that defines expectations which aren’t just cosmetic but machine verifiable. They can also include operational SLOs for APIs as contracts describe not just the shape of the data but operational characteristics and guarantees.

Document graph markup language (DGML), a new open source specification from Docugami, promises provenance down to individual data points automatically extracted from documents.

“It’s critical to know the validity and provenance of the information your AI is relying on,” Docugami CEO and XML co-creator Jean Paoli says. “Establishing the validity of data right from the start, at scale, is vital and far more efficient than trying to clean up bad data later.” DGML combines semantic tags describing what content means in its business context with bounding boxes showing exactly where in the document the content comes from, with attestation to prove it.

AI demands provenance

All this context is the kind of metadata Anthropic’s context engineering guide recommends feeding to agents for accuracy. Developers are already used to giving coding agents more context, Das argues. “The same thing is happening with data, as when people realize when AI agents can’t make sense of what they’re doing, hallucinations happen,” he says.

Kumar agrees that organizations realize agents need context to provide better insights. “In many cases, it has to do with ensuring the existing data had clear semantics and relationships,” she says.

If you want to make sure the purchase return window an AI chatbot promises customers is based on your own policy, not a wish list from a user forum, you need rich context. It’s not just metadata. Organizations need to have semantics, data lineage, and ontologies. “Many are also building knowledge and ontology graphs,” adds Kumar. “By ensuring the systems understand what the data means, it’ll be able to provide a better response.”

And if you’re going to the expense of fine tuning, which needs relevant and domain- or task-specific examples, you don’t want noise, duplication, or irrelevant content in your data. You can, of course, exclude poor data if it’s annotated and verified earlier, but you can also improve model performance with extra information, Aronchick points out. “The augmentation of the existing data makes the data you pull out more valuable,” he says.

Expanso recently won an Edge AI award for fine tuning a base level model with only about 3,200 images by augmenting them with metadata. “The reason it worked on that few is because I could tell it deterministically what was in the frame,” he adds. “It’s labeling at the point of capture instead of paying somebody to label it later. What if I developed models for predictive analytics of store behavior on a per city, region, or country basis? If I’m able to take the raw point of sale information and augment it with additional metadata, I’m turning this into a much easier thing to fine tune.”

Or you might even avoid the expense of fine tuning entirely, suggests Das. “You get the short-term advantage by fine-tuning and getting great performance at much cheaper cost on a smaller model, and it gets stripped away in a couple of months as a new model shows up,” he says. “You have to always run that calculus of when’s the right threshold to fine tune an existing model, distil it, and then run it for a fair amount of time to recoup the costs of fine tuning.”

Although regulated or slow-moving industries will see benefits from fine tuning a model they can run for six to 12 months on data with higher quality and better provenance, many organizations may use the improved data quality to get good results without fine tuning.

“We’re taking a more knowledge graph-oriented approach to grounding the model, and betting on the fact that because the knowledge graph is changing often, it’s better to keep it as a runtime artifact than a baked-in one.”

Your AI model isn’t the problem. Your data was never ready for it

The meeting that changed my perspective

I remember sitting there and realizing I wasn’t thinking about the model at all. I was thinking about the data feeding it.

One discussion stands out in particular. We were evaluating how predictive analytics could improve sales forecasting for a national portfolio of opportunities. Leadership wanted greater confidence in projected outcomes so resources could be prioritized earlier in the sales cycle. As conversations turned toward model accuracy, we discovered something more important. Different teams weren’t consistently recording opportunity stages, probability scores and client attributes. The model wasn’t struggling because it lacked sophistication. It was learning from business processes that had never been standardized in the first place. That meeting changed how I approached every AI initiative that followed.

Throughout my career leading enterprise business intelligence initiatives, I’ve repeatedly watched organizations blame the algorithm when the real issue was inconsistent data, fragmented ownership across departments and business definitions that meant different things to different teams. AI doesn’t distinguish between disciplined and inconsistent business processes. It learns from both with equal confidence.

I’d built and defended executive dashboards for years before that meeting, and dashboards had trained me to believe imperfect data was a manageable, even routine problem. Experienced leaders read a dashboard with context. They know which numbers to trust, which ones need a caveat and which gaps to mentally fill in based on what they already know about the business. Predictive AI doesn’t have that judgment. Machine learning assumes the historical data it’s trained on represents reality as it actually is. If two departments define “active customer” differently, or if a critical field has been silently incomplete for two fiscal years, the model doesn’t notice or compensate. It learns the inconsistency as ground truth, and it repeats that mistake at scale, with confidence, every single time it runs.

That moment fundamentally changed how I approach every AI initiative. I stopped starting with technology and started with data integrity instead.

5 questions I ask before any AI platform conversation

Today, I rarely begin AI discussions by talking about technology. Before any conversation about platforms or vendors, I ask five questions of the leadership team. Can we explain, in plain language, where this data actually comes from? Do the business leaders in the room agree on what our core definitions mean, or does “revenue” or “active account” shift depending on who’s presenting? Would we rely on this data to make a multimillion-dollar decision without a human manually double-checking it first? Is there a specific, named person accountable for every critical dataset, or does ownership dissolve the moment something goes wrong? And underneath all of it, are we actually solving a business problem, or are we chasing a technology because it’s the thing everyone else is talking about this quarter?

I remember one initiative where these questions prevented us from moving too quickly. During an early assessment, we discovered that two operational systems treated the same customer differently because each had evolved around separate business processes. Executive reports appeared consistent because manual reconciliation had become part of the monthly reporting routine. Once we identified the inconsistency, the project paused while business stakeholders agreed on common definitions and ownership. That decision delayed the AI initiative by only a few weeks, but it likely prevented months of troubleshooting after deployment. More importantly, it strengthened confidence in every analytics initiative that followed.

These conversations reveal far more about whether an organization is genuinely ready for AI than any vendor demonstration ever will. A polished proof-of-concept can make almost any dataset look production-ready for the ten minutes it’s on screen. These five questions don’t have that luxury. They tend to surface, quickly and uncomfortably, where an organization’s data confidence actually breaks down, and that’s the information leadership needs before committing budget and reputation to a rollout.

This lines up with what the NIST AI Risk Management Framework has argued for a while now: governance and accountability belong at the foundation of an AI initiative, not layered in after a model is already in production. Governance built in retroactively tends to be theater, built to explain a failure that’s already happened rather than to prevent one.

Leadership before technology

The organizations I’ve seen actually succeed with AI invest first in governance, ownership and shared business definitions, and only then in the platform itself. They clean up master data before they scale a model against it. They remove duplication in customer and product records. They assign accountability for datasets the same way they’d assign accountability for a budget line, with a name attached and consequences if it slips. This work rarely shows up in a demo, which is probably why it gets skipped so often in the rush toward deployment.

One lesson I’ve seen repeatedly is that assigning ownership changes behavior almost immediately. Once business leaders understood they were accountable for the quality of specific datasets, not just the reports generated from them, conversations shifted. Instead of asking why dashboards looked different, teams began discussing why the underlying business process produced inconsistent information. Governance stopped being viewed as documentation and became part of everyday decision-making. The improvements weren’t dramatic overnight, but they were sustainable, and that consistency ultimately mattered more than any individual technology upgrade.

I saw this firsthand during an executive reporting initiative where multiple leadership teams relied on the same performance dashboard but interpreted one KPI differently, because ownership had never been clearly assigned. Once the business designated a single owner for the metric and standardized its definition across reporting systems, disagreements disappeared almost overnight. More importantly, that same governance work later allowed predictive analytics to be introduced with confidence, because everyone was working from the same version of the truth.

I’ve learned that AI projects rarely fail in the data science team. They fail months earlier, when leadership assumes the organization already understands its own data.

McKinsey’s research on scaling AI reinforces this pattern at scale: the organizations that generate lasting value from AI are consistently the ones that pair the technology with real changes to their operating model and governance, rather than simply layering AI on top of how things already worked. That finding matches what I’ve observed leading enterprise analytics initiatives directly. The technology was rarely the constraint. The organization’s relationship with its own data was.

It’s tempting to frame AI adoption as an engineering problem with a leadership footnote, when in practice it’s closer to the reverse. CIOs reporting on rebuilding an AI-ready data strategy make a related point: treating data ownership as a purely IT issue stops working once business units, product teams and AI platforms are all generating and transforming data continuously, which is exactly why accountability has to sit with named business leaders, not a technical team working in isolation. A related piece on building an AI-ready data culture puts it more bluntly: an organization can’t scale AI without first scaling trust in its own data, and that trust starts with culture and ownership, not tooling.

I no longer ask whether an organization is AI-ready. I ask whether its leaders would bet on their own data without a human checking behind the model first. If the honest answer is no, the next investment shouldn’t be another AI platform or a more sophisticated model. It should be a stronger data foundation, built deliberately, with clear ownership, before a single additional AI use case gets greenlit.

If another executive asked me for one piece of advice before approving a major AI investment, I’d tell them this: spend one day interrogating your data before spending another dollar on your model. What that conversation reveals will tell you more about your organization’s readiness than any vendor demonstration ever could.

Organizations rarely fail because their AI isn’t intelligent enough. They struggle because they ask AI to learn from data that was never prepared to support intelligent decisions in the first place. The organizations that lead in this next era won’t be the ones with the most advanced models. They’ll be the ones that got their own house in order first, and knew it.

This article is published as part of the Foundry Expert Contributor Network.
Want to join?

Why your data layer is AI’s most critical climate technology

I have spent enough time building cloud and AI infrastructure to know one thing: Efficiency problems never show up where teams expect them. They tend to sit just beneath the surface, quietly shaping outcomes long before they appear in the metrics anyone is tracking. Most enterprise conversations still center on token pricing. That focus makes sense. The per-unit cost of using models has fallen quickly, and organizations want to understand whether AI can scale without pushing budgets out of bounds.

That question has become more complicated than it looks. At the same time, the nature of each interaction has changed. What used to resemble a straightforward exchange now involves retrieval and ongoing reasoning that unfolds across multiple steps. Even with lower unit costs, total consumption continues to rise because each request requires more underlying work. This creates a contradiction. Tokens are cheaper, yet overall compute demand keeps increasing. That compute demand has to go somewhere, and where it goes is a growing problem. 

The answer, increasingly, is the power grid. The International Energy Agency has warned that AI and data centers are becoming a major source of electricity demand, and Goldman Sachs Research has projected that data center power demand could rise 165% by 2030 compared with 2023 levels. Today, electricity consumption from data centers already amounts to roughly 415 terawatt hours, or about 1.5% of global electricity use, and has been growing at roughly 12% annually over the past five years. Much of the response has focused on models or hardware. Both matter, but they do not explain the full picture. In enterprise environments, a meaningful share of inefficiency begins before a model processes anything at all. It begins in the data layer.

The paradox: Cheaper tokens, more total compute

I have seen this pattern play out before. When unit costs drop, usage expands and everyone is surprised, every time. AI is following that same trajectory.

Lower token pricing has made it easier to integrate models into more workflows. Those workflows have also become more involved. Systems retrieve broader context and revisit outputs before returning a result. The additional activity outweighs the savings from lower pricing. In many cases, a system prompt alone can consume 2,000–5,000 tokens before a user even contributes input, which means the baseline cost of each interaction has already expanded. This is no longer just a financial consideration. Increased compute demand brings increased energy consumption, which is becoming a central constraint for enterprise AI.

Why agentic AI changes the energy equation

Average load: It rarely tells the full story. A platform can appear stable while quietly absorbing constant internal work that never surfaces to users. Agentic AI introduces that kind of persistent demand.

An agent continues operating in the background even after a user interaction ends. It checks for new inputs, revisits its internal state, then prepares the next action. That ongoing loop changes how infrastructure is consumed. Instead of handling activity in bursts, systems begin to carry a steady baseline load. Compute usage becomes continuous rather than intermittent.

This shift exposes inefficiencies that might otherwise go unnoticed. A slow query or a fragmented data source affects far more than a single request. The same issue repeats continuously, which increases both cost and energy consumption over time. This matters even more as inference workloads scale, with projections suggesting they could account for more than 40% of data center demand by 2030.

The overlooked source of energy growth: The data layer

In almost every architecture review I’ve sat through, attention gravitates toward model performance or infrastructure spend. The data layer gets treated as an afterthought. That assumption falls apart the moment you move into retrieval-heavy AI. 

Agents depend on access to reliable context. They pull information from operational systems as well as analytical platforms. When those sources are disconnected, each interaction requires additional effort to assemble a usable view.

That effort accumulates quickly. Data must first be located before it can be used. It often needs to be moved into another environment, then reshaped into a format the model can process. Only after that does it become useful for decision-making. Each step introduces overhead that consumes compute and energy. In practice, engineers can spend up to 40% of their time preparing data before it is even usable for downstream systems. The model remains the most visible part of the system, but the surrounding data work often determines how much energy the system ultimately uses.

Fragmentation as the root cause

Enterprise architectures evolve over time, and fragmentation usually reflects a series of reasonable decisions rather than a single mistake. The challenge appears when AI systems need to operate across all environments at once.

Fragmentation forces repeated work. Systems retrieve overlapping datasets because no single source is trusted as authoritative. Pipelines reprocess information that already exists elsewhere. Teams build parallel structures instead of relying on shared ones.

This pattern increases demand in ways that are easy to overlook. When retrieval becomes inconsistent, applications compensate by sending more context than necessary. Models then process larger inputs, which increases token usage without improving the quality of the outcome. What appears to be a model efficiency issue often traces back to data architecture. It also helps explain why as many as 60% of AI projects are abandoned before reaching production, often due to gaps in data readiness rather than model capability.

Why sovereign, unified architecture changes the math

I have seen organizations improve efficiency without changing models simply by reducing friction in how data is accessed. The good news is that the fix is closer than most teams assume. 

A more unified architecture shortens the path between a question and the data needed to answer it. When systems can access authoritative data directly, they avoid repeated transformations and unnecessary duplication. Retrieval becomes more precise, which allows models to operate with less excess input.

This does not require consolidating everything into a single system. Enterprises will continue to operate across multiple environments. The objective is to reduce unnecessary movement and make data easier to use wherever it resides. Sovereignty and control play an important role as well. Organizations operating across different environments need a clear understanding of where data resides and how it can be used. When governance is built into the architecture, systems spend less time reconciling access and more time producing results. Reducing friction at this layer has a direct effect on both cost and energy use.

What 2026 will expose

The next phase of enterprise AI will shift attention from access cost to operating cost. Early efforts focused on whether organizations could use advanced models. The next stage focuses on whether those systems can run continuously across real workflows without creating unsustainable demand.

That shift will make energy consumption more visible. It will also make inefficiencies harder to ignore. Some increase in demand reflects real value. Systems that improve decision-making or streamline operations will naturally require more compute. The more difficult question is whether additional consumption reflects useful work or avoidable overhead.

When systems repeatedly move and reprocess the same data, the increase in energy use does not correspond to better outcomes. It reflects architectural inefficiency. CIOs will need to examine that distinction more closely.

The takeaway for tech leaders

Before adding infrastructure, ask an honest question: Is this capacity supporting real growth, or is it just hiding inefficiencies that should have been fixed first? Start with the data layer.

Review how systems retrieve context. Look for duplication across environments. Identify where data is repeatedly transformed before it becomes usable. These patterns reveal whether the architecture is enabling efficient AI or creating unnecessary demand.

Enterprise systems will always involve complexity. The objective is to ensure that complexity does not translate into avoidable work. In the agentic era, the efficiency of AI systems is closely tied to how much work happens before the model produces an answer. Data architecture plays a central role in determining whether that work remains controlled or expands beyond what is necessary.

This article is published as part of the Foundry Expert Contributor Network.
Want to join?

❌