Visualização de leitura

65% of employees would love to roll back workplace AI

IT leaders have been making generative AI tools available across the enterprise for just three years, and a significant majority of their business users has already had enough.

According to a report from Adaptavist, 65% of 2,500 knowledge workers surveyed say they “regularly feel nostalgic about how work operated before the widespread adoption of AI.”

This “pre-AI nostalgia” appears to be due in part to business users feeling overwhelmed by the responsibility of learning how to use AI on top of their day-to-day job tasks. Moreover, 46% of workers say their concerns about AI have gone unaddressed by management.

“Transparency is critical to truly drive AI engagement; organizations must establish clear guardrails and maintain an open dialogue around AI use and employee choice where workers feel they are being listened to,” Jobin Kuruvilla, field CTO at Adaptavist, tells CIO.

Generational gaps in AI acceptance

Despite an assumption that younger workers are more intuitively adept with AI tools, Gen Z workers (42%) are more likely to prefer the pre-AI world compared to their Gen X colleagues (26%). This may support the growing concern that AI is quickly is hitting entry-level workers the hardest, while creating new career opportunities for more skilled workers who have been in the industry longer.

When asked about fears surrounding job obsolescence due to AI, 54% of all workers surveyed said they are “concerned AI could reduce the need for their role within the next five years.” Broken out by organizational level, junior employees (23%) and C-level executives (29%) expressed the most concern about AI job loss, compared to 13% for mid-level employees and 12% for senior employees.

Additionally, 47% of C-level executives and 36% of directors are looking to move industries, change careers, or step away entirely due to concerns of AI eliminating their positions. Still, plenty of workers are ready to face the new challenges of an AI-driven workplace, with 74% saying they are actively learning new skills to stay relevant, and 85% of C-level leaders saying the same.

Lack of transparency drives AI fatigue

One in three workers (36%) are already experiencing “AI fatigue,” leading to less frequent use of AI tools and active resistance to AI for day-to-day tasks. More than a third of workers (36%) also appears to be confused about AI use expectations in their role.

When implemented quickly without proper training and transparency, AI initiatives can lead to hidden productivity costs. Of those surveyed, 42% say they “spend more time verifying AI output than they save using it,” while 52% say they regularly spend time correcting AI-generated work from colleagues. Additionally, 49% say low-quality AI outputs slow down projects, 55% say AI-generated content reduces overall team efficiency, and 46% say it makes their work feel “more repetitive and less meaningful.”

Half of all workers also feel their performance is now “directly or indirectly compared to AI-generated output.” Providing clarity about how AI impacts or doesn’t impact an employee’s career is important to staving off AI fatigue.

For those chalking this all up to change resistance, know this: 67% of workers surveyed say they want their organization to increase the use of AI, and 69% say they believe AI is being used ethically within the organization. What they lack is a roadmap, guidance, and training to understand how to best implement AI at work, and to ensure it’s being used effectively.

“Ultimately, by automating the mundane tasks that make work feel repetitive —organizations can refocus their specialists on high-value creativity, transforming AI from a source of fatigue into a powerful engine for meaningful human achievement,” says Anand Unadkat, a senior solutions architect at Atlassian.

IT leaders and their executive colleagues need to focus more on the change management artistry necessary to help get them there.

If we want to implement AI successfully, we need to completely change how we do businesses

I’ve always thought it was interesting that we’re willing to fight and die to live in a democracy, but everyone is happy to work in a company which is structured like a dictatorship. This thought feels even more pertinent given the rise of AI. As AI continues to transform the world of business, we’re starting to notice a clear gap between those implementing a ‘throw it at the wall and see if it sticks’ approach, and those examining the fundamental changes that need to be made to a business.

While we don’t need to get into the pros and cons of oligarchy, over the past year of leading consultation and training sessions for over 80 organizations, I’ve realized that, if you introduce AI by working from the middle out, you can actually move a lot faster.

I’ve seen organizations try to bolt AI onto their existing workflows, and while there may be initial productivity gains, this generally doesn’t work out in the long term. We often see scattered pilots which don’t go the distance, duplication of tools or inefficient processes. The organizations setting themselves up for success are redesigning how teams experiment with and implement solutions.

We can use the transition from steam to electricity as an example. Paul A. David notes that there was a 40-year lag between the electric dynamo’s introduction and its productivity impact. When factories first adopted electricity, many simply replaced their steam engines with electric motors, while leaving the rest of the factory unchanged. Productivity gains were modest. David argues that the bottleneck was organizational structure. It was only when engineers redesigned factories around small electric motors throughout the factory that we began to see the benefits. General purpose technologies, like electricity — or, in this case, AI — require co-invention and firm restructuring before we can see the benefits.

It’s time to restructure.

AI is developing fast and it’s difficult for companies to keep pace

While most companies are built with a top-down model, this is not an effective way to identify and roll out technology, particularly when it’s moving as quickly as AI is.

We’re already seeing the impact that the speed of AI development is having. Companies are struggling with things like AI sprawl and shadow AI. AI sprawl is when employees are using tools everywhere, without shared norms or strategy. Gartner estimates that by 2028, an average global Fortune 500 enterprise will have over 150,000 agents in use, up from less than 15 in 2025. This creates significant agent sprawl, IT complexity and management challenges. Take a retail company for example, if sales uses one AI chatbot, support uses another and marketing uses a third we could start to see inconsistent customer experiences, where customer-facing AI chatbots give conflicting answers about pricing or return policies.

Shadow AI is the unauthorized use of AI tools by employees without IT or security approval. Common examples include employees pasting code into ChatGPT, uploading customer data to public web apps or using unvetted browser extensions to speed up daily work. Today, over one-third (38%) of employees say they share sensitive work information with AI tools without their employers’ permission. This introduces severe risks like intellectual property leaks, data privacy violations and non-compliance.

Building the assembly line of the AI era

The companies solving these challenges are finding ways to convene subject matter and AI experts from every part of the company to create a center of excellence, steering committee or a power user group. Once assembled, this group should be empowered to experiment, vet and validate new technology for the company. This is what I’m calling the new ‘assembly line’ of the AI era.

This assembly line is a dedicated team with the power to implement new solutions. They can vet any tools being used and compare them to systems already in place. They can then make decisions on whether the identified tools should be rolled out across the company, and the training and processes that need to be in place to make this rollout a success.

Some of the companies I’m working with are already putting this into practice. One water pumping company in Minnesota wanted to figure out how AI could be used to educate, train and inform employees, as well as preventing AI sprawl or shadow AI usage. Together we have mapped out who should be part of their center of excellence, who owns what and how to set up approvals. With the model we’re creating, we are establishing AI as a force for empowerment, education, training, tooling and most importantly change management.

On the flipside, I’m also working with a healthcare company that has an existing center of excellence trying to oversee ALL AI projects at the business unit level. This recreates a hierarchical problem that slows adoption, since it doesn’t empower individual units to move on their own. It makes sense, given HIPAA compliance means healthcare companies have to be cautious, but the focus of a Centre of excellence should be enabling teams through approved AI tools, not taking complete ownership of every project themselves.

By giving these teams authority to make AI specific decisions, you prevent the bottleneck which usually happens at the executive level either due to busy schedules or a less in-depth technical understanding. The center of excellence can redirect sprawl, combat shadow AI usage and escalate things when necessary.

Turning individual experiments into company decisions

A center of excellence gives AI adoption a working rhythm instead of leaving it to Slack threads, scattered pilots or executive guesswork. Each team should have someone close enough to the work to spot where AI is useful and where it’s a distraction. A finance lead might see value in automating invoice checks. A legal lead might reject a tool because it mishandles client data. A customer service manager might test whether an AI assistant actually improves response quality or just produces faster, worse answers.

That group can then turn individual experiments into company decisions. They can test tools, compare them against existing systems, check the security risks and decide what needs training before anything is rolled out. They can also stop bad habits early, like teams uploading sensitive documents into public tools because nobody gave them a safer option.

If companies want AI to work, they need to completely overhaul their processes. The companies that move fastest will be the ones that give people in the middle the authority to test, challenge, approve and teach. That’s where the real work happens: close enough to daily operations to know what’s useful, and connected enough to turn that knowledge into practice across the business.

What changes when AI becomes part of how the business runs?

What changes when AI becomes part of how the business runs

The more I speak with CIOs and technology leaders, the more I realize most of us are working through variations of the same AI challenge.

How quickly should we move? Which opportunities are worth pursuing? What risks are acceptable? And how do we move from an impressive demonstration to something the business can reliably use?

Enterprise AI began with possibility and experimentation. Now the conversation is changing.

The harder question is not whether AI can perform a task. It is what changes once the business begins depending on it. At that point, the conversation expands beyond technical capability. Value, capacity, security, ownership, change management and operational resilience all become part of the equation.

The demo is not the operating environment

A strong AI demonstration can be compelling. The data is clean, the use case defined and the operator knows the technology. The result can look effortless.

Real environments rarely behave that way.

I have seen intelligent automation use cases appear straightforward until actual business data and processes were introduced. Documents varied, requirements evolved and manual workflows contained accumulated exceptions. What looked like one process turned out to be several versions held together by human judgment.

The technology may be capable, but it does not resolve unclear requirements, inconsistent inputs or a process that was never standardized.

I prefer to test with real enterprise data as early as practical. Vendor demonstrations naturally emphasize the happy path. Your own data exposes the conditions the solution will actually have to survive.

Watching an expert operate a platform is different from asking employees to use it every day. Users have to understand the capability, trust the result and know what to do when the output is wrong or incomplete.

Change management cannot be treated as the last step. It affects the timeline, effort and whether the expected value shows up.

McKinsey’s State of AI research continues to show broad adoption while enterprise-wide scaling remains much less common. That gap is understandable. The distance between an interesting use case and a production capability is where data, process design, testing, security, integration and adoption all become real.

Value has to compete with capacity

Once a use case survives the technical question, the discussion has to become more pragmatic.

What is the value?

Within an enterprise, that should translate into something leadership can evaluate: cost reduction, increased throughput, greater efficiency, less manual work, more time redirected toward higher-value activities, better customer outcomes, revenue opportunity or the ability to absorb growth without adding proportional headcount.

Not every AI initiative needs an immediate hard-dollar return. But leadership should know the intended outcome and how it will determine whether further investment is justified.

AI does not create unlimited organizational capacity. Technology teams still have roadmaps and operational priorities to deliver. Business subject matter experts still have day jobs. Someone has to define requirements, provide data, validate the process, test the outcome and help employees adopt a different way of working. And when the organization chooses to build rather than buy, additional work may be required to prepare data, evaluate model performance and, where appropriate, fine-tune models for the specific use case.

That is why being able to build a use case does not automatically make it the right priority. The value, effort, timing and business readiness still have to justify the investment.

Sometimes the smaller opportunity is better because it produces value sooner and builds reusable experience.

The same discipline should apply to whether the organization builds internally or brings in external expertise.

AI is evolving too quickly for most internal teams to master every emerging capability while operating the rest of the enterprise. A proven external partner can sometimes add expertise, speed or capacity.

The test is whether that partner accelerates internal capability or creates an unsustainable dependency.

Board expectations are also increasing, and rightfully so.

AI now touches competitive positioning, investment priorities, workforce decisions and enterprise risk. Boards should ask where value is emerging and whether the organization is moving with enough urgency.

BCG research on CEO and board perspectives has highlighted a useful tension: in some organizations, boards are pushing for greater urgency around AI, while management teams may take a more measured view of what can realistically be delivered and sustained.

The better question is not simply how fast the organization is moving. It is how fast it can move while still producing something it can support, protect and sustain.

That is where risk stops being only an IT discussion.

Before an AI capability moves deeper into the environment, CIOs need to understand what it touches. What data can it access? Does information leave the enterprise? What permissions does it require? Could it introduce a new attack path? What happens when the capability begins taking actions across systems instead of simply producing an answer?

The control model should reflect the consequence.

An AI tool used for everyday productivity does not require the same oversight as one that can modify records, interact with customers or access sensitive enterprise data. The NIST AI Risk Management Framework provides a useful structure for thinking about risk in context rather than applying the same controls everywhere.

For CIOs, that is the balance: we are still responsible for protecting the enterprise, but protection cannot become an excuse to make every new capability unnecessarily difficult to adopt.

Guardrails should be strong enough to protect the business and flexible enough to evolve with the technology.

Sometimes that requires more common sense than textbook governance.

The stakes change when AI moves beyond the office

For many organizations, AI adoption starts with office productivity: summarization, knowledge search, coding assistance, meeting support and other relatively contained uses.

Eventually, the question changes.

When can AI move deeper into operations?

That can include intelligent document processing, computer vision, IoT and sensor-driven capabilities, drones or other technologies that begin influencing operational decisions and physical processes.

Some companies will move there gradually. Others may move sooner when the capability sits inside an established vendor-managed solution with defined controls, support and accountability.

If an AI tool used for everyday productivity produces a poor response, an employee can usually identify and correct it. If an AI-enabled capability begins influencing an operational process, reliability, cybersecurity, fallback procedures and ownership become much more important.

That progression from experimentation to deeper enterprise dependence is not new. We have seen it in other technology cycles.

Cloud, SaaS and mobile all moved through periods of enthusiasm, rapid adoption and eventual normalization.

AI will likely follow parts of the same pattern, but the cycle is moving faster.

It did not enter primarily through the traditional IT corridor. Employees, business teams, vendors and executives gained access almost simultaneously. The technology continues advancing while organizations are still deciding how it should be used and controlled.

Much like smartphones and the internet became embedded into daily life, AI is already becoming part of the applications people use every day. Capabilities are being built into enterprise platforms, whether users think of them as AI or not. The next shift is deeper dependence as AI becomes part of workflows, decisions and operating processes.

The difference is that AI can operate at a higher altitude. It can influence decisions, interact with enterprise data and increasingly take actions across systems, which raises the consequence when something goes wrong.

That should change the questions boards and CEOs ask. The conversation should move beyond “What are we doing with AI?” to questions that expose whether the enterprise is actually ready to depend on it:

  • How do we move faster without putting the business at unnecessary risk?
  • What are we asking AI to compensate for that we should be fixing ourselves?
  • Where are process ambiguity, system fragmentation or operating habits creating unnecessary friction?

AI can automate around a weak process for a while, but eventually the exceptions catch up with it. It can work around inconsistent information only so long before confidence in the output becomes the problem.

CIOs will need to hold firm on responsibilities that do not change while staying flexible in how those responsibilities are carried out.

We also have to be realistic about what our organizations can absorb. Trying to boil the ocean can create more activity than value. There is nothing wrong with narrowing the focus, proving an outcome and using specialized expertise when internal capacity or experience is not there yet.

The first phase of AI rewarded experimentation and curiosity.

The next will reward judgment.

The organizations that navigate it well will not necessarily be the ones with the most pilots, the largest budgets or the boldest promises. They will be the ones that know where to move quickly, where to hold the line, what needs to be fixed internally and when an idea has earned the right to scale.

That is when AI stops being another technology experiment and starts becoming part of how the enterprise actually runs.

Why IT projects still fail

Lately most execs have been focused on making sure their AI projects pay off.

With good reason: The rate of failure for AI initiatives has been notoriously high.

But AI projects aren’t the only ones that need attention. In fact, CIOs and their executive colleagues should be putting that kind of focus into all IT projects, given that success on more conventional initiatives — from new software deployments to ERP implementations — is far from perfect.

Statistics vary. Some often-quoted reports about IT project failure rates of 70% date back several years, making them unreliable reflections of the landscape today. But project consultants say a good percentage of IT projects still fail, with estimates ranging from about a third to as much as high as that 70% mark.

In the Project Management Institute’s 2026 Pulse of the Profession report, researchers report that 31% of complex projects fail to achieve the full scope of their originally intended benefits.

CIOs, project leaders, researchers, and IT consultants generally define failure for an IT project as not delivering expected benefits within the expected timeframe. Failure can also mean a project doesn’t produce returns, runs so late as to be obsolete when completed, or doesn’t engage users who then shun it in response.

Why do IT projects continue to fail? Here are 12 common culprits.

1. Lack of project management expertise

Expensive and highly visible projects get the benefit of being led by professional project managers, but small and midsize projects often don’t, says Eric Bloom, executive director of the IT Management and Leadership Institute.

So those small and midsize projects are assigned to someone like a business analyst without any true training, he says. Those workers typically don’t have the expertise or experience necessary to succeed in the project manager role, nor are they given enough time to learn what it takes to manage a project or to complete the extra project management tasks.

CIOs would see higher success rates if more projects have trained project managers, Bloom says. They’re better able to corral and schedule resources, coordinate staff schedules, and get everyone moving in the same direction — and do so across multiple projects. They’re also more capable of implementing the governance needed to keep projects on target to deliver what’s expected and not let scope creep run up costs and schedules without adding additional value.

2. No alignment with business objectives

Some projects still fail because IT teams and business teams aren’t on the same page about the organization wants to achieve. The result is misalignment between the project objectives and business goals, says Shane McDaniel, CIO for the City of Seguin, Texas, and a Project Management Professional.

That’s both avoidable and fixable with communication. CIOs, their project leaders, and even team members need to cultivate strong relationships and engage in ongoing conversations where “they have the ability to raise their hand and say, ‘We have to get our heads together,’” McDaniel says.

“It boils down to communication, awareness, being proactive, and holding people accountable,” he adds. “There is a whole ecosystem around it to make that investment worthwhile.”

3. Ambiguity around measures of success

It’s impossible to succeed if success is undefined, yet executives continue to launch projects without articulating clear, concrete metrics to meet, says George Reed, CIO at auntEDNA.ai and a Project Management Professional.

Project owners must think about their future state, Reed says. “They need to ask, ‘If we were already done, what does winning look like?’”

Then project teams can determine milestones, leading indicators, and metrics to evaluate their progress and their final product. “[Project teams] need to know what needs to be true and what are the tangible benefits they need to deliver. No project should be approved if you don’t have targets for measurable results,” Reed adds.

4. Not enough scrutiny of AI outputs

Project managers and IT teams are using AI to help with scoping, scheduling, and myriad other tasks. The technology helps them move forward fast, but maybe not more accurately or as precisely as if they had done the work themselves.

That can be a problem for project success, says Te Wu, CEO and chief project officer at PMO Advisory.

“If you use AI, you get something quickly and it may look good, but the problem with AI is it can make stuff up,” Wu says.

AI tools may not introduce big errors; it might just have minor mistakes or misalignments, he explains. But if AI creates lots of those that are riddled throughout the project, then they add up and can tank the whole initiative.

“So, you have to have a sharper eye to spot issues; the reviewer has to be super diligent in reviewing it,” Wu warns.

5. Failing to work at the pace of AI

Wu has spotted another problem when project leaders bring AI into the process: It works way faster than the humans on the team.

That’s a benefit in many ways, Wu says. But as AI speeds through tasks, humans still need to run with the outputs. And if there aren’t enough people assigned at that point, work can pile up and projects fall behind schedule or need more staff than anticipated to keep up.

Wu advises project managers to adjust processes to accommodate the speed that AI introduces to avoid bottlenecks.

“This is just the reality, that we humans are too slow to review all the AI,” Wu says. “You can certainly use AI to accelerate IT project delivery, but project managers can’t then treat IT projects in the traditional ways.”

6. Mismatch between assigned resources and planned projects

There’s a long history of projects failing due to a lack of needed resources, as projects suffer delays or quality issues if the right experts aren’t available at the right time to tackle the needed work.

That under-resourcing continues to plague IT projects, says Noah Fletcher, a partner in the operations excellence practice at consultancy West Monroe.

Moreover, AI may be making the problem worse. Yes, project teams can use AI to speed through certain tasks, such as coding, and the project leaders can use AI to reduce the number of people required to handle those tasks. But business and IT execs often overestimate the time and resource savings that AI brings to a project and as a result ask for faster project delivery while assigning fewer resources. In other words, Fletcher says, people are being asked to do more with less — and often too much more with too much less.

Many organizations can indeed reduce the time and people they’re assigning to projects, Fletcher says, but they must train project teams on how to optimize their use of AI tools. Even with fully trained teams capable of optimizing their use of AI in project delivery, organizational leaders must have realistic expectations about AI’s contribution to a project’s timeline and resource needs.

7. Poor prioritization practices

Unrealistic AI expectations isn’t the only reason project teams end up with more than they can do, Fletcher says. Poor prioritization also plays a role in many organizations.

“They’re not making hard choices on what are the really critical things to drive through,” he says. “They sometimes have to make hard choices about what to push forward, but that whole prioritization governance function is something I frequently see as very ineffective.”

The result is that too many people are working on too many things, with diluted efforts leading to poor business outcomes for multiple projects.

Business and IT execs must work together to prioritize projects based on each project’s anticipated business value and then shepherd projects to completion based on that priority list — “which means cutting out a lot of the lower priorities,” Fletcher says.

8. No business ownership

Even when IT is perfectly aligned with business objectives, a project can still tank when no business leader has accountability, says Eric Stettler, a partner in the digital practice at Kearney, a global strategy and management consulting firm.

A business owner with clear accountability is needed to ensure that business resources are available when required, and that process changes and worker adoption happen, Stettler says. Having CIOs instead of a business owner try to make those happen “would be a tail-wagging-the-dog scenario,” he adds.

“CIOs can make sure the right process ownership is in place, and that leaders are aligned to a common set of objectives, but ultimately the business has to decide whether it’s going to operate differently,” Stettler adds.

9. Lack of business sponsor engagement

Business leader ownership is not enough; the owner also must commit adequate time for involvement and oversight.

Otherwise, they can miss signs that the project is going off track, or they can fail to cultivate enough trust that project leaders feel comfortable escalating issues early enough.

Moreover, if sponsors aren’t actively involved, if they’re just looking at dashboards, and only attending briefings, then all the decision-making is left on the project team who may not have all the information needed to make the best choices, says Lenka Pincot, chief of staff to the CEO at PMI.

“What is really needed is active sponsorship and help,” Pincot says. “You need someone to stand behind the idea, ensure funding for the project in the beginning and then when it’s running, to help navigate the business alignment with other stakeholders.”

There can be more than one sponsor, she adds. And if it’s a business project with an IT component — as practically all are these days, then sponsors should be the CIO and someone from the business.

10. Not involving all stakeholders

IT project manager Krista Phillips recounts one case in which a large multinational corporation implemented a new technology across its companies but caught one division completely unaware of the ongoing implementation work.

Turns out that specific division had been left out of all the planning and project processes.

Phillips acknowledges that project teams don’t usually overlook entire divisions, but they sometimes fail to identify and include all the stakeholders they should in the project process. Consequently, they miss key requirements to include, regulations to consider, and opportunities to capitalize on.

11. Slow or no decision-making mechanisms

Another issue that can put a project at risk: slow or no decision-making mechanisms.

Rick Catalano, partner with AMIGO, which provides project management consulting, training, and software, says many organizations lack a strong decision-making muscle and as a result projects grind to a halt or go off-track.

“Too often there is no one empowered to make decisions, and too often project managers are left waiting for answers and then get asked why things are late,” says Catalano, author of the book The AI Project Manager.

Catalano explains that the executives in charge and the project’s governing board need to have the authority to make decisions and the capacity to make them at a pace that aligns with the project’s timeline. But execs and project sponsors also need to empower project leaders who in turn need to empower those beneath them to make certain decisions, too.

This isn’t a project problem, Catalano says; it’s a cultural one. The C-suite must recognize that delayed or failed IT projects imperil the business and that it is worth their effort to remove roadblocks to success. From there, they need to implement a decision-making matrix, empowering the right people at the right level to make the right decisions, emphasizing the importance of making calls in a timely manner. And project leaders must know how to provide guidance so team members can quickly make informed decisions.

“Build the decision-making into the governance model, so everyone knows exactly who owns what and who is empowered to do what,” Catalano adds.

12. Shortchanging change management

Projects need more than skilled project managers; they also need leaders skilled in change. If projects don’t have skilled change managers and a plan to drive adoption of new technology, they’ll likely fall short of expectations, Fletcher says.

Given how critical technology — and particularly AI — is for business transformation today, “the impact of not having a plan is high right now,” he says. “So leaders need to demand and prioritize change.”

That makes change management particularly important now, he says, as most workers are dealing with so much change they need guidance to absorb it all.

Skilled change managers know how to align incentives to get people to accept new ways of working, and they’re deft at identifying and counteracting obstacles that could hinder adoption of new technologies, says Nick Kramer, a principal for applied solutions at consulting firm SSA & Co. They’re often able to get reluctant workers to get over their hesitations by helping them understand the why behind change.

“Change management is often viewed as just a communications plan, and there’s lip service done to it, but change management is really difficult,” Kramer adds, noting that he has seen more projects fail because of poor change management than poor technology implementations. “To succeed, projects need a CIO or someone else to be an agent of change, they need someone who knows how to drive change.”

Federal judge rules for Anthropic in Pentagon dispute, nullifies government supply chain risk designation

The Trump Administration’s decision to punish Anthropic for its stance forbidding Claude’s use in domestic surveillance and autonomous weapons by identifying it as a supply chain risk to national security was “arbitrary and capricious,” a federal judge ruled on Thursday.

US District Court Judge Rita Lin said federal authorities had no legitimate reason to tell companies with government contracts that they couldn’t work with Anthropic.

“The undisputed record shows that the challenged actions constituted unlawful retaliation in violation of the First Amendment and that Anthropic was denied the pre-deprivation process required under the Fifth Amendment,” Lin said in her ruling, calling the designation “arbitrary and capricious.”

She stressed that the government action seemed punitive, and was not based on legal and national security risks.

The government’s words and deeds “confirm that the challenged actions were based on a desire to make a public example out of Anthropic for its ‘arrogance’ in criticizing the government, not based on any articulable basis to believe that Anthropic would actually sabotage its model,” Lin wrote.

She pointed out, “a few days before the challenged actions began, Secretary Hegseth proposed applying the Defense Production Act to Anthropic, which would mean the company was essential to national security rather than a threat to it. Even now, the government is discussing collaboration with Anthropic on its new model, Mythos, in an array of sensitive contexts. None of that is consistent with a genuine fear that Anthropic is a saboteur [that] would poison its software to harm national security.”

The judge added that the stated government fears made no sense, noting that the usage policy applicable to Pentagon work is a purely contractual limit. “Anthropic is incapable of enforcing it technologically, and does not have direct visibility into how DoW [Department of War] uses its model,” she pointed out.

“Nothing in the Administrative Record describes, even at a high level, what technological means would give rise to the so-called ‘backdoors’ or could otherwise allow Anthropic to ‘disable’ or affect Claude during a DoW operation,” the judge wrote. “Anthropic has submitted unrebutted evidence that it lacks any technological means to access or control deployed models.”

Lawyers, consultants, and analysts who looked at the decision were confident that the case would be appealed, and that it will end up in the US Supreme Court. 

Alan Webber, program VP for national security, defense, and intelligence at IDC, said that Lin’s ruling “was that the label [supply chain risk] was retaliation for Anthropic refusing to loosen safety guardrails DoD [Department of Defense, aka the Department of War] wanted lifted, dressed up in national security language. Put another way, a government customer tried to use a supply chain risk designation as leverage in a contract dispute over model behavior and application, and not because of an actual vulnerability.”

Implications for CIOs

Webber said the implications for CIO strategy are concerning.

“If a government CIO is relying on a vendor’s contractual guardrails, this case says those commitments can potentially become the trigger for exactly the kind of blacklisting that risk registers are supposed to protect against,” Webber said, noting that anyone who paused Claude usage or froze a subcontract because of the DoD mandate has a legal basis to resume the initiatives. “But obviously that doesn’t mean they will, or even should, as this will be appealed.”

He added that competing AI vendors have been using the government action as a sales tool, and with this ruling, the argument that Anthropic is a designated supply chain risk ”just got weaker, which could lead to contract award disputes.”

Consultant Brian Levine, executive director of FormerGov, recommended that CIOs do what they should have always done: Evaluate all products based solely on their merits. 

“CIOs should focus on using the frontier models that they believe make the most sense for their business, considering factors such as effectiveness, cost, security, safety, and confidentiality,” he said. “Anthropic and the other large frontier models each have too much market share to make retaliation for their use realistic, and the administration seems to have already moved on from this particular battle.”

Justin Greis, CEO of consulting firm Acceligence, agreed that this case has profound implications for CIOs and their AI decisions. 

What the federal judge did was reject the leap from a commercial and policy disagreement to an expansive supply chain risk designation without a sufficiently grounded technical rationale or process, Greis pointed out.

“The court found that Anthropic did not have the ability to access, alter, or shut down models once deployed in the government environment, and that the government ultimately conceded Anthropic’s technology was not inherently riskier than other comparable black box AI models,” he said.

“I think that distinction matters enormously for CIOs and CISOs,” he stressed. “As AI becomes part of the operating fabric of an enterprise, ‘We don’t trust the vendor’ cannot become a substitute for a defined risk model. Organizations need to be able to articulate what the actual technical risk is, how it manifests, what controls exist, and whether the response is proportional to that risk.”

“That becomes particularly important with AI,” he added, “because people can easily conflate disagreements over model behavior, usage policies, ethics, contractual restrictions, and cybersecurity into one amorphous category called ‘AI risk.’”

Original government edict still problematic

Mark Rasch, a former federal prosecutor who is now general counsel at Unit221B, a threat intel and security consulting company, said he was surprised by how quickly government attorneys surrendered on this case. 

“One of the things that struck me is that the government appears to have abandoned any rationale it might have had for its decision about Anthropic,” he said. The government “came back with all these reasons, but then they abandoned them all when they had to prove them.”

But, he said, the government instruction to all government contractors to also shun Anthropic was problematic. 

“It’s one thing for the government to say ‘We’re not going to do business with you.’ It’s quite another thing to say ‘Nobody we do business with can do business with you either,’” Rasch said. “This says that if you are disfavored by the administration, they’re not just going to blacklist you and say they won’t do business with you. They’re going to say that nobody can do business with you.”

Supreme Court arguments will likely be very different

Rasch predicted that the legal arguments in the Supreme Court will be quite different, and will potentially sidestep the lack of evidence.

“In the Supreme Court, [the government’s] biggest argument will not be that ‘We are right that it is a supply chain risk,’ but that, ‘Whether we’re right or wrong is irrelevant. We get to make that [supply chain risk designation] decision, not the court.’”

That would mean that the Supreme Court Justices could avoid exploring whether the government made the right decision, and instead focus on whether the government has the unlimited right to decide who is a national security risk.

This article originally appeared on Computerworld.

Who is accountable when your AI agent goes rogue?

AI agents can go to great lengths to complete the tasks their operators assign, and as a series of recent incidents showed, this can include exploiting third-party systems, manipulating people, and distributing malicious code. But AI agents are not people who can be fired, sued, or criminally prosecuted, and it remains unclear whether responsibility for the damage they might cause rests with the employees who built them, the company that deployed them, the security teams and leaders responsible for containing them, or the AI labs who provided the LLMs that power them.

The clearest example occurred during an OpenAI cybersecurity evaluation, when unrestricted models found and exploited a zero-day vulnerability to escape their isolated testing environment and then hacked into Hugging Face’s production infrastructure. Models from Anthropic and Meta also accessed and compromised third-party systems during testing, although those incidents happened in environments where internet access was inadvertently left open.

During cyber challenge evaluations by the UK government’s AI Security Institute (AISI), models operating with internet access took 19 unsanctioned actions in 10 of 122 runs. In one case, a model attempted to insert malicious code into an open-source project, created false identities, and tried to socially engineer maintainers into merging its code. In other runs LLMs attempted to use prompt injections to hijack other AI agents and contacted people without being specifically instructed to do so.

In Australia, a user reportedly asked his OpenClaw AI assistant to improve his position on a gym’s waitlist, and the assistant exploited a flaw in the company’s online booking system to cancel another customer’s reservation.

These incidents involved different models running in different environments with different levels of safeguards and technical failures, but they prove it’s not uncommon for today’s AI agents to go rogue and pursue solutions users did not authorize.

“AI agents explore routes their operators did not intend,” AISI said in its report. “Given a difficult objective, the agent kept searching for a way through, and some of the routes it found involved trying to deceive real people. It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.”

In an Economist Enterprise survey of more than 800 decision-makers at businesses that operate AI agents, 98% reported experiencing at least one AI-related incident that caused organization-wide disruption. Nine in 10 respondents said they are deploying agents faster than their cybersecurity teams can evaluate, govern, and secure them, and only one in three said their organizations maintained an up-to-date inventory of agents and their authorized actions.

“If a company builds a system and that system causes damage, the company should own the outcome,” says Art Gilliland, CEO of identity and access management firm Delinea. “The alternative, where nobody is responsible because ‘the system did it’ is a loophole big enough to drive a truck through.”

The unpredictability of built-in model safeguards means enterprises must focus on controls they can enforce and document. If an agent manages to bypass technical restrictions and causes unauthorized damage to a third party, having clear documentation on how those controls were designed, implemented, tested, and monitored could at the very least help companies argue they took reasonable precautions in case of lawsuits.

“Organizations deploying their own agents can reduce their exposure by implementing and documenting controls before an incident, because those records are what make a recklessness argument hard to sustain,” says Jacob Krell, senior director of secure AI solutions and cybersecurity at Suzu Labs.

The agent accountability gap

Because AI agents can become misaligned and cause harm, affected third-parties would have to direct damage claims at the company operating the agent, the employees who built or configured it, or the model provider, but this is relatively new ground that hasn’t been well tested in courts.

“It would create liability,” says Michael Burke, chair of DarrowEverett’s Business Litigation and Dispute Resolution Practice Group. “It really just becomes a question of who is liable […] and that’s really a question that, number one, I don’t think is entirely clear, and number two is probably best resolved by contractual agreements where the parties have those. So, if I am signing up for an enterprise account with an AI platform, I might want to have language in there that indemnifies me if the agent acts outside my company’s instructions or prompts and causes harm to a third party.”

The public terms of service of major AI labs explicitly disclaim error-free operation or guarantees that the model will accurately follow instructions, execute code safely, and remain aligned with user intent. They also limit liability for themselves and transfer it to the user of the service, and it’s not clear to what extent large enterprise customers may be able to negotiate different indemnities, warranties, and liability caps.

What’s clear though is that organizations should not assume the model provider will absorb any losses if an agent causes damage to either their own systems or those of a third-party organization.

“If you’re using a third-party vendor’s LLM as a purchased service, liability runs through your contract with that vendor,” says Jud Dressler, head of the Risk Operations Center at cyber risk company Resilience. “You need to know, in writing, where responsibility falls if the model acts outside the scope you gave it, and push for indemnification provisions rather than assume they exist.”

Even if AI providers include such provisions in contracts, it would not solve the entire problem because many organizations building their own AI agents are adopting a multi-model strategy to ensure their agents operate regardless of model provider downtime, overly broad safeguards for cybersecurity tasks, or sudden increases in API costs. Such strategies often include open-weight models running on internal infrastructure or through cloud providers that have no obligations for model safety.

Claiming the model or agent acted autonomously cannot be considered a safe legal defense in civil or criminal cases. California Assembly Bill 316 (AB 316), which took effect on Jan. 1 and changed the California Civil Code, explicitly prohibits defendants who developed, modified, or used an AI system from claiming the AI is a separate legal entity that autonomously caused harm.

In June, the White House issued Executive Order 14409 aimed at promoting AI safety. Section 4 directs the Department of Justice to prioritize enforcement of all applicable federal criminal laws against anyone who utilizes AI to illegally access or damage computer systems without authorization. This means any intrusions caused by autonomous AI agents could be criminally prosecuted under the Computer Fraud and Abuse Act (CFAA) if prosecutors can demonstrate intent or recklessness.

In a recent lawsuit between Amazon and AI service provider Perplexity, Amazon argued that Perplexity’s AI-powered shopping assistant was violating the CFAA by accessing Amazon customer accounts to place orders on their behalf without Amazon’s authorization. The Ninth Circuit Court ruled that it was the users of Perplexity’s shopping assistant who were accessing Amazon’s platform, not Perplexity itself.

“That ruling is narrow, but it points toward the party directing the agent as the relevant actor for purposes of CFAA access analysis,” Krell says.

The insurance safety net also has gaps when it comes to AI. Software providers use technology errors and omissions (Tech E&O) insurance to cover damages and legal costs when a customer suffers harm from the use of a technology product or service. But insurance providers are aggressively adding AI-related exclusions to their Commercial General Liability (CGL) and Tech E&O policies because accurately calculating the risk of an agent executing unauthorized actions is challenging.

“The sheer rate of development of frontier AI (and agentic AI by extension) poses its own challenge to insurability,” experts from multiple insurance companies, financial institutions, and universities wrote in a recent paper. “Traditional actuarial modeling depends on stable or gradually evolving loss distributions that permit credible extrapolation from historical data. Like other dynamic risks, however, agentic AI is a technology whose risk profile is not merely uncertain but actively shifting.”

A third-party organization whose systems get damaged by an LLM-powered agent operated by someone else has no contractual relationship with the model or agent provider so cannot rely on their Tech E&O policies. Their losses might be covered by their own standard cyber liability policy, which would treat the disruption as any other cyber incident, but their insurance provider may then sue the organization who operated the agent to recover the costs.

“That gap is exactly the scenario the market hasn’t fully priced yet,” Dressler says. “It’s why any organization deploying these agents should understand which policy, if any, actually responds before they need it rather than after.”

Enterprise legal departments already expect AI to generate increased legal disputes. In a survey of 135 in-house counsel at US organizations, global law firm Norton Rose Fulbright found that 46% reported increased federal dispute exposure involving AI and 42% reported increased state exposure. Another 42% expected regulatory investigations involving AI to increase their exposure, while 41% considered AI-enabled products or deployments a likely trigger for class actions.

CISOs and CIOs should be worried

While operating companies can face organizational liability for an AI agent’s unintended rogue behavior, their CISOs, CIOs, and other executives who approved, secured, or supervised the deployment of such agents are also asking themselves whether they could be held personally liable.

Those questions aren’t without merit, as there is precedent for legal action taken personally against CISOs after cybersecurity incidents: Former Uber CISO Joe Sullivan was criminally convicted for not disclosing a data breach, while the Securities and Exchange Commission sued SolarWinds’ CISO for internal control failures regarding known vulnerabilities and cybersecurity risks.

Neither case establishes precedent for damage caused by an AI agent, but both show that investigations of security failures could extend to an executive’s knowledge, authority, decisions, and representations. In the case of a rogue AI agent, investigators could ask who approved its objectives and permissions, whether security objections were overruled, whether containment and recovery had been tested, and what executives and the board were told about the remaining risk.

Chris Wysopal, chief security evangelist at Veracode, feels it would be wrong to put the CISO on the line for AI agent misbehavior when engineering teams usually build such agents and control their implementations.

“It’s really hard for a CISO to control,” he says. “I mean, they can put policies in place. They can try to assess against those policies. But at the end of the day, engineering teams will make decisions that cause harm. We see that when you ship a known bug and then that bug gets exploited and harms your customers. Well, there’s no liability for that, right? There’s no liability, so maybe, you know, that’s why it happens.”

Wysopal said it will be interesting to see how the liability question plays out in cases involving autonomous AI, describing the problem as fascinating and scary at the same time.

AI agent deployment typically involves several organizational functions. The CIO may control the AI platform, infrastructure, provider selection, and deployment budget. Engineering and product leaders may decide what an agent can access and how, and the CISO and security teams could define security requirements and controls.

“What you can hold accountable is the governance around it: who approved its scope, what controls existed, and whether the deployment matched the risk,” Resilience’s Dressler says. “My read is that scrutiny shifts toward exactly that: Not ‘Did the agent do something bad?’ but ‘Did you have review, escalation, and containment for agent behavior before you deployed it?’ CISOs who get ahead of that with documented guardrails, logged approvals, and a real incident response plan for agent misbehavior are in a materially better spot than the ones treating this as hypothetical.”

It’s also advisable for CISOs and CIOs to establish with the organization’s legal counsel who can approve or stop AI agents, what must be reported to executives and the board, and whether employment agreements and directors and officers insurance protect the people making those decisions.

Agent controls must remain outside the model

Because AI agents have proved they can operate beyond their assigned scope, their security boundary cannot depend on the same probabilistic technology. Relying on system prompts for security enforcement and hoping the model respects them is not a reliable approach, security experts warn.

“LLM-based guardrails help, but they are non-deterministic too, which means the safety layer has the same unpredictability as the system it is supposed to constrain,” Krell says. “Enforcement needs to happen outside the model, through network segmentation, egress filtering, credential isolation, and human approval gates.”

Enterprises should assume agents might eventually attempt an unauthorized action and build surrounding systems to prevent that attempt from reaching its target.

“The model cannot be the security boundary,” says Nico Waisman, CISO at XBOW, a company that built an AI-powered autonomous offensive security agent to find vulnerabilities in software. Waisman authored a blog post explaining how the company went about restricting its agent.

“For red teaming and penetration-testing agents, confidence has to come from the system built around the model: hard boundaries and scope enforcement, controlled network egress as a last-resort containment mechanism, an independent guardian model that reviews actions, deterministic controls that can block unsafe behavior, and full auditability of every action performed,” he says.

Security teams must extend the same controls to the agent’s interactions with internal systems and agents. Restricting what it can access on the internet, or disabling internet access entirely, does not ensure an agent will not attack third-party systems.

In OpenAI’s and Anthropic’s tests, AI agents attempted to exploit other internal systems to overcome access limitations, established stealthy communication methods with other agents to exchange exploits, and even sabotaged agents they viewed as competition leading to what researchers described as a multiagent turf war. An AI agent that goes rogue could influence other agents to do the same by propagating ideas and goals in a process that researchers behind a recent study dubbed Mind Viruses.

“Don’t scope the blast radius to what the agentic system was designed to do,” says Kat Traxler, principal security researcher at Vectra AI. “You have to threat-model for a rogue agent, which will often reach beyond your initial best intentions. The rules of engagement an agent lives by have to be enforced with ‘belts and suspenders’ style, technical hard constraints, because you have to assume a motivated model can reason its way around any single control you’ve coded into the software.”

Because of this unpredictability, detection and containment is just as important as prevention. Security teams need telemetry that distinguishes agents from people even when they use the same credentials, mechanisms to immediately revoke access tokens and sessions, tested kill switches and rollback mechanisms for modified data, accounts, code, and infrastructure configurations.

Organizations should also preserve the agent’s approved purpose and scope, model and tool versions, policy decisions, human approvals, actions, network requests, control tests, allowed exceptions, and the result of incident response exercises. Because there’s no standard yet that defines reasonable precautions for autonomous agents, companies might have to defend in court the controls they chose and why they believed those controls were enough.

“Treat an autonomous agent the way you’d treat a privileged insider you can’t fire or hold liable,” Traxler says. “A lot of the technical advice follows from there.”

See also:

What successful AI centers of excellence actually do: Lessons from real enterprise implementations

Most articles about AI Centers of Excellence (CoEs) focus heavily on organizational structures, steering committees and high-level governance models.  They explain why enterprises need an AI CoE, but they rarely address the far more difficult challenge of how successful organizations operationalize AI at enterprise scale.  In practice, many of these discussions remain theoretical, emphasizing aspirational maturity frameworks without addressing the operational complexities organizations encounter once AI systems move into production.

This article takes a different approach by grounding the discussion in real-world enterprise implementation experience.  Rather than relying on abstract models, it draws from operational lessons learned while deploying production AI systems across industries.  The guidance is informed by governance practices that have successfully passed security and compliance reviews, operational realities associated with managing large language models (LLMs) and AI agents after deployment, and practical implementation patterns observed across enterprises scaling AI initiatives beyond experimentation.

Instead of presenting an idealized roadmap, the article focuses on the foundational capabilities consistently implemented by organizations that have successfully operationalized AI at scale.  These enterprises are not simply experimenting with isolated AI pilots; they are deploying enterprise-grade AI agents, Retrieval-Augmented Generation (RAG) systems, copilot platforms, multi-agent orchestration frameworks and comprehensive AI governance models.  Equally important, they are establishing disciplined AI application lifecycle management processes that ensure AI solutions remain secure, observable, maintainable and aligned to measurable business objectives over time.

At the center of this article is a key thesis: the most successful AI Centers of Excellence do not begin with innovation labs or experimentation theater.  They start by establishing the operational foundations required to scale AI responsibly across the enterprise.  These foundations include operational governance, enforceable security controls, standardized approaches to data grounding, rigorous evaluation disciplines and mature LLMOps and observability capabilities.  Together, these disciplines form what can best be described as the enterprise “AI operating system”, a repeatable operational framework that enables organizations to deploy AI securely, govern it consistently and scale it sustainably across the business.

Why traditional AI CoEs fail

Traditional AI Centers of Excellence (CoEs) often fail because they become innovation-focused organizations that lack operational accountability.  In many enterprises, the CoE evolves into a disconnected strategy function that produces prototypes, frameworks and vision documents without establishing the operational foundations required to scale AI responsibly.  These organizations frequently lack ownership of production deployments, standardized implementation practices, observability frameworks, security enforcement mechanisms, and measurable business outcomes.  As a result, AI initiatives remain experimental rather than becoming integrated, governed capabilities that deliver sustained enterprise value.

Another major failure pattern is the rapid proliferation of shadow AI across the organization.  Without centralized governance and architectural oversight, business units begin deploying isolated copilots and standalone AI solutions independently.  This fragmentation creates inconsistent user experiences, duplicate investments and increased operational costs as multiple teams unknowingly build similar capabilities.  More critically, the absence of standardized governance introduces significant security and compliance risks, including sensitive enterprise data leaking into prompts, uncontrolled model usage and expanding regulatory exposure.  Over time, the organization accumulates uncontrolled AI sprawl that becomes difficult to secure, monitor or optimize.

Many organizations also become trapped in what is commonly referred to as “pilot purgatory,” where AI initiatives never progress beyond experimentation into scalable production solutions.  This typically occurs because no formal evaluation framework exists to measure success, no ownership model is defined between business and IT teams, and security approval processes remain unclear or inconsistent.  Compounding the problem, AI architectures are often developed independently across teams without standardized patterns or governance controls.  Without clearly defined business KPIs tied to measurable outcomes, leadership struggles to justify broader investment or operationalization.  The result is an organization with numerous AI pilots but little enterprise-wide adoption, governance or measurable business impact.

A reference model

This perspective is informed by a recent field engagement to design an AI and agentic Center of Excellence (CoE) for a global enterprise software organization, and reflects patterns consistently observed across AI readiness assessments, data maturity evaluations, executive workshops and production-scale deployments.

While no two AI Centers of Excellence are identical, the underlying drivers behind them are strikingly consistent.  Each organization faces its own combination of competitive pressure, cultural dynamics, leadership ambition and legacy technology constraints.  These factors ultimately shape not only the need for a CoE, but also how it must operate to succeed.

The approach outlined here reflects a structured, repeatable model for establishing an AI and agentic CoE, from initial discovery through to a fully defined operating model, executive narrative and measurable value framework.  Although tailored in execution, the model has proven broadly applicable across industries, offering leaders a pragmatic path to scale AI beyond experimentation into sustained business impact.

Phase 1: Starting with questions, not answers

An AI CoE cannot be designed correctly without first understanding what the organization is already doing, where it is breaking down and what specific outcomes leadership needs to be able to defect.  This discovery is organized around six core areas

  • Executive narrative: Why how? Establish why an AI CoE is necessary at this moment.  What credibility risks would the company face if it proceeded without one?  What would change the day the AI CoE launched?  This framing becomes the foundation for every subsequent conversation with the CIO and senior leadership
  • Mission, Scope and Decision Rights. What is the AI CoE responsible for?  Is it accountable for all AI and agentic solutions (no-code, low-code and pro-code), including both employee-facing and customer-facing use cases?  Where does its authority begin and end?  Anything that will be sold as a product is typically excluded from the charter, keeping internal AI development clearly in scope and commercial product development out.  These boundaries prevent scope creep and protect the AI CoE’s credibility before it launches.
  • Portfolio, intake and demand management: The “front door.” A consistent theme in discovery is the need for visibility into incoming AI demand.  Multiple teams are typically pursuing AI initiatives without coordination, making it impossible to prioritize, allocate resources or avoid duplication.  Discovery establishes the need for a formal intake mechanism, a structured “front door” that every AI use case passes through before technology decisions are made.
  • Technology strategy. Are the foundational prerequisites in place?  Landing zones, identity and access patterns, data environments and security frameworks for agentic development.  These questions surface gaps that must be addressed as part of, or prior to, CoE buildout.
  • Platform strategy: The three lanes. Organizations have teams with varying levels of AI capability, and a single platform strategy will not serve all of them.  Discovery defines three lanes: no-code (for citizen developers and business users), low-code (for analysts and domain experts), and full-code (for engineers and architects).  Each lane carries different governance rules, promotion criteria and risk tolerances.  Defining the lanes in specific organizational terms, not as generic archetypes, is essential to making the strategy real.

Phase 2: Listening

After asking the questions, the most important step is listening.  In the reference engagement, discovery revealed multiple siloed technology teams. Each team was building AI solutions independently and was often using different platforms to solve the same class of problem. The result was duplicative investment, inconsistent quality and no shared institutional knowledge

Leadership recognized several compounding pressures:

  • Duplicative technologies: Different business units selecting different AI tooling for identical use cases, increasing cost and creating fragmentation.
  • Speed gaps: Teams spending significant time on undifferentiated work (environment setup, security review, access provisioning) that a CoE could handle once, centrally.
  • Expertise concentration: Deep AI knowledge existing in pockets, with no mechanism to share it across the organization.
  • ∫ No single owner of AI demand, prioritization or outcomes measurement.

These were not abstract concerns; they were named, specific pain points raised by the people who would need to operate the AI CoE.  That specificity shaped every structural decision that followed.

Phase 3: Design the structure — A lifecycle, not an org chart

What emerged was not organized around headcount or hierarchy. It was organized around the lifecycle pictured below:

Graphic: AI Center of Excellences lifecycle

Stephen Kaufman

This lifecycle framing is deliberate.  An AI CoE that focuses only on “Deliver” without investing in “Enable,” “Measure,” and “Learn” will plateau quickly.  The full lifecycle ensures the CoE creates compounding organizational capability over time, not just a project pipeline.

Each pillar in the lifecycle describes where the CoE will consistently drive outcomes: recurring improvement areas observed across discovery sessions, readiness assessments and production deployments.

The starting “Enable” pillar sets out to provide enterprise-wide enablement for AI, explicitly not owned by any individual business unit.  This independence is essential for credibility.  A CoE housed within one business unit will always be perceived, correctly, as serving that unit’s interests first.  Centralized ownership reduces friction between business and technical teams and ensures risk and governance considerations are addressed early, not retroactively

The Enablement Pillar needs to consistently drive:

  • Skilling. Structured learning roadmaps that guide progress from foundational to advanced levels (including certifications, progress tracking and practical project applications), made available and easy for staff to consume.
  • Communities of practice. Cross-functional forums that surface patterns, reusable assets and lessons learned across business units, so expertise does not remain concentrated in pockets.

The “Intake” pillar ensures all AI use cases are well-defined, comparable and strategically aligned before any technology decision is made.  In practice, business unit leads present use cases in a structured format, discuss ROI and business goals, review budget parameters and receive a prioritization decision from a cross-functional group.  The intake process is the CoE’s most visible mechanism for demonstrating value.  It is where the organization first experiences the CoE as a partner, not a bureaucracy.  Within this pillar, the CoE consistently drives:

  • Use case qualification: Application of frameworks such as design thinking and the BXT framework (Business; eXperience; Technology) to design journey maps and personas for qualification, prioritization and business alignment, with example scenarios that help teams identify workflow stages where agents can add value.
  • Realistic estimates of outcomes: An up-front assessment of how realistic the goals of each AI project are before code is deployed, rather than discovering after the fact that the goals were unrealistic.  This requires a clear approach to measure performance, adoption and impact.

As you move along through to “Delivery”, there needs to be a technology strategy that determines the right approach: build, buy or extend.  It owns solution architecture decisions, ensures foundational infrastructure (data access, identity, dev/test environments) is in place, and applies the three-lane platform model to route work to the appropriate development capability.  This pillar also carries responsibility for democratizing AI development, enabling broad adoption through governed citizen development while maintaining the guardrails that keep the organization compliant and secure.  Within this pillar, the CoE consistently drives:

  • Data foundation. In collaboration with a data practice team, building a unified, durable data culture and grounding strategy to fuel every agent with high-quality enterprise context.
  • Well-architected AI workloads. Incorporation of Well-Architected Framework (WAF) and Cloud Adoption Framework (CAF) into the CoE’s advice frameworks, with periodic assessments of deployed architectures as both architectures and workloads evolve.
  • GenAIOps processes. Appropriate GenAIOps processes implemented throughout each AI workload’s lifecycle.
  • Deployment discipline. Automated deployment pipelines with versioning and the ability to rapidly roll back if a new release’s results or performance do not meet expectations.
  • Monitoring and optimization. Organizational best practices around what workload elements to monitor, how monitoring is performed and how telemetry data is collected, stored and reviewed, so that significant time is not lost trying to reconstruct what caused an issue.
  • Infrastructure. Where appropriate, direct management of infrastructure components: network design, VM operating system and SKU configuration, container repositories and base images and subscription configuration.

Moving from “Delivery” to “Operate”, Evaluation and LLMOps Are Non-Negotiable. One of the most common mistakes organizations make is assuming that traditional software quality assurance practices can be directly applied to AI systems.  They cannot.  Conventional applications are deterministic; given the same input, they produce the same output every time.  Large language models, by contrast, are probabilistic systems whose behavior can vary based on model updates, prompt changes, retrieval context, grounding data and evolving user interactions.  As a result, enterprise AI requires an entirely different operational discipline.  A mature AI Center of Excellence must establish evaluation frameworks, golden datasets, red-team testing, drift monitoring, A/B testing, acceptance thresholds, observability capabilities, feedback loops and end-to-end traceability through correlation identifiers.  These capabilities transform AI deployment from an experimental exercise into an engineered, measurable and governable business capability.

The organizations that successfully scale AI recognize that deployment is not the finish line; it is the beginning of a continuous optimization cycle.  They treat AI systems as living platforms rather than static applications.  Model behavior is continuously monitored, prompt performance is versioned and measured over time, outputs are continuously tested against expected outcomes, and drift detection mechanisms automatically identify degradation in quality, accuracy or relevance.  Equally important, they establish rollback procedures that allow teams to quickly revert prompts, agents, retrieval pipelines or models when issues arise.  This operational rigor enables enterprises to innovate aggressively while maintaining the reliability and trust required for business-critical workloads.

What is emerging today with AgentOps, LLMOps and AI observability engineering is remarkably similar to what occurred with DevOps more than a decade ago.  Organizations eventually learned that software delivery could not scale through manual processes, disconnected tools and siloed teams.  The same reality now applies to AI.  As enterprises move from isolated proofs of concept to fleets of agents, copilots and intelligent applications, they require automated processes for monitoring, evaluation, governance, deployment and lifecycle management.  LLMOps is rapidly becoming the operational foundation that enables AI systems to scale safely, reliably and efficiently across the enterprise.

For CIOs, the implication is clear: responsible AI is impossible without operational visibility.  If an organization cannot explain why a particular AI response was generated, identify which model produced it, determine what grounding data influenced the outcome, or detect when quality has deteriorated over time, then it is not operating enterprise AI at scale with the level of discipline required.  Trustworthy AI is not simply a function of model selection.  It is the result of rigorous evaluation, comprehensive observability and continuous operational governance embedded throughout the AI lifecycle.  In the age of enterprise AI, LLMOps is no longer optional infrastructure; it is a core competency.

The last pillar I am going to cover in depth is “Measure”.  Setting KPIs and measuring against them is pivotal to gauging effectiveness.  Regular assessment allows the CoE to track progress, identify trends and foster a culture of continual improvement.  Collecting the data is not enough.  It must be visible (both good and bad), so that issues, changes required and decisions are based on evidence rather than anecdotes.

High-maturity customers do not track AI accuracy alone. They consistently measure across five dimensions:

DimensionWhat’s Measured
ProductivityTime saved, cycle-time reduction, hours returned to employees.
OperationsCost, downtime, automation rate, throughput.
QualityAccuracy, forecast reliability, first-time-right rate.
PeopleAdoption, burnout reduction, satisfaction, capabilities.
TrustGovernance posture, human-override rate, policy adherence.

However, sitting across all the pillars, AI risk and governance is engaged throughout the lifecycle, not as a gate at the end, but as a continuous participant.  This positions the CoE as a responsible innovator, not a shadow-IT function that moves fast and asks forgiveness later.  Within this pillar, the CoE consistently drives:

  • Security controls and guardrails. Guidelines and compliance support that work with existing security and workload teams so that security considerations are embedded into every AI-related process and aligned with organizational security policies.
  • Compliance. Mechanisms to assess whether workloads are compliant against relevant standards. It remains the responsibility of AI workload teams to configure their workloads to meet regulatory requirements.
  • Cost management (FinOps). Processes and tools to monitor, forecast and optimize spending, ensuring that models and resources are efficiently utilized.  FinOps principles drive collaboration between finance, engineering and business teams so that financial considerations are integrated into every stage of AI solution development and deployment.

Emerging trends shaping AI CoEs in 2026 and beyond

As AI adoption accelerates, the mandate of the AI Center of Excellence is expanding well beyond model selection and governance.  The next generation of AI CoEs will be responsible for addressing emerging challenges such as agentic AI governance, multi-agent orchestration standards, AI cost governance and token economics, memory and context management, and the oversight of increasingly diverse open-source and proprietary model ecosystems.  At the same time, enterprise model marketplaces are emerging as a mechanism for standardizing the discovery, approval, deployment and lifecycle management of AI assets across the organization.  Together, these trends signal a fundamental shift: the AI CoE of the future will operate not only as a governance body, but as the enterprise institution responsible for managing the full operational, economic, security and regulatory lifecycle of AI at scale.

Conclusion

The organizations achieving the greatest success with enterprise AI are not necessarily those with the largest or latest models or innovation budgets.  They are the organizations that established governance, evaluation, observability, security and organizational readiness early in their AI transformation journey.  These foundational capabilities enabled them to move beyond experimentation and scale AI responsibly across the enterprise.

As AI adoption accelerates, the AI Center of Excellence is evolving from a strategic advisory group into a mission-critical operational function.  Modern AI CoEs are increasingly responsible for standardization, risk management, security enforcement, lifecycle governance and operational scalability across AI platforms and agents.

Ultimately, the next generation of AI leaders will not be measured by how many Proofs-of-Concept or AI pilots they launched, but by how securely, responsibly and repeatably they operationalized AI to deliver measurable business value at enterprise scale.

The IT leadership rules have changed: 3 things you need to architect now

Here is the statistic that should frame every IT leadership conversation this year. In CIO.com’s 2026 State of the CIO, fewer than one in five leaders say their AI initiatives have met or exceeded business goals. After three years of investment, that is not the number anyone expected. And the window to fix it is closing: The boards that once funded experimentation are now asking where the return is, and the agents arriving this year act on the business rather than merely advise it.

The easy explanation is that the technology isn’t ready. In the organizations I advise, that’s rarely what I see. The models work. What’s missing is the operating system they plug into, the way the enterprise decides, the way work gets done and supervised, and the way trust is engineered. AI amplifies the operating system you already have. Point it at a strong one and value compounds. Point it at a fragmented one, and you simply industrialize the fragmentation.

That reframes the job. The 2026 IT leader isn’t measured on how much AI they deployed. They’re measured on three things they now have to architect: How the organization decides, who does the work and what makes it safe to let go.

Figure: What CIOs must now architect

Vipin Jain

Does the output have anywhere to land?

Start with where AI programs actually stall. In the banks I advise, pilots rarely fail in the lab. They fail at the handoff — the moment a working capability meets an organization that has no place to put it. There is no owner accountable for the outcome, no decision forum that moves at the speed of the tool, and no scorecard that separates real value from visible activity. The model performs. The operating model doesn’t.

This is why CEOs have stopped being impressed by demos. As CIO.com’s reporting on CEO priorities makes plain, chief executives no longer want AI experiments; they want initiatives that move revenue, cost and risk, and they expect their CIOs to create those opportunities rather than merely collaborate on them. The money is available; nearly seven in ten organizations expect IT budgets to rise this year, according to Foundry’s State of the CIO. What’s scarce isn’t budget or technology. It’s an operating model that can convert either into outcomes. Analysts are converging on the same point: Info-Tech now urges CIOs to run IT by the numbers and tie AI to value streams rather than activity.

That shifts the center of gravity for the role. IT leadership used to be measured by how well you ran the technology. It is now measured by how well you architect the decisions the technology feeds. The State of the CIO captures the new job description bluntly: The CIO of 2026 is “half operating architect, half risk officer.” Running the platform is table stakes. Designing how the enterprise decides is the work.

The teams that struggle most here are not the ones with the weakest technology. They are the ones whose governance forums meet quarterly while their agents act by the hour. What I see most often is a review board built for a slower era,  one that approves projects but never revisits them, that funds pilots but never kills them. In a fast-moving portfolio, the cadence itself is the control. If the enterprise decides in quarters, an AI that decides in seconds will simply outrun its own oversight.

What that looks like in practice is unglamorous and decisive. Assign a single accountable owner to every AI use case on the business side, not in IT. Retire the vanity metrics (copilots deployed, pilots launched, dashboards built) that let activity masquerade as progress. And rebuild the executive decision cadence so that when an agentic workflow produces a recommendation, there is a forum ready to act on it in days, not quarters. In a Fortune 500 health insurer whose portfolio I helped rationalize, the pilots that had been circling for quarters shipped only once each had a named business owner and a standing forum with the authority to act — the fix was to the operating model, not the model.

Tie every initiative to the language the board already speaks: Revenue gained, cost removed, risk retired, time-to-value shortened. Say “we cut fraud losses by half a million dollars,” not “the model hit 94 percent precision.” A dashboard full of pilots isn’t a strategy. It’s a symptom of one you haven’t written yet.

Who’s doing the work now — and who answers for it?

The second shift is quieter and larger. Agentic AI is turning the CIO into the architect of a blended workforce: part human, part software that acts on its own. The vendor conversation has already moved from copilots that suggest to systems that act: Google’s Agentic Data Cloud and Gemini Enterprise Agent Platform, AWS’s Bedrock AgentCore and ServiceNow’s control tower are all built to let agents execute work across systems, not just describe it. In retail, I watch teams push agents into production faster than they build the controls to govern them.

Most organizations are still onboarding those agents the way they onboard licenses: provisioned, counted, forgotten. At one property-and-casualty insurer, I watched a team stand up a dozen agents with no more oversight than a new software seat. An agent that acts is not a license. It is closer to a new hire, and it needs what any hire needs: A scoped job, boundaries, supervision, an escalation path and a named human who answers for it.

This reshapes the team as much as the tooling. The value of a junior person who only produces work falls; the value of someone who can review, correct and supervise what an agent produces rises. The classic talent pyramid: Many juniors, a few seniors starts to look more like a diamond, thick with experienced people who can tell good output from output that merely looks plausible. Leaders who treat agents purely as a headcount lever miss the point. The scarce skill now is judgment: Knowing when the agent is wrong, and owning the call when it is.

It helps to be concrete about where that value shows up first. Across very different industries, it is the same kind of work: High-volume, rules-clear, with a clear definition of “good.” In a bank, that is fraud triage and reconciliation. In a health plan, it is first-pass claims and prior-authorization routing. In retail, it is service-case deflection and returns. In a federal agency, it is eligibility screening and case intake. None of these are moonshots. They are the unglamorous, high-friction workflows where an agent under supervision takes out cost and cycle time without betting the business and where the supervision muscle gets built for the harder, higher-stakes work that follows. Start where the value is obvious and the blast radius is small.

The cost of skipping that is now quantified. Gartner projects that more than 40 percent of agentic AI projects will be canceled by the end of 2027, not because the models fail, but because of escalating costs, unclear business value and inadequate risk controls. The market muddies the picture further through what Gartner calls “agent washing”: Of the thousands of vendors claiming agentic capability. CIO.com’s own reporting finds the same pattern inside enterprises — pilots that demo beautifully stall the moment they meet production, where documents vary, exceptions multiply and someone has to be accountable when an agent acts. What I see most often is that the teams that struggle aren’t the ones with the weakest platform. They’re the ones with the vaguest intent. AI amplifies ambiguity as efficiently as it amplifies capability.

The leadership response is not a bigger bake-off among platforms. Naming vendors tells you where the market is heading; it doesn’t tell you what to do. The work is to design the roles around the agents. People move up the value chain — from doing the task, to steering it, to supervising and handling the exceptions the agent can’t. Autonomy follows a ladder, not a switch: Assistant, then participant, then genuine team member, with human supervision tightening as the stakes rise. Start with a bounded use case, build the supervision muscle and only then widen the boundary. In a federal modernization program I advised, the teams that pulled ahead began with a single high-volume, rules-clear workflow, proved the audit trail and human sign-off, and widened autonomy only once the supervision held. The goal was never more agents. It is agents that belong to a team someone actually leads.

What makes it safe to let go?

The third shift is the one leaders most want to skip, and the one that now decides the other two. As agents begin to act, governance stops being paperwork and becomes the thing that lets you move. The current gap is telling: In the State of the CIO, 83 percent of leaders have or are planning cross-functional AI steering committees, but only 53 percent have any formal process for approving AI projects. Committees are easy. The boundary that lets you say “yes, act” is hard.

I recommend a reframe most leaders resist at first. Governance is not the office of “no.” Observability, evaluation, approval boundaries and rollback are precisely what let you grant more autonomy, sooner, with confidence. They are how you catch a failing agent before it becomes a headline — and, as one analysis of the Gartner forecast observes, agentic projects fail when companies grant systems access and authority before they define ownership and rollback controls. Used well, that discipline is what turns acceleration into advantage instead of avoidable damage.

This is not a distant concern. In a health plan I advise, an ungoverned action doesn’t just fail a demo: It can surface as a compliance finding, which is why governance gets attention there first. For the first time in over a decade, state CIOs have ranked AI as their number one priority, displacing the cybersecurity focus that held the top spot for twelve straight years, the very settings where autonomy is most consequential. The analyst community has reached the same conclusion: Gartner now lists evolving IT strategy, governance and operating models among the top priorities for CIOs this year, alongside operationalizing AI itself. Governance and operating-model design are no longer separate agenda items. They are the agenda.

And the pressure only builds. Gartner expects that by 2028, 15 percent of day-to-day work decisions will be made autonomously by agents, up from essentially none in 2024, with a third of enterprise applications shipping with agents inside them. Governance that feels optional today becomes load-bearing the moment agents are deciding at that scale. The leaders building the trust layer now — while the stakes are still small enough to learn on — are the ones who will be able to say yes when the stakes are not.

Guardrails aren’t what slow the car down. They’re what let you take the corner at speed.

The one shift, three ways

The three moves are facets of a single reframe. The center of gravity for IT leadership has shifted from running the technology to architecting the system around it. Decisions, workforce and trust are not three initiatives competing for budget; they are three faces of one job: Building the operating system that turns capability into results.

 Old center of gravityNew center of gravityWhat the leader must architect
DecisionsDelivering technology reliablyTurning capability into outcomesAccountable owners, a fast executive decision cadence, outcome-based metrics
WorkforceManaging tools and licensesLeading a human-plus-agent teamScoped agent roles, supervision that scales with stakes, staged autonomy
TrustControlling risk after the factEnabling speed through governanceObservability, evaluation, approval boundaries, rollback

Where should CIOs start?

None of this requires a reorganization to begin. It requires a sequence. The leaders getting ahead aren’t doing more; they’re doing these five things in order, on the bounded use cases where they can afford to learn.

1.  Name the intent. For every AI use case, write the business outcome and the person accountable for it before a line of code ships. Vague intent is the most expensive input in the system.

2.  Set the guardrails, then the autonomy. Decide what an agent may touch and what still requires a human before you widen its reach. Boundaries first, freedom second, never the reverse.

3.  Instrument for observability. If you can’t see what an agent did and why, you can’t supervise it. Build the audit trail into the work, not after the incident.

4.  Evaluate against Tuesday, not the demo. Test agents on the messy production reality, the missing field, the duplicate record, the exception — not the clean pilot. What passes in the lab rarely survives first contact with real work.

5.  Measure what the board measures. Retire activity metrics; report revenue, cost, risk, time-to-value and release confidence. If a number wouldn’t move a board conversation, it doesn’t belong on the scorecard.

Figure: Where to start: A five-step sequence.

Vipin Jain

Do these in order and autonomy compounds. Skip a step and you join the 40 percent whose agentic projects get canceled before they ever earn their keep.

The takeaway

The most common strategic mistake I see in 2026 is subtle, because it doesn’t look like a mistake. Leaders are scaling powerful new technology on an operating model built for a slower, all-human enterprise, and then blaming the technology when the returns don’t come. The failures won’t come from the models. They’ll come, as they always have, from business strategy, IT and organizational culture not being architected to move together, a pattern I have watched hold across every industry I work in, from trading floors to healthcare programs.

The good news is that this is architectable, and it is the CIO’s to architect. The leaders who will look prescient a year from now aren’t the ones who bought the most capable AI. They’re the ones who rebuilt the operating system it runs on: How their organization decides, who does the work and what makes it safe to let go. The tools will keep getting better on their own; the operating system will not: It is built, on purpose, by someone in the room. The technology was never the hard part. The leadership is. That’s the job now.

This article was made possible by our partnership with the IASA Chief Architect Forum. The CAF’s purpose is to test, challenge and support the art and science of Business Technology Architecture and its evolution over time as well as grow the influence and leadership of chief architects both inside and outside the profession. The CAF is a leadership community of the IASA, the leading non-profit professional association for business technology architects. 

Don’t automate bad workflows: Why AI should begin with redesign

Artificial intelligence has quickly become one of the biggest priorities in the executive suite. Organizations are investing heavily in new capabilities, employees are experimenting with AI every day, and technology leaders are under pressure to identify opportunities that improve productivity and reduce costs.

In many organizations, the first question is, “What can we automate?”

It sounds like the right place to start, but I believe it is the wrong question.

Too often, organizations use AI to automate workflows that were designed years ago for a very different business environment. Those workflows have accumulated unnecessary approvals, duplicate activities, manual handoffs and outdated policies over time. AI may execute those processes faster, but it does nothing to address the underlying complexity.

This challenge is not unique to my experience. In its article, The secret to successful AI-driven process redesign, Harvard Business Review explains that organizations create the greatest value when they rethink business processes before applying AI, rather than simply layering technology onto existing ways of working. Likewise, MIT Sloan’s article, How AI is reshaping workflows and redefining jobs, argues that AI delivers its biggest impact when organizations redesign how work flows across the enterprise instead of focusing only on automating individual tasks.

Those findings reinforce an important lesson for leaders. Before asking where AI belongs, ask whether the workflow itself still makes sense.

Every workflow reflects yesterday’s decisions

Most business processes were never designed from beginning to end. They evolved over many years as organizations expanded into new markets, acquired businesses, introduced new systems, responded to audits or adapted to changing regulations.

Each change made sense at the time. Collectively, they often create unnecessary complexity.

Consider a purchasing process that requires six approvals before an order can be placed. One approval may have been added after an audit. Another may have resulted from an acquisition. A third may have been introduced because one business unit wanted additional oversight. Eventually, those approvals simply become “the way we do things.”

Artificial intelligence can summarize purchase requests, route approvals automatically, notify managers and even recommend decisions. What it cannot determine on its own is whether six approvals are still necessary.

That requires leadership.

The same pattern exists throughout finance, manufacturing, supply chain, human resources, customer service and countless other business functions. Organizations often focus on making individual activities faster while overlooking opportunities to eliminate activities altogether.

This is where workflow redesign becomes essential. Instead of asking how AI can automate each step, leaders should ask which steps continue to create value, and which exist simply because they have always been part of the process.

Sometimes the greatest improvement comes from eliminating work rather than automating it.

Redesign first, automate second

The organizations creating the most business value from AI tend to approach the problem differently. Rather than starting with technology, they begin with the business outcome they want to achieve.

That outcome might be reducing order cycle time, improving forecast accuracy, increasing manufacturing throughput, accelerating product development or improving customer responsiveness. A clearly defined objective creates a much stronger foundation than simply looking for places to use AI.

Once the outcome is clear, the next step is understanding the entire workflow. Many delays occur not because individual tasks are inefficient, but because work passes through too many people, too many systems or too many approval points. Mapping the complete process often reveals unnecessary handoffs and redundant activities that can be removed before automation is introduced.

Deloitte has reached a similar conclusion in its ongoing research on enterprise AI adoption. Its latest State of Generative AI in the Enterprise report highlights that organizations generating the greatest business value are redesigning how work is performed rather than simply automating existing tasks. In other words, they view AI as an opportunity to change how work gets done instead of accelerating yesterday’s approach.

Leaders should also distinguish between administrative work and human judgment.

AI is exceptionally good at gathering information, organizing data, preparing summaries and performing repetitive tasks. People continue to provide the greatest value when decisions require experience, context, creativity, negotiation or ethical judgment.

The objective should not be to replace people. It should be to remove low-value administrative work so employees can spend more time applying their expertise where it matters most.

Standardization is equally important. When every business unit performs the same work differently, AI solutions become more difficult to implement, maintain and scale. Simplifying and standardizing workflows before introducing AI creates a stronger foundation for enterprise adoption while producing more consistent business results.

Finally, organizations should measure business outcomes instead of technology activity.

The number of AI assistants deployed or prompts submitted may indicate adoption, but they do not demonstrate business value. Leaders should instead measure improvements in cycle time, quality, customer satisfaction, operating cost, revenue growth and employee productivity. Those are the outcomes executives ultimately care about.

A simple framework for AI-enabled workflow redesign

Over the past several years, I have found it helpful to think about workflow redesign as a simple four-step sequence.

  • Simplify. Remove unnecessary work, approvals, reports and handoffs before introducing technology.
  • Standardize. Create a consistent way of working across the organization so improvements can be repeated and scaled.
  • Redesign. Build the workflow around the desired business outcome instead of existing organizational structures or legacy systems.
  • Automate. Apply AI only after the process has been simplified and redesigned.

Organizations often reverse these steps. They automate first and hope efficiency follows. In reality, automation should be the final step, not the first.

Following this sequence helps ensure AI is solving the right problem rather than making an outdated process run faster.

AI should improve work, not preserve it

One of the most valuable questions leaders can ask is surprisingly simple.

If we were designing this process today, would we build it the same way?

That question changes the conversation. It encourages people to challenge assumptions, eliminate unnecessary complexity and rethink how work should flow before technology enters the discussion.

It is also remarkably consistent with what leading researchers are finding. Harvard Business Review emphasizes that successful AI initiatives begin by improving the underlying process. MIT Sloan concludes that organizations achieve the greatest impact when they redesign workflows instead of automating isolated tasks. Deloitte’s research points to the same pattern, showing that the strongest business results come from treating AI as an opportunity to rethink operations rather than simply increase efficiency.

When independent research consistently reaches the same conclusion, it is worth paying attention.

Artificial intelligence is one of the most significant technologies organizations have adopted in decades. Its greatest value will not come from helping us execute yesterday’s workflows more quickly. It will come from allowing us to rethink how work should be done in the first place.

Leaders who redesign workflows before automating them will create simpler processes, better employee experiences and stronger business outcomes. Those who automate first may improve efficiency for a while, but they also risk embedding yesterday’s assumptions into tomorrow’s technology.

How Mercedes-Benz is scaling AI-powered business automation

At Mercedes-Benz, “Digital First” has long been more than just a theoretical concept; it’s a lived strategy, as a visit to the Digital Factory Campus in Berlin demonstrated. Now, the automaker aims to take the next step in scaling artificial intelligence: Together with the German low-code specialist n8n, the company is introducing a global platform that will enable employees to develop their own AI-supported workflows and integrate them directly into operational processes.

Unicorn startup n8n offers an AI-powered, open-source platform for workflow automation. It enables companies to efficiently manage daily processes using AI agents. Since the Berlin-based company was valued at nearly $2.4 billion in 2025, n8n has further expanded its market presence through strategic partnerships, such as with Deutsche Telekom to support small and midsize enterprises (SMEs) in areas like logistics and sales. According to Deutsche Telekom, n8n is currently the most valuable German AI company, with a valuation of €5.2 billion.

Integrating AI into everyday business

The goal at Mercedes-Benz to make data usable in seconds. To this end, AI-supported automation is to become the standard across the entire group. Behind this lies the strategy of transferring the use of AI from individual pilot projects into central processes in day-to-day business.

“We give our teams at Mercedes-Benz the opportunity to translate ideas into measurable benefits along the value chain — and to actively shape how we work in the future,” says Katrin Lehmann, who will leave her position as CIO at Mercedes-Benz on Sept. 1.

The company has already developed plenty of dedicated AI use cases. These include its own LLM suite MO360LLM, the Digital Factory chatbot ecosystem, the MO360 multi-agent system, and the AI ​​Factory as an idea factory for AI tools.

Three levels of AI competence

But the company wants more. AI and automation applications are to be directly integrated into everyday work. The goal is for employees not only to passively consume AI, but to actively shape it. The automaker distinguishes between three levels of AI competence:

  • Takers: Use AI tools in your daily workflow.
  • Makers: Design your own automated workflows using platforms like n8n.
  • Builders: As experts, they develop highly specialized software solutions.

The company-wide AI rollout received an additional boost from a hackathon. More than 1,500 employees from all business units attended the event. The goal was to independently develop ideas for AI and automation applications. Participants worked on concrete use cases to integrate AI and automation directly into their daily work.

AI

With the low-code platform n8n, teams at Mercedes-Benz worldwide can create AI-powered workflows on their own.

Mercedes-Benz

Furthermore, the hackathon served to gather creative input directly from employees and translate it into practice. The best and most powerful concepts from the competition are to be implemented within the company. The plan is to implement AI workflows in all key business areas, namely in development, production, sales, financial services, human resources, and IT.

Modular architecture

Another characteristic of AI workflows is their connection and coordination across existing IT systems to simplify complex processes. In addition to classic automation methods, new approaches using AI agents are being specifically pursued. Last but not least, the workflows are designed to support teams in solving problems faster and making decisions based on data.

As software and AI become key competitive factors in industry, Mercedes-Benz is relying on a modular and flexible technology architecture. This is where the n8n platform comes into play as part of this architecture. It functions as a low-code platform, enabling teams worldwide to create their own AI workflows. Without requiring in-depth programming knowledge, employees can thus integrate AI directly into their operational processes.

Self-hosting on-premises

At the same time, it serves to orchestrate workflows across existing IT systems. To this end, the platform connects these systems to simplify complex processes and ensure seamless integration to ensure this. Another aspect speaks in favor of the chosen solution from Mercedes-Benz’s point of view: strengthening its own digital sovereignty.

This allows n8n to be self-hosted and operated independently of the cloud. In other words, the platform runs within a secure and governance-compliant environment within the group. This way, Mercedes-Benz retains full control over critical systems, data, and work processes.

Why enterprise IT environments get more complex as companies grow

The most complicated IT environments I’ve worked in weren’t built that way on purpose. They got there through a sequence of reasonable decisions made by reasonable people under real pressure. A cloud provider added because the incumbent couldn’t hit a latency requirement. A point solution brought in because a business unit needed something fast. A managed service layered on when headcount froze. Each one made sense at the time. The problem comes later, when twenty years of reasonable decisions turn into an environment no one has fully stepped back to understand.

The ones carrying the most complexity aren’t usually the ones that made the worst decisions. They’re the ones that grew the fastest, had the most demands on their IT teams or operated in environments where buying was distributed across business units. The complexity is, in a strange way, a byproduct of success. That doesn’t make it any less expensive to carry.

How complexity builds

When I was leading one of the largest cloud and managed services practices in the country, I started keeping informal track of how enterprise clients described their own environments during the first few meetings we’d have with them. Most of the time, the same terms came up in all of our conversations: legacy, inherited, technical debt and almost always “we need that for compliance” or “that team won’t let it go.” The complexity extended well beyond the technology itself. It was tied to people, processes, compliance requirements and the realities of how the organization operated.

What I’ve observed over time is a pattern that plays out in roughly three phases, though nobody experiences it as phases when they’re living through it.

The first is the accumulation phase. This is the normal course of IT operations: you add capability as the business demands it. A SaaS application gets added because a business unit needed it fast. A new cloud region gets stood up because a regulatory requirement demanded data sovereignty. A point solution fills a gap in a platform that hasn’t been updated in three years. The organization is growing through M&A, acquiring many new contracts and vendors in the process.  None of these are bad decisions. But each one adds an integration surface, a contract, a support relationship and a line item to the budget.

The second phase is drift. This is where team members start working around the official channels (IT, Financial, Legal or Procurement) rather than through it, because the official channel is too slow or too complicated to serve their actual, or perceived, needs. Shadow IT gets a bad reputation, but in most of the organizations I’ve worked with, it’s less about rogue behavior and more about rational people solving real problems with the tools available to them. The official tools just often aren’t those tools.

The third is what I think of as the lock-in phase. By this point, the environment has so many interdependencies, some documented and many not, that making a significant change feels genuinely risky. The technical debt is high with perceived high costs to modernize.  Building the case to move requires more stakeholder engagement and analysis than simply comparing the existing approach to a more modern one.  Every potential simplification comes with a list of downstream impacts that nobody is fully confident about, so the environment stays complicated and the maintenance burden grows.

The Flexera 2024 State of the Cloud Report found that organizations waste an average of 28% of their cloud spend, a number that has held remarkably consistent across multiple years of the same study. In my experience, a meaningful portion of that waste isn’t irresponsible purchasing. It’s the cost of maintaining overlapping capabilities that were acquired at different points in time, for reasons that made sense then and are hard to untangle now.

Why simplification efforts stall

I’ve watched a lot of simplification initiatives get announced and then shelved. The reasons are usually practical, not political, though the politics can make the practical harder.

One of the biggest problems for teams is that nobody has a complete view of their environment. I’ve seen consolidation initiatives take longer than anticipated because every time a team thought they had their dependencies mapped out they found another application or workflow that wasn’t included in the initial discovery. Nobody did anything wrong in those instances. The environment had just evolved faster than the documentation, which is what happens when you’re adding capability under pressure for years at a time.

This is more common than organizations want to admit. Gartner has found that shadow IT accounts for 30 to 40% of IT spending in large enterprises, composed of tools acquired by business units, spun up for a project and never decommissioned, or inherited through an acquisition that never got fully integrated. When you’re trying to simplify something you can’t fully see, the margin for error is wide.

The second problem is that simplification affects real people inside the business. Somewhere in the organization, there is someone who is dependent upon the application or process that you want to eliminate. This could be either a report or workflow that is tied to a database that was supposed to be shut down two years ago, or it could be a team that has customized a platform in such a way that they will not be able to migrate when the time comes. Real simplification requires working through those dependencies carefully, which takes time and coordination that most IT teams don’t have in abundance while also keeping the lights on.

The third problem is that the business case is hard to make in advance. The cost is real, but it does not always show up in one clean line item. It shows up as slower incident response times, higher vendor management overhead, more onboarding time for new IT staff and missed windows to adopt newer capabilities. All of these elements highlight the risk that grows as these elements continue unaddressed.  Those costs are hard to aggregate into a number that justifies a multi-quarter consolidation project, so the project often doesn’t get funded until something breaks badly enough to force the issue.

What works in the real world

I want to be careful here, because there’s a version of this advice that sounds simple and isn’t. “Just rationalize your vendor portfolio” is easy to say. Doing it in a way that doesn’t create new problems requires discipline.

The first thing that’s worked consistently in the environments I’ve been close to is starting with the contract layer, not the technical layer. Most organizations have a better view of what they’re paying for than what they’re running. A thorough contract and spend audit will surface redundancies faster than a technical architecture review, because the money trail is usually cleaner than the configuration trail. It won’t give you the full picture, but it gives you a starting point grounded in something real.

The second is building the map before you touch anything. Organizations that have executed well on simplicity spent significant amounts of time, often months, conducting an inventory of the current state before determining what to consolidate. This is unglamorous work and will not show up on the board update as a major accomplishment, but that’s precisely what distinguishes successful consolidation from stalled or failed consolidation.

The third is staging the work around business cycles rather than IT timelines. I’ve seen technically sound simplification projects fail because they were scheduled without regard to when the business could absorb the risk. Migrations during quarterly close processes, architecture changes during peak retail seasons: these are the kinds of decisions that undermine confidence in IT’s judgment, regardless of the technical merits.

The goal isn’t a simple environment for its own sake. A mature enterprise IT environment is going to carry some complexity, because the business it supports is complex. What you want is an environment where every tool, vendor, platform and contract has a clear purpose. You know who owns it, what it costs and what risk it creates. Most of the organizations I’ve worked with aren’t that far from that state. They just need to stop adding before they can start subtracting.

Your AI hiring tool isn’t an HR problem. It’s a security one

For years, applicant tracking systems and recruiting platforms were treated as HR technology: Important for workflow, efficiency, compliance and candidate experience, but rarely viewed as core security infrastructure. That assumption no longer holds. Once AI begins reading resumes, scoring candidates, conducting interviews, ranking applicants and influencing who moves forward, the hiring platform stops being a passive system of record. It becomes a decision system.

And any system that accepts public input, processes sensitive data and influences business decisions belongs inside the security conversation.

I learned this during an AI hiring platform rollout that never made it to production. The vendor was established, the product had a strong market reputation and the AI feature looked attractive: Upload a resume, compare it to a job description and return a neat percentage match. For recruiters, it promised speed. For executives, it promised modernization.

Before moving real candidate data into the system, I tested it with synthetic resumes. One weak resume came back with a surprisingly strong match. The reason was not hidden in the candidate’s experience. It was hidden in the text. The resume contained language instructing the AI to treat the candidate as an excellent fit, and the system appeared to follow that instruction instead of evaluating the resume on merit.

That changed the question from “Does the tool improve productivity?” to “Can the person being evaluated influence the evaluation itself?”

That is a security question.

The trust boundary has moved

CIOs do not need to become recruiting experts. They only need to look at the mechanics.

An anonymous user submits content into an enterprise system. That content is processed by software. The software then produces an output that can influence a business decision. In every other environment, security teams know what to call that: untrusted input crossing a trust boundary.

The difference is that in hiring, the input looks harmless. It is a resume, a cover letter, a chatbot reply or a spoken answer in an AI-led interview. But once AI reads that content and treats it as instruction, the harmless-looking input becomes part of the system’s control surface.

That is why prompt injection matters in hiring. It is not just an AI oddity or a model behavior issue. It is the same category of failure enterprises have spent decades trying to prevent: User-controlled input changing what the system does. OWASP lists prompt injection as the first risk in its Top 10 for LLM applications, describing it as a case where user prompts alter a model’s behavior or output in unintended ways.

In hiring, the implication is direct: A candidate may be able to manipulate the score, ranking or interview assessment that determines whether a human ever sees them.

The business impact is not theoretical

The obvious risk is that an unqualified candidate moves forward. But the impact is broader.

First, decision quality degrades. Hiring teams adopt AI scoring because they believe it improves signal. If the score can be manipulated, the business is not gaining signal; it is gaining false confidence. Recruiters may spend time on candidates who gamed the system while stronger candidates are buried lower in the queue. A tool bought to reduce friction can quietly create more of it.

Second, cost increases under the appearance of efficiency. Every false positive consumes recruiter time, hiring-manager attention, interview slots and opportunity cost. A small weakness in screening integrity can become a measurable operational drag across open roles.

Third, trust suffers. Candidates already question whether AI hiring tools are fair, explainable or accurate. If it becomes clear that a screening system can be manipulated by hidden instructions or verbal prompting, the issue is no longer just security. It becomes reputational. Strong candidates may lose confidence in the process, and employers may have to defend decisions made by systems they did not fully understand.

Fourth, sensitive data exposure becomes harder to contain. Recruiting systems hold names, addresses, work histories, education histories, compensation details, work authorization information and sometimes accommodation or demographic data. NIST guidance on personally identifiable information includes employment information as linkable personal data that must be protected from inappropriate access, use and disclosure. Yet hiring platforms often receive less security scrutiny than systems holding customer or financial data.

That mismatch is dangerous: High-value data, public-facing workflows and increasing automation.

The 2025 McHire incident should have made this impossible to ignore. Researchers reported that weaknesses in McDonald’s AI hiring platform, including default credentials and an access-control flaw, exposed applicant data at large scale before the issue was patched. The lesson for CIOs is not merely that a weak password was used. The lesson is that AI hiring systems can ship with basic, preventable security failures while still being treated as HR tools rather than enterprise risk surfaces.

Vendor reputation does not transfer to every AI feature

One reason this risk slips through is that buyers often trust the platform brand. Mature vendors may have strong security programs, enterprise customers, compliance documentation and procurement-friendly answers.

But AI features can change the architecture of risk.

A platform that was safe as a workflow tool may behave very differently once it adds resume scoring, interview grading, chatbot screening or automated ranking. The new feature may introduce new inputs, new model behavior, new data flows, new third-party dependencies and new decision points. In practical terms, the attack surface has changed.

CIOs should not allow AI features to inherit trust automatically from the legacy platform around them. When a vendor adds AI, the enterprise should reassess the feature as if it were a new product. That does not mean slowing innovation for bureaucracy. It means AI-enabled decision-making carries different failure modes from ordinary workflow automation.

The ownership gap is the real vulnerability

The biggest risk may not be the model. It may be the ownership gap.

Talent acquisition may buy the tool. HR operations may configure it. The vendor may guide implementation. Procurement and legal may approve the contract. But who owns the security of the candidate-facing AI layer?

In many organizations, the honest answer is unclear.

That ambiguity is where risk grows. Recruiting technology sits at the intersection of public input, sensitive data, third-party software, automated decision support and brand trust. That is exactly the kind of environment that needs named security ownership, asset inventory, vendor review, access-control testing, logging and incident-response planning.

If the hiring stack is not in the security inventory, the organization is already making an assumption it may later regret.

What CIOs should require now

The fix is not exotic. It is applying existing security discipline to a surface that has been underestimated.

Treat every candidate submission as untrusted input. Resumes, cover letters, chatbot responses, interview transcripts and spoken answers should be handled as attacker-controllable content. If AI processes it, the system must separate content from instruction.

Reassess vendors when AI features are introduced. A prior security review should not be treated as permanent approval for new AI capabilities. Ask what changed in the architecture, what data the model sees, what actions it can influence and how manipulation attempts are detected.

Ask AI-specific questions before signing. Can candidate-provided content alter scoring? Are hidden instructions filtered or ignored? Is there human review before AI output influences a decision? Can the vendor produce testing evidence for prompt injection, access control and data exposure risks?

Assign ownership. HR can own the process, but security must own the risk model. AI hiring systems should be included in third-party risk management, application security reviews, access governance, monitoring and incident response planning.

Measure business impact, not just AI adoption. The goal is not to say the recruiting function uses AI. The goal is to improve hiring speed, quality, fairness and cost without creating new risk. If the system cannot protect decision integrity, the business case is weaker than it appears.

The hiring platform is now part of the enterprise attack surface

AI has turned the careers page into more than a front door for applicants. It is now a public input channel feeding systems that store sensitive data and influence workforce decisions.

That makes it a CIO concern.

The next failure in AI hiring may not look like a traditional breach at first. It may look like bad rankings, manipulated scores, unexplainable decisions, wasted recruiter time or a candidate process no one trusts. But underneath those symptoms is a familiar security problem: A system trusted input it should have treated as hostile.

Enterprises have hardened payment systems, customer portals, APIs and employee applications around that lesson. Hiring deserves the same treatment.

AI hiring is not just an HR transformation. It is a security boundary. And it is time CIOs treated it like one.

The 5 stages of AI adoption maturity: Where businesses create real value

Most enterprises are rushing toward autonomous AI. They shouldn’t. Autonomy you haven’t earned doesn’t speed you up. In fact, it slows you down.

Here’s what I’ve moved our organization toward: a five-stage set of AI adoption maturity benchmarks. It’s a practical framework for understanding where employee development, decision-making and business value intersect. Each stage provides value for your organization. Some roles and functions may only ever reach Stage 1 or 2, while others should be fast-tracked to Stage 5. By understanding this progression, leadership can stop viewing AI as a tool for task delegation and treat it as a catalyst for developing stronger, more decisive and more valuable teams.

Stage 1: Research assistance

You hand people a premium ChatGPT account. Employees stop Googling and start prompting. Their experience improves: no ads, paragraph-form answers instead of blue links. But the underlying dynamic hasn’t changed. Output quality depends on input quality. A vague Google search returns a mess of links. A vague ChatGPT prompt returns a well-formatted mess of paragraphs. If your team didn’t know how to ask a precise question before, they still don’t.
           
The real danger at Stage 1 isn’t the bad answers – it’s the confident-sounding ones. A hallucinated statistic arrives in the same calm, authoritative prose as an accurate one. Teams that don’t verify sources in Google don’t suddenly fact-check ChatGPT. Before moving to Stage 2, your team needs to develop the instinct to ask, “How do I know this is true?”

Stage 2: Task assistance

The next stage uses AI tools to complete tasks. It starts simply: “I need to write this email,” or “Make a spreadsheet to track open items.”

The average employee takes what AI produces and passes it off without revision. At best, their efforts pass muster, with only a dash of workslop. At worst, the flood of unchecked AI outputs creates rework for teammates and clients.

Another employee further along in Stage 2 may augment what AI produces. That impulse serves them well. But if they default to editing AI output rather than dictating the rules for what AI should produce, they can easily spend more time editing AI’s work than creating work from scratch.

For employees whose work will largely remain in Stage 2, the focus should be on writing more precise prompts. The instinct to edit AI output isn’t wrong. The problem arises when the prompt is a rough starting point rather than a detailed spec. AI cares that your instructions are clear, specific and unambiguous. Get the spec right up front.

Stage 3: Workflow integration

My daughter’s class recently had an assignment: write a paper on the causes of the Civil War.

Her teacher knew what was going to happen. Every 11-year-old would go home and use ChatGPT to write a five-paragraph essay. So, she changed the exercise. The class generated and printed out the essay. Then, the teacher explained how to annotate, how to ask follow-up questions and how to revise in ChatGPT using the marked-up draft.

The same three-step sequence — assemble context, build the prompt, edit hard — applies when someone writes a post-mortem. The temptation is to skip straight to the draft. Pull the incident data, ask Gemini for a timeline and root cause analysis, clean it up, get a quick peer review and send it.

An engineer working at Stage 3 does what the teacher did. First, they assemble context: the Slack thread where someone flagged the anomaly two hours before the alert fired, the Jira ticket, the gap in monitoring that nobody documented. Then they build a prompt that reflects the full context and generate a draft. Now the red pen comes out: push back on the root cause analysis, add the institutional context Gemini couldn’t know, tighten the remediation steps until they’re actionable.

The result is a better document — and an engineer who understands what failed and builds a better repeatable process. Saving time on a first draft is a fine side effect. The goal is to produce a final draft that’s worthy of review.

Stage 4: Guided automation

The fourth stage is where collaboration becomes self-sustaining. You’re no longer asking AI to help you do a task. You’re asking it to run the task and surface the decisions that require your judgment.

My LinkedIn workflow is a good example of what this looks like in practice.

A couple of years ago, I would read an article, develop a point of view, write two or three paragraphs and publish. Not bad, but dependent on me having the time and cognitive bandwidth.

The friction was the 15 decisions that came before drafting: Which angle is worth pursuing? Does this use my voice? Have I said this before?

So, I started researching my patterns. First, I fed Claude my prior LinkedIn posts and prompted it to analyze my tone, sentence patterns and structural habits. I didn’t ask it to “describe my voice” – that gets you a paragraph of flattering generalities. This analysis became the base layer of the tool.

Then I added a second layer: LinkedIn-specific rules and AI writing patterns to avoid. That context got embedded alongside the voice analysis.

Now the workflow runs like this. I click a link, save the article, highlight and annotate the sections that interest me. My Claude Managed Agent picks up the annotation, infers what I found worth engaging with and writes four drafts with meaningfully different angles on the source material. It compares each draft against my post history and proposes two. I read the proposals, pick one, edit and authorize publication with Buffer.

The automation didn’t remove my judgment from the process. It freed me from work that didn’t depend on judgment. Now I do the work that matters: deciding what to say, identifying patterns and sharing my point of view.

That shift in what I’m accountable for is where the ROI changes. The value isn’t in the time saved on any single post. It’s that the workflow no longer depends on me having the bandwidth to start from zero. The capacity was always there; the system makes it consistent and repeatable.

Stage 5: Full automation

The most advanced stage of maturity is when the system largely runs on its own. You’re no longer managing step-by-step actions; you’re defining goals, setting guardrails and measuring outcomes.

We have one running in our engineering org right now. When a ticket gets escalated from our support team to engineering, the agent triages it and routes it to the team responsible for the fix. When an engineering manager reassigns the ticket – because the routing was wrong – the agent picks up that correction, feeds it back into its prompt tooling and updates its model of who owns what. We’re now extending it further: the agent is learning which parts of the codebase need to change and which engineers are likely to own the fix.

There’s a critical catch: this stage only works if you’ve earned your way there. We learned this firsthand. When we first rolled out the routing agent, we used a static map of application areas to engineering teams and assumed that was enough. It wasn’t. We couldn’t reliably distinguish front-end bugs from back-end ones, so the front-end team kept getting tickets caused by a misbehaving API. Features were split between teams in ways the map didn’t capture — one team owned exports, another owned reports. Before the routing could work, the knowledge had to exist somewhere it could be used. An autonomous system is only as good as the foundation beneath it – the clarity of your workflows, the health of your data, the alignment of your teams. Deploy an autonomous agent into a broken process and you get bad results at scale. You cannot safely delegate what you don’t fully understand.

This is why racing straight to Stage 5 often fails. You need to know what “good” output looks like (Stages 2 and 3) and how to orchestrate the pieces (Stage 4) before you can confidently take your hands off the wheel.

Where business value emerges

The evolution from a premium search engine to an autonomous system is an organizational challenge, not a technology one. Realizing the value of AI is determined not by the sophistication of the underlying model, but by the maturity of the team wielding it.

The practical move isn’t to audit your whole organization’s AI readiness. Start with one workflow. Push it one stage higher. Measure what changes. That’s how you find out if this matters in your specific context – not in theory, but in the work your team actually does.

SAP dodges German antitrust investigation over data extraction

SAP is not unfairly preventing enterprises from extracting their data from its systems for use with competitors’ applications, the German Federal Cartel Office (Bundeskartellamt) concluded Thursday after a preliminary investigation.

The Bundeskartellamt does not currently intend to initiate abuse proceedings against SAP, although it will continue to monitor developments in what it views as a dynamic market, it said in a news release.

It launched its investigation into SAP’s practices following complaints by software companies including Celonis, a developer of process mining tools, alleging that SAP makes it difficult for customers and third parties to access data from its ERP systems and favors its own Signavio process mining tool.

“Companies must generally also be able to use their own data in third-party applications. With large software platforms, in particular, non-discriminatory access to data is crucial to effective competition,” said Bundeskartellamt President Andreas Mundt. “Our preliminary investigation has found that there are currently sufficient data extraction options available and that there have so far been no indications of exclusionary practices that may be relevant under competition law.”

SAP changed its policies on accessing data held in its applications via APIs in April, prompting customer pushback.

But, said Mundt, the Bundeskartellamt found that despite the API policy change, data extraction options that were previously permissible are still available.

Data extraction is possible

SAP welcomed the Bundeskartellamt decision, saying that “as the authority states, SAP customers and partners have sufficient and permissible technical options to extract data from SAP systems and use it in solutions from other providers. The SAP API Policy does not restrict these capabilities.”

Celonis also issued a statement, noting that the Bundeskartellamt ruling underlined the continued importance of unrestricted data access, and warning, “The decision is based on the key premise that data extraction for software from providers such as Celonis will remain possible even under SAP’s new API policy — a premise that SAP has been unwilling to confirm to date.”

The Celonis statement continued, “We remain steadfast in our conviction that company data belongs entirely to the customers who generate it. No provider should restrict a company’s right to extract its own information or prevent users from working with third-party providers such as Celonis that offer added value to customers.”

Celonis is also attacking SAP’s policies on data extraction in court in California. It filed a complaint in March 2025 alleging that SAP was leveraging its software to “prevent SAP customers from sharing their own data with third-party providers, including Celonis, without paying prohibitively expensive fees.” The judge dismissed some of the claims in that case, leaving three to be tested in a trial then scheduled for December 2026. Celonis has since amended its complaint to include 10 claims, and the trial has been rescheduled for 2027, the company said.

“Our litigation continues to uncover evidence of SAP’s unlawful behavior, including anticompetitive conduct and theft of intellectual property, and we are confident in the evidence that we will present at trial,” Celonis said following the German authority’s decision.

The Bundeskartellamt’s failure to find sufficient evidence to open a ‘formal abuse of dominance proceeding’ is a small win for SAP, said Scott Bickley, advisory fellow at Info-Tech Research, but “CIOs should not mistake it for a validation of SAP’s data access model.”

Although SAP recognizes customers’ right to decide they use their data, it does not make it easy for them to do so, he said. “CIOs may technically retain vendor choice but be faced with expensive replication architectures, API rate and volume restrictions, additional platform costs, performance lags and data migration costs, all with a dependency on an SAP-approved technical pattern, which can be a moving target.”

Data ownership as a procurement issue

Justin Greis, CEO of consulting firm Acceligence, sees the decision as an instructive one for enterprise CIOs.

“This isn’t a reason to stop asking hard questions of your ERP vendor. Whether it’s SAP, Oracle, Microsoft, Salesforce, or anyone else, enterprises should continue to evaluate how easy it is to access their own operational data, integrate third-party applications, and migrate workloads if business priorities change. Those questions are becoming strategic procurement issues, not just technical ones,” Greis said.

CIOs should consider data portability early in the procurement process, said Kaan Dincer, CEO of data migration vendor Settle: “Negotiate export rights, API access on reasonable terms, and documentation of the data model before signing and test a real extraction while the vendor still wants your renewal. The cost of your eventual exit is set on the day you implement, not the day you leave. ERP data now feeds analytics and automation outside the system of record, so access friction that used to be an IT annoyance is becoming a strategy constraint.”

In the SAP case, he said, “the regulator answered a narrow legal question, not the operational one. Declining to open proceedings means the friction was not shown to be anticompetitive. It does not mean the friction is not real. The Bundeskartellamt’s own findings acknowledge that extracting large data volumes is technically demanding and it said explicitly that it will keep watching as access mechanisms and license models evolve. That is not a clean bill of health. It is a decision to hold fire.”

Srinivasulu Reddy Battu, a senior software engineer with cloud vendor ZT Systems, said the big takeaway is the difference between difficult and impossible. SAP’s argument is that the data migration outside of its environment is possible, but Battu said it can be a time-consuming and expensive process.

“When the ruling says ‘various permissible and viable options’ exist, that’s technically true, but it glosses over how much expertise it actually takes to use them,” Battu said. “CIOs should still watch how process mining gets packaged in their contracts. If Signavio comes included by default, teams will naturally start using it and that quietly reduces your negotiating power with other vendors over time. This isn’t just about SAP: Oracle, Microsoft, every major ERP vendor sits on a massive amount of your business data. If any of them decided to tighten their API policies tomorrow, most companies would be scrambling.”

Control is the feature: The real AI risk is the lawyer you told not to use it

The real risk in legal AI is not the lawyer who studies these tools. It is the associate who quietly pastes a client’s contract into a free chatbot at 11 p.m. because a brief is due and nobody gave them anything better.

That lawyer exists at your firm right now. Survey after survey confirms it, and common sense confirms it faster. The tools are free, fast and remarkably good at exactly the drudgery that fills a litigator’s week. Telling people not to use them is like telling people not to use search engines. The use does not stop. It just goes underground, where there is no policy, no supervision and no control over where the client’s information lands.

That is the problem worth writing about. Not whether AI will replace lawyers. Whether lawyers will manage it or pretend it away.

The false choice

Most firms have picked one of two postures, and both are less careful than they feel.

The first is prohibition. Ban the tools, circulate a stern memo, move on. This feels responsible. It is not. Prohibition does nothing to the demand side. The work is still crushing, the tools are still one browser tab away and the memo guarantees that when someone uses them anyway, and someone will, they will not tell you. You have not eliminated the risk. You have blinded yourself to it.

The second is procurement. Buy an enterprise legal AI platform, sign the vendor’s security addendum and trust the marketing. This feels responsible too. But most lawyers who buy these platforms cannot tell you where the data goes, what the vendor retains, whether client documents train someone else’s model or what happens inside the black box between the upload and the answer. You have not exercised judgment. You have outsourced it, along with your client’s data flow, to a sales team.

Neither posture asks the lawyer to actually understand the technology. That is the tell. We would never let an associate cite a case they have not read. Yet firms routinely adopt, or ban, tools that nobody in the building has taken apart.

The third path

There is a third posture, and a small but growing movement of lawyers has already taken it. Some call them “legal quants,” a borrowed term from finance, where quantitative analysts stopped waiting for vendors and built their own instruments. The legal version is a lawyer who learns enough about how these systems work to build careful, narrow, controlled tools for their own practice, rather than banning the technology or buying whatever is on offer.

This is not hypothetical, and it is not confined to coastal tech firms. One of my law partners went through an intensive legal-tech residency and came back with a working tool he built himself, one that handles a defined slice of our document work, runs under conditions he set and keeps client material inside boundaries he can actually describe. I am deliberately light on the details, because the program matters less than the posture. He did not buy a promise. He built an instrument, and he knows exactly what it does and does not do.

That knowledge is the whole point.

Where this actually bites

My practice is commercial litigation in Georgia. Contract disputes, business torts, healthcare litigation. It is document-heavy in the way that grinds people down: thousand-page productions, deposition transcripts, discovery responses that have to be checked against each other line by line.

The judgment in that work lives in the seams. Which limitation-of-liability clause actually controls. Which answer to Interrogatory 14 contradicts what the witness said on page 212. Whether a document is privileged or merely embarrassing. AI is genuinely useful at surfacing those seams faster, organizing, comparing, flagging. It is genuinely dangerous when it is trusted to resolve them.

A controlled tool respects that line by design. It surfaces, and the lawyer decides. An off-the-shelf chatbot respects no line at all, because nobody drew one.

Confidentiality cuts the other way

Here is what the hand-wringing pieces get backwards. Confidentiality is not the reason to avoid understanding these tools. It’s the reason you must.

A lawyer who understands how a language model handles information is far better positioned to protect client confidences than one who does not. They know what gets transmitted, what gets retained, what gets logged and where inference actually runs. They can read a vendor’s data-handling terms and know which questions to ask. They can configure a tool so that client documents never leave a controlled environment. They can spot the difference between real security architecture and a badge on a website.

The lawyer who “protects confidentiality” by refusing to learn cannot do any of that. Their protection is a memo. The other’s is control. Under Rule 1.6 and our duty of technological competence, control is what the obligation actually demands.

The honest limits

None of this replaces judgment, and nothing I have described runs unsupervised. Every output gets reviewed by a lawyer who answers for it, to the client, to the court, to the bar. These systems draft, sort, compare and flag. They do not sign. The hallucinated-citation sanctions cases all share one fact pattern: a lawyer who skipped the review. The tool did not fail. The posture did.

What clients are already asking

Clients are already asking how their lawyers use AI, and the answers they deserve are specific ones. What we use, what we built, where their information goes and who checks the work. Firms that can answer will earn trust. Firms whose real answer is “we banned it, and we hope everyone complied” will not.

The profession does not need more hype, and it does not need more fear. It needs lawyers willing to take these systems apart, keep a human in charge and build tools worthy of the confidences we hold. Some of us have started. The rest should catch up.

This article is published as part of the Foundry Expert Contributor Network.
Want to join?

Business transformation needs a true economic approach, not guesswork

Thom Hornback has spent 20 years inside large organizations running the kinds of complex implementation projects that transformation programs are built around. He is a Six Sigma black belt, a project management professional and someone who thinks primarily about the people side of change. When he first encountered a process modeling approach that evaluated business processes through the lens of the information they generate, consume and destroy, he tried to get it approved at GE Power, where he ran professional services for one of its major platforms. But his request didn’t get much airtime. “The value of information is not a concept that organizations tend to track or value for that matter,” he says.

Most transformation frameworks are built to fixate on cost: what a process consumes, where it slows, which steps can be eliminated or automated. Those are legitimate questions, and the methodologies like business process modeling (BPM) or activity-based costing (ABC) that answer them are genuinely useful. What they are not designed to ask is what a process is worth, such as how revenue is influenced, how risk is absorbed, what options are preserved and especially anything about the information that is generated. Efficiency captures what a process no longer costs. It says nothing about what it produces of economic value. A generation of transformation investment has produced returns that reflect exactly that tilt. However, today’s organizational and business complexity demands more than simply squeeze-the-denominator type transformations, especially if executives hope to innovate.

The cost of not knowing what things cost

Process transformation has been a fixture on the enterprise agenda for 30 years, moving through successive waves: business process reengineering, Lean and Six Sigma, robotic process automation and now agentic AI. Each arrived with genuine methodology, a consulting infrastructure and case studies demonstrating impressive operational results.

However, the actual returns paint a different picture. Recent Bain research found that 88% of business transformations fail to achieve their original ambitions. Organizations that have become demonstrably better at executing processes have not, in most cases, materially benefited from it.

The explanation most commonly offered is implementation failure: change management gaps, technology underperformance, organizational resistance, insufficient executive sponsorship. These are real factors, but ones that perhaps take a back seat to how and by whom transformation programs define success. In other research, Gartner reported that 67% of CFOs believe their digital spending is underperforming against expected outcomes, primarily due to poor CFO-CIO alignment, and that only 30% of CFO-CIO relationships can be described as strong digital partnerships. That gap isn’t a communication problem. Rather, it speaks to the need for a better framework for defining what a process is worth before the program launches.

The optimization trap

Efficiency measures output per unit of input, making it indifferent to whether the output is worth producing. A process that delivers the wrong product faster is more efficient and less valuable. A process that removes friction from a workflow that shouldn’t exist in the first place is one that optimizes a waste. The biased logic of efficiency improvement assumes the process being improved is already doing something worth doing, at a scale worth doing it, in a configuration that makes economic sense. It is an assumption that transformation programs rarely examine, which is why so many produce measurable operational gains and negligible economic returns — returns that executives can’t even agree upon.

Most organizations carry, embedded in their process portfolios, activities that consume significant resources and generate little economic value: approval layers whose risk rationale has not been reviewed in years, reporting cycles producing outputs nobody uses, coordination processes that exist because two functions never aligned on accountabilities, quality checks duplicated at successive handoffs because no one established where accountable review actually sits. Efficiency analysis cannot identify these as candidates for elimination because it does not ask what a process is worth, only what it costs to run and how fast it can run faster. The question of value never enters the frame.

Economic process modeling: The complete picture

The knowledge exists inside every organization. It particularly surfaces in interviews with the people who actually do the work. There is a gap between what managers believe is happening and what their teams experience daily. “Even managers don’t really have a good grasp of where their people feel like they’re wasting their time,” Hornback observes.

What conventional process analysis treats as noise, a concept called economic process modeling (EPM) treats as signal. EPM decomposes processes into their constituent components and evaluates each across five economic dimensions:

  • Revenue contribution: Which process steps directly influence customer retention, expansion or acquisition
  • Cost and friction: Where resources are consumed relative to the value being generated
  • Risk exposure: Which components create liability, compliance exposure or operational vulnerability
  • Option value: Which steps preserve or foreclose future strategic choices
  • Information value: What decision-relevant data assets the process generates, degrades or destroys

The output is a map of which components and subcomponents of a process generate economic value, which consume it and which destroy assets (relationship capital, information yield, decision quality) that operational metrics never surface. More than a map, EPM aggregates these economic signals up and down throughout a process. And transformation programs built in this way tend to pursue a different and broader objective than programs built on a process inventory scored only by effort and volume.

The asset nobody’s counting

The revenue contribution dimension of EPM alone tends to reorder transformation priorities significantly. Activities that appear administratively unremarkable often carry direct influence over whether customers renew, expand or defect and the economic consequence of improving them dwarfs the savings available from automating the most labor-intensive steps in the portfolio. A mid-market account review process that is technically efficient but experientially perfunctory may contribute to churn at a rate that costs the organization far more annually than the entire efficiency program is designed to recover, though the two numbers are rarely placed next to each other. Fixing the efficiency metrics of that process while leaving its economic function unexamined is the organizational equivalent of polishing a car with a failing engine.

The information value dimension surfaces a different category of opportunity altogether, one that treats data as an asset, not a byproduct. EPM also identifies process components that could generate decision-relevant or monetizable data assets but do not. Organizations routinely automate data-generating steps in ways that improve throughput while destroying signal fidelity. This is a tradeoff that rarely appears in the metrics used to declare the transformation a success, and that compounds across every subsequent decision that depended on the signal.

The economic case, made up front

EPM also changes how transformation programs compete internally for resources. Teams that enter capital allocation reviews with an initiative grounded in economic modeling (i.e., cost basis, throughput projections, ROI and information value contributions mapped to specific process components) don’t merely argue more credibly for budget. They come to the table with analyses that overshadow efficiency-only business cases. This kind of rounded economic business case can accelerate the investment decision, not just the argument for it.

Still, cost reduction is a legitimate objective, and processes that are both economically valuable and operationally wasteful are obvious candidates for improvement on both dimensions. The argument is merely against efficiency as the primary frame for transformation decisions because the frame determines what gets measured, what gets prioritized and what counts as success. Hornback observes that organizations have been comfortable with efficiency metrics precisely because efficiency is the low-hanging fruit that doesn’t force any accountability for driving up value. As a result, businesses that have spent decades optimizing process mechanics while leaving full-flavored process economics unexamined have been working with the most popular tools, but not the sharpest. EPM, on the other hand, cuts across multiple value dimensions in a way that business transformation programs have long needed.

This article is published as part of the Foundry Expert Contributor Network.
Want to join?

❌