Visualização normal

Ontem — 8 de Setembro de 2026Stream principal
  • ✇Security | CIO
  • Harnessing unleashed AI agents
    In Northeast Greenland, where temperatures can plummet to -40°F, security officials rely on the Sirius Dog Sled Patrol, led by well-trained canines that guard the sprawling, weather-beaten coastline – tundra territory where snowmobiles commonly fail. Tethered together with the right harness that efficiently channels their collective energy toward a shared mission, the sled dogs are more than up to the challenge. But left to run free without the leashes and human guidance,
     

Harnessing unleashed AI agents

8 de Setembro de 2026, 08:00

In Northeast Greenland, where temperatures can plummet to -40°F, security officials rely on the Sirius Dog Sled Patrol, led by well-trained canines that guard the sprawling, weather-beaten coastline – tundra territory where snowmobiles commonly fail. Tethered together with the right harness that efficiently channels their collective energy toward a shared mission, the sled dogs are more than up to the challenge. But left to run free without the leashes and human guidance, they naturally become a pack of wild animals bent on following their instincts.  

Enterprises relying on AI could learn a thing or two from this scenario. In recent years, organizations have depended on copilots and chat-based assistance designed to answer questions or summarize information. These systems have advanced to include autonomous agents increasingly capable of executing workflows, accessing tools, interacting with software and making decisions with limited human oversight. AI has been enabled to serve as a true workforce partner.

It’s an evolution that promises significant productivity gains but requires a more advanced foundation. Even the smartest agents need clear directives and the right connections to successfully maneuver sophisticated enterprise systems and maximize their potential.

This concept has been coming up pretty frequently in conversations I’ve been having with tech leaders lately. When I was in Nashville not long ago for the Insurance Innovators USA conference, and later over a few cocktails with former colleagues near San Francisco, I quickly tuned into a growing trend. Instead of talking about predictable topics like which foundation model was the most intelligent, the conversation veered toward a more thought-provoking challenge: How do we connect and amplify these increasingly autonomous AI systems to yield the greatest value more safely?

The answer to that question represents enterprise AI’s next major opportunity. Organizations are now realizing that capability and raw intelligence are only the beginning:  Building the infrastructure that enables agents to perform dependably at scale matters even more.

Operationalizing intelligence

Autonomous agents are a different animal from traditional AI assistants. That’s because they don’t simply generate text; they take resonant action. A self-directed AI agent can, for instance, update customer records, trigger software workflows, initiate financial transactions and coordinate with other AI agents. These proficiencies significantly up their value and turn them into vibrant operational resources. But these assets require a structured environment to succeed. An agentic system can have the necessary tools but lack the right controls to navigate compliance and privacy rules. To tap their full potential, the architecture that effectively directs their actions must exist.

Traditional guardrails weren’t designed for this kind of autonomy. Prompt filtering, simple permissions and basic access controls do the job for conversational AI. But they don’t cut it when it comes to enabling software that makes decisions and interacts with enterprise systems independently. That requires a new level of orchestration.

Enterprises need a standardized control layer for agent behavior, regardless of which underlying model powers them. We have to recognize that intelligence by itself isn’t enough – control is just as important.

Which brings us back to those trusty sled dogs. Think of each dog as a large language model (LLM) task. We often run several LLM tasks within a harness, often involving different models, comparable to a sled team. Just as each dog is positioned for what it does best, from lead dog to wheel dog, a “mixture of experts” delegates each part of the problem to the LLM task best suited to handle it. Without a harness guiding their powerful capabilities for a common purpose and enabling better performance, those LLM tasks, like the dogs, can’t effectively pull the sled. An AI model needs this same type of surrounding governance to reliably perform enterprise work and accomplish its objectives.

An agent harness provides the necessary infrastructure to contain and channel agent capability safely. It securely defines permissions and access boundaries, determines rigid tool usage limitations, manages workflow sequencing, human approval workflows and approval logic and creates audit and observability trails. The LLMs provide raw power, but the harness enables the coordination and audit trails needed to transform AI intelligence into reliable operations.

AI tools are progressing into increasingly dynamic autonomous agents. It’s encouraging to see that organizations have mostly moved beyond experimentation and are finally incorporating AI into production workflows that impact customers and revenue. But that means regulators are paying closer attention, particularly to organizations in insurance, financial services and other highly regulated industries. The architecture facilitating these agents has to be resilient enough to both comply with requirements and foster speedy innovation.

Autonomous AI agents signal a new era of speed and capability, creating exciting prospects for executive leaders ready to scale operations. To take advantage of this momentum, they should ensure that early deployments have strategic guardrails and a clear operational runway for these agents to thrive. The right infrastructure and the ability to interact with multiple software systems enable agents to orchestrate complex, multi-system workflows with precision and high-impact efficiency. That means enterprise-grade governance around agentic systems must improve.

Major foundation model providers are increasingly implementing proprietary harness capabilities directly into their ecosystems. These exclusive harnesses often provide better performance optimization, more seamless coordination and enhanced access to model-specific capabilities. The prevailing industry sentiment is that these environments will consistently deliver the best results. Case in point: If you want the strongest performance from Claude, you’re better off using Anthropic’s surrounding ecosystem and harnessing infrastructure rather than treating the model as a standalone component.

That said, there’s also value in maintaining the freedom to jump between models on a daily basis. Most developers, me included, switch between something like six models daily, whether that’s Claude, Gemini, Muse or an open-source option, depending on the task. That flexibility gets much harder to preserve once a company builds on a provider-specific harness, such as Anthropic. While this will likely improve performance and cut costs, the trade-off is increased vendor lock-in.

This creates a strategic choice for organizations: Fully embrace a vendor ecosystem for immediate performance, or maintain ownership of your own orchestration layer? Use the harness provided by the model provider, or build your own custom harness tailored to your business requirements?

I remain hopeful that many enterprises will leverage vendor innovations, while ensuring their core business logic remains portable instead of embedded within closed proprietary systems. But only time will tell.

The many benefits of harnessed agents

A carefully designed agent harness does more than merely decrease risk. It also lays a foundation for implementing autonomous agents with better confidence. You can count on the safe deployment of autonomous agents in production environments. No more wondering whether or not an agent will exceed its authority: Your enterprise can define exactly what it is permitted to do. A robust harness also delivers fine-tuned control over agent actions and access to tools, including which APIs, enterprise systems and software resources that each agent can invoke. Compliance-ready auditability is equally important for regulated industries.

The bottom line is that you can rely on the right harness to provide better peace of mind, transforming your AI into a transparent operational system that ensures reduced operational risk while seamlessly amplifying automation. The result is scalable AI systems that companies can actually trust.

Trust isn’t guaranteed just because a model scores well. It’s earned via system predictability. As my friend and former Google colleague Ben Mathes warned me, crafting custom rules around today’s models is risky. That’s because every few months, new foundation models make yesterday’s engineering workarounds extinct. We should instead prioritize building robust frameworks that can adapt as models progress.

I believe lasting advantage comes from fat skills – modular, detailed instruction sets that tell an AI agent how to perform a specific task without cluttering its core system – and fat prompts that capture institutional knowledge, along with rigorous backends that meticulously organize enterprise data. This enables the harness to evolve alongside improving models without needing to be completely rebuilt, which means business expertise can remain the primary fuel that powers AI success.

Actionable steps for enterprises

So, what are the best practices going forward? CIOs and CTOs should treat agent governance as a core infrastructure decision. Procurement focus needs to expand from models to platforms to, ultimately, control systems. And enterprises need to understand that competitive advantage will be dependent on three factors:

  1. Safety – Does the model safely do what you wanted to do?
  2. Performance – Does it do it well?
  3. Costs – Does it do it with relatively low expense?

Professionals in this space now face the strategic decision I mentioned earlier: use vendor-provided harnesses and maximize performance, or build proprietary internal harnesses to preserve flexibility and avoid vendor lock-in.

Without a resilient harness, you risk slower adoption due to security concerns. For example, Tesla is rolling out a $200 token-per-month cap on employee spending on third-party AI tools at around the same time a new Claude model debuted with lower per-request token costs. Yes, safety continues to be nonnegotiable. But once you meet that threshold, optimizing performance and expense becomes the Pareto Frontier problem your organization should be closely watching.

The AI arms race is no longer merely about smarter models. Instead, it’s about safely deploying autonomy at minimal cost. That’s why implementing an appropriate agent harness is so crucial. It becomes the critical operating layer that allows intelligent agents to reliably function inside an enterprise.

As we transition to the next phase of AI adoption, control is going to matter as much as capability to executives. The LLM also matters, of course, but without the proper framework, it can’t operate effectively. The organizations that dominate won’t necessarily have the best model; instead, they’ll have the most effective framework for deploying and governing autonomous agents.

Antes de ontemStream principal
  • ✇Security | CIO
  • Who gets to decide? The CIO and the new architecture of enterprise authority
    For most of my career, technology governance began with a familiar set of questions: Is the system secure? Is it resilient? Does it meet the architecture standard? Can we afford it? Those questions still matter. But they are no longer enough. AI is moving rapidly from producing content and recommendations to initiating actions. It can route work, change code, approve exceptions, communicate with customers, trigger transactions and coordinate other systems. In that environm
     

Who gets to decide? The CIO and the new architecture of enterprise authority

1 de Setembro de 2026, 06:00

For most of my career, technology governance began with a familiar set of questions: Is the system secure? Is it resilient? Does it meet the architecture standard? Can we afford it? Those questions still matter. But they are no longer enough.

AI is moving rapidly from producing content and recommendations to initiating actions. It can route work, change code, approve exceptions, communicate with customers, trigger transactions and coordinate other systems. In that environment, the most important question may not be what the technology can do. It is who, or what, has the authority to do it.

That distinction is becoming urgent. Stanford University’s 2026 AI Index reports that 88% of surveyed organizations used AI in 2025, while deployment of AI agents remained in the single digits across nearly all business functions. The gap matters. Enterprises have gained broad experience using AI as a tool, but many are only beginning to understand AI as an actor inside an operating model.

I call the resulting challenge the enterprise authority gap: The distance between the speed at which intelligent systems can act and the enterprise’s ability to define, constrain and account for that action. Closing that gap will require more than an AI policy. It will require an architecture for decision rights.

The authority gap is becoming an operating risk

Traditional systems execute permissions. AI-enabled systems increasingly interpret intent. That is a fundamental change.

A conventional application may allow an employee to approve a payment up to a defined limit. An AI agent may evaluate the request, assemble supporting information, communicate with another system, recommend an exception and initiate the next step. Each individual action may appear legitimate. The combined sequence may create an authority that nobody explicitly granted.

I learned a version of this lesson long before generative AI. In banking and payments environments I led, the most consequential risks were rarely contained within one application. They emerged where business rules, identity, workflow, vendor dependencies and operational exceptions met. A payment platform could be technically sound and still create exposure if decision rights were unclear during an exception, outage or recovery event. The control was not simply in the code. It was in knowing who could act, under what conditions and with whose accountability.

AI compresses those seams. It can traverse data, applications and organizational boundaries in seconds. If the enterprise has not made authority explicit, the system will inherit whatever permissions, defaults and informal practices already exist. Automation then turns ambiguity into scale.

This is why I believe consequence, not activity, should set the control boundary. The same technical action can carry very different enterprise consequences. An agent rescheduling an internal meeting is not equivalent to an agent changing a customer credit decision, releasing software into production or moving money. Governance that treats all AI activity alike will either obstruct low risk work or insufficiently control high-risk work.

Regulators and standards bodies are already pointing in this direction. The NIST AI Risk Management Framework organizes AI risk work around govern, map, measure and manage, with governance operating across the lifecycle. The European Union’s AI Act requires high-risk systems to support effective human oversight, including the ability to monitor, interpret and override their operation. These are important foundations. For the CIO, however, the operating question remains practical: How are those principles translated into enforceable authority inside the architecture?

Build an architecture for decision rights

Enterprises need an enterprise authority architecture: A deliberate model connecting business decisions, human accountability, machine autonomy and technical enforcement. I would build it around four disciplines.

  1. Define the decision before selecting the technology. Teams often begin with a model, platform or agent and then search for a use case. I have found the reverse sequence to be more durable. Start with the business decision or workflow being changed. Identify its economic value, affected stakeholders, existing control owner and consequence of failure. Only then determine whether AI should inform the decision, recommend an action or execute it.
  2. Separate capability from authority. A system may be capable of completing a task without being authorized to complete it independently. That difference should be visible in design. I use a simple progression: observe, recommend, prepare, execute within limits and execute with exception authority. Moving from one level to the next should require evidence, not enthusiasm. Accuracy is one input, but so are reversibility, explainability, financial exposure, customer impact and recovery time.
  3. Make authority technically enforceable. Policy statements do not stop an agent from calling an API. Authority must be expressed through identity, entitlements, transaction limits, segregation of duties, approval gates, runtime monitoring and kill mechanisms. Every consequential action should leave an attributable record: what the system knew, what rule it applied, what it did and which accountable owner accepted that operating boundary. This is where architecture and operations must reconnect. In several cloud and infrastructure transformations I led, we learned that control-plane resilience mattered as much as workload resilience. A service could be available while the organization had lost the ability to govern or recover it. The same principle applies to AI. An autonomous capability is not enterprise-ready if the organization cannot constrain it, observe it and reclaim control when the normal management path fails.
  4. Measure the economics of authority. Leaders should move beyond model accuracy and unit cost to what I call liability-adjusted autonomy: the value created by delegated execution after accounting for oversight, error correction, compliance, recovery and potential harm. A faster decision is not automatically a better decision. If it increases the cost of exceptions, transfers hidden work to employees or creates an unbounded downside, the apparent productivity gain is misleading.

The measurement question should therefore be: What is our cost per successful, governed outcome? That connects technology performance to business value without pretending risk is external to the calculation.

Consider a fraud-alert workflow. AI may be highly effective at prioritizing cases, assembling evidence and recommending disposition. Those capabilities can reduce analyst effort and improve response time. But authority to block an account, decline a transaction or communicate suspected fraud to a customer carries a different consequence. The architecture should assign separate thresholds, evidence requirements and escalation paths to each decision rather than treating the workflow as one automation opportunity.

The same logic applies outside financial services. In healthcare, recommending a scheduling change is different from changing a treatment pathway. In manufacturing, predicting equipment failure is different from stopping a production line. In human resources, drafting a job description is different from filtering candidates. The relevant boundary is not whether AI is present. It is how much consequential authority the enterprise has delegated.

Who gets to decide? Capability is not authority. The CIO’s role in the new era of enterprise architecture

Rajjie Sarmey

The CIO’s next mandate is enterprise authority

This mandate changes the CIO’s relationship with the rest of the enterprise. Decision rights cannot be owned by IT alone because the consequences do not remain in IT. Business leaders own outcomes. Risk and legal leaders interpret obligations. Security establishes trust boundaries. Human resources shapes workforce practices. Audit tests whether controls operate as intended. The CIO’s unique role is to make those responsibilities coherent and executable across the technology estate.

That work should begin with an authority inventory. Most organizations can produce an application inventory, and some can produce a credible AI inventory. Far fewer can show where machines influence or execute consequential decisions, which identities they use, which systems they can reach, who approved that reach and how authority is withdrawn.

I would ask every leadership team five questions. Which decisions are we allowing AI to influence? Which actions can it execute without human approval? What is the maximum consequence of a wrong or manipulated action? Who is accountable when several systems contribute to the outcome? Can we stop, reverse and reconstruct the decision within the time the business requires? If those answers are fragmented across policy documents, vendor configurations and tribal knowledge, the enterprise does not yet control its autonomy.

Boards should ask a related question. Are we governing AI as a portfolio of experiments or as a new distribution of enterprise authority? The first view focuses on investment, adoption and risk reporting. The second recognizes that AI can alter how the company makes commitments, treats customers, allocates capital and exercises judgment. That is a governance issue, an operating-model issue and increasingly a fiduciary issue.

My experience across architecture, operations, modernization and enterprise transformation has taught me that accountability cannot be added after scale. By then, the most expensive decisions have already been embedded in platforms, permissions and process design. Authority must be designed at inception, tested before deployment and monitored throughout operation. Human review at the end of a poorly bounded system is not meaningful oversight. It is often an expensive illusion.

The organizations that lead in the next phase of AI will not be those that automate the most decisions. They will be those that know which decisions deserve automation, what evidence earns greater autonomy and where human judgment must remain nondelegable.

The CIO has an opportunity to lead that transition. Not as the owner of every decision and not as the enterprise’s technology gatekeeper, but as the architect who connects intelligence to authority, authority to accountability and accountability to measurable value.

The next generation of CIO leadership will not be defined by how much intelligence the enterprise deploys. It will be defined by how wisely the enterprise distributes authority.

  • ✇Security | CIO
  • Bedrock, Vertex or build it yourself: The AI infrastructure decision most CIOs get backwards
    Across dozens of enterprise procurement reviews, I see technology executives make the same expensive mistake. They start their cloud AI strategy with the wrong question: “Which provider offers the smartest model today?” I sat through a meeting where a client’s leadership team listened to a slick 45-minute vendor pitch highlighting benchmark scores, processing limits and exclusive model access. By the end of the presentation, the executives were ready to sign a multi-
     

Bedrock, Vertex or build it yourself: The AI infrastructure decision most CIOs get backwards

31 de Agosto de 2026, 08:00

Across dozens of enterprise procurement reviews, I see technology executives make the same expensive mistake.

They start their cloud AI strategy with the wrong question: “Which provider offers the smartest model today?”

I sat through a meeting where a client’s leadership team listened to a slick 45-minute vendor pitch highlighting benchmark scores, processing limits and exclusive model access. By the end of the presentation, the executives were ready to sign a multi-year, multi-million-dollar commitment just to secure priority access to that single model.

I watched experienced leaders prepare to make a permanent infrastructure commitment based entirely on a temporary technological lead. Signing a long-term contract based on a six-month feature advantage treats a rapidly commoditizing utility service as a permanent asset, while surrendering control over the true intellectual property of your business.

The central strategic axiom

Raw computational intelligence is a rented utility overhead. Proprietary corporate context is owned enterprise capital. Never tie the permanent location of your corporate capital to the temporary rental location of a utility.

The economics of rented intelligence vs. owned capital

The top-performing commercial model on the market today will inevitably be matched or surpassed shortly by a cheaper, faster alternative. As Sequoia Capital detailed in its analysis of market economics, massive capital continues to pour into underlying processing infrastructure, driving the baseline cost of raw intelligence steadily downward toward commodity pricing.

When I evaluate technology investments with CFOs and CIOs, we strictly separate variable operational utilities from durable intellectual property across four strategic dimensions:

  • Market nature: Rented processing capabilities operate on fast-changing, highly commoditized and declining price curves. Owned corporate context forms unique, proprietary and highly defensible business positions.
  • Enterprise assets: Rented utilities encompass raw processing power, external models and third-party cloud infrastructure. Owned context includes customer ledgers, internal business rules, compliance frameworks and institutional memory.
  • Commercial strategy: Rented capabilities require a pay-as-you-go, unbundled approach that embraces maximum supplier churn. Owned context requires total asset ownership, isolated environments and zero vendor lock-in.
  • Financial objectives: The financial goal for rented capabilities is minimizing marginal cost per transaction. The financial goal for owned context is maximizing long-term enterprise valuation.

Raw processing power should be managed like electricity: your systems connect to the provider, consume what is required for the task and retain total freedom to switch utility suppliers if pricing or performance dictates a change.

Your corporate context, however, is a permanent capital asset. As Harvard Business Review has demonstrated across past technology cycles, lasting competitive advantage is built on proprietary data, unique operational workflows and institutional memory (never on shared infrastructure). A commercial model possesses zero understanding of your firm’s private pricing structures, key client nuances or regulatory boundaries until you feed it your context.

The mechanics of vendor capture

In my architecture reviews, I constantly see how managed cloud platforms naturally blur the line between rented processing power and owned corporate context.

Integrated cloud environments rarely position processing power as a standalone, interchangeable utility. Instead, platform architectures naturally encourage corporate engineering teams to bundle processing power with proprietary storage formats, closed management tools and native operational frameworks.

I have watched this architectural design trap enterprise teams in three distinct phases:

  1. Data entanglement: Corporate knowledge becomes formatted to fit a specific vendor’s environment, making future extraction costly and complex.
  2. Workflow dependence: Business rules and approval logic are built directly into vendor-owned management software, tying daily operations to their platform.
  3. Loss of leverage: During contract renewals, the enterprise cannot credibly threaten to switch providers because moving away requires a multi-month operational migration.

Once your business rules and customer records are deeply bound to a single vendor’s ecosystem, your negotiating leverage vanishes.

The architectural mandate: Vendor-neutral gateways

To preserve commercial leverage and maintain operational agility, I advise technology leaders to mandate an internal management layer between core corporate applications and external technology providers. Gartner’s strategic cloud planning research projects that the vast majority of enterprise organizations will require a multi-provider strategy specifically to prevent commercial lock-in and control long-term operating costs.

An internal control gateway acts as a central management point. Instead of allowing individual applications to establish direct connections to an external cloud vendor, every application communicates exclusively with your internal gateway.

This gateway enforces three mandatory executive controls:

  • Cost-optimized task routing: The gateway evaluates incoming tasks and automatically routes them to the most cost-effective external provider available. Routine administrative tasks go to low-cost utility systems, while complex tasks go to high-capacity options.
  • Centralized data protection: Before corporate data leaves the enterprise perimeter, the gateway strips sensitive customer details and logs the transaction for compliance verification.
  • Commercial agility: Because company applications interact solely with your internal gateway rather than directly with a vendor, you retain the ability to switch cloud providers instantly. If a vendor raises prices or a competitor releases a superior option, your team simply updates a routing rule within your internal system.

Rent the computational processing power as a temporary utility, but retain total ownership and control over the corporate nervous system.

  • ✇Security | CIO
  • The enterprise AI race will be won by platform teams, not prompt engineers
    Every enterprise AI conversation I hear eventually turns to prompts. Which prompting technique produces the best results? Which large language model reasons more effectively? Which framework generates more accurate answers? These are worthwhile questions, and I understand why they dominate the conversation. Prompt engineering has become one of the most visible aspects of enterprise AI because it delivers immediate, tangible results. A better prompt can transform an
     

The enterprise AI race will be won by platform teams, not prompt engineers

25 de Agosto de 2026, 07:00

Every enterprise AI conversation I hear eventually turns to prompts.

Which prompting technique produces the best results? Which large language model reasons more effectively? Which framework generates more accurate answers? These are worthwhile questions, and I understand why they dominate the conversation. Prompt engineering has become one of the most visible aspects of enterprise AI because it delivers immediate, tangible results. A better prompt can transform an average response into an exceptional one within seconds.

But after spending years building enterprise platforms, leading cloud modernization initiatives and operating mission-critical systems, I have reached a different conclusion.

The organizations that ultimately win the AI race will not be distinguished by who writes the best prompts. They will be distinguished by who builds the strongest enterprise platforms.

Prompt engineering can improve the quality of an AI interaction. Platform engineering determines whether AI can become a trusted, scalable capability that transforms an entire business.

This aligns with Gartner’s view that platform engineering is becoming a foundational discipline for improving developer productivity and standardizing enterprise software delivery, creating the operational foundation that AI initiatives increasingly depend upon.

Prompt engineering is only the beginning

Prompt engineering deserves its popularity. It lowers the barrier to entry for AI, allows teams to experiment quickly and helps organizations discover new ways to improve productivity. Business users can automate repetitive tasks, developers can accelerate coding and analysts can uncover insights faster than ever before.

These early successes are important because they build confidence in AI.

However, I have noticed that many organizations mistake successful experimentation for enterprise readiness.

McKinsey has reached a similar conclusion in its research on agentic AI, arguing that lasting enterprise value comes from redesigning workflows, operating models and governance around AI rather than simply deploying increasingly capable models.

Creating a useful AI demonstration is relatively straightforward. Turning that demonstration into a secure, reliable business capability is considerably more difficult.

The real questions begin after the pilot succeeds.

Where does the AI retrieve its information? How is sensitive data protected? Which systems can the AI interact with? How are responses validated? Who owns the workflow when something fails? How are changes deployed safely? How do we measure accuracy over time? How do we maintain governance while allowing innovation?

These are not prompt engineering problems.

They are platform engineering problems.

Enterprise AI is an infrastructure challenge

In my experience, enterprise AI behaves much like every major technology transformation that preceded it. Whether organizations were adopting cloud computing, enterprise integration, DevOps or platform engineering, long-term success rarely depended on selecting the newest technology. It depended on building an operational foundation that could support continuous growth.

AI follows the same pattern.

In an earlier CIO.com article, I argued that the next AI bottleneck would not be the model itself but the enterprise infrastructure surrounding it. The same principle applies here because platform teams are responsible for building that infrastructure at scale.

A large language model does not operate in isolation. It depends on APIs to access business applications. It requires secure identity management before performing actions on behalf of users. It needs clean, governed data to produce reliable responses. It relies on messaging systems, event streams, monitoring platforms, deployment pipelines and security controls to function consistently within enterprise environments.

Every AI interaction touches dozens of enterprise services that most users never see.

When AI performs well, the model often receives the credit.

When AI fails, the root cause is frequently somewhere else entirely.

I have seen situations where outdated data, unavailable APIs, inconsistent permissions or unreliable integration services create failures that appear to be AI problems but are actually infrastructure problems.

The model simply exposes weaknesses that already existed within the enterprise architecture.

Platform teams build enterprise trust

One of the most overlooked aspects of AI adoption is trust.

Employees will only embrace AI if they believe it consistently provides accurate, timely and secure information. Business leaders will only automate critical processes if they understand how decisions are made. Security teams will only approve broader deployment when governance is embedded into the platform itself.

Trust cannot be created through prompts.

It is built through architecture.

Platform engineering teams establish standardized APIs, reusable services, identity controls, deployment automation, observability, logging, auditing and policy enforcement that make AI predictable rather than experimental.

Instead of every business unit creating its own AI implementation, platform teams provide reusable capabilities that allow innovation to scale without creating operational chaos.

That is the difference between isolated success stories and enterprise-wide transformation.

Integration will separate leaders from followers

Throughout my career, I have consistently found that integration determines whether technology delivers business value.

AI is no different.

Every meaningful AI workflow eventually becomes an enterprise integration workflow.

An AI assistant may retrieve customer information from a CRM system, verify inventory through an ERP platform, initiate an approval workflow, update a service ticket, notify a collaboration platform and record every action for auditing.

None of these activities depend solely on prompt engineering.

They depend on reliable APIs, event-driven architecture, secure messaging, resilient infrastructure and well-designed automation.

Organizations that already possess mature platform engineering capabilities have a significant advantage because they can integrate AI into existing operational processes instead of building disconnected point solutions.

The conversation should no longer be, “How do we deploy another AI assistant?”

It should be, “How do we make AI another trusted service within our enterprise platform?”

Platform engineering is becoming AI engineering

I believe one of the biggest organizational shifts over the next several years will be the evolution of platform engineering.

Historically, platform teams focused on developer productivity, cloud infrastructure, automation, observability and operational reliability.

Those responsibilities are expanding rapidly.

Today’s platform teams are increasingly responsible for AI gateways, model orchestration, retrieval services, vector databases, prompt management, policy enforcement, cost optimization and AI observability.

They are becoming the teams that connect AI to the rest of the enterprise.

This evolution requires new skills, but it builds upon capabilities many platform organizations already possess.

They understand automation.

They understand reliability.

They understand governance.

Most importantly, they understand how to create standardized services that hundreds or thousands of developers can safely consume.

That expertise will become one of the greatest competitive advantages in enterprise AI.

Governance must scale with innovation

The NIST AI Risk Management Framework reinforces this approach by encouraging organizations to govern, measure, manage and continuously monitor AI risks throughout the system lifecycle rather than treating governance as a final checkpoint.

As AI evolves from answering questions to executing business processes, governance becomes inseparable from innovation.

Organizations cannot afford to treat governance as a review step that occurs after deployment.

Identity management, access controls, auditability, observability, policy enforcement and regulatory compliance must become foundational components of the platform itself.

The most successful enterprises will not slow innovation through excessive controls.

Instead, they will build platforms where secure innovation becomes the default experience.

Developers should not have to reinvent governance every time they build a new AI capability.

The platform should provide those guardrails automatically.

That is how organizations innovate at scale.

The organizations that win will think beyond models

The AI industry will continue producing larger models, better reasoning capabilities and more sophisticated agents.

Those advances will matter.

But I believe the lasting competitive advantage will belong to organizations that invest equally in the platforms surrounding those models.

Five years from now, I do not think enterprise leaders will remember which company wrote the most sophisticated prompts.

They will remember which organizations built AI platforms that employees trusted, security teams approved, developers embraced and business leaders could confidently scale across the enterprise.

Models will continue to evolve.

Prompts will continue to improve.

The organizations that separate themselves from everyone else will be the ones whose platform teams quietly made AI reliable, secure, integrated and operational.

In the enterprise AI race, that is where the real competitive advantage will be built.

  • ✇Security | CIO
  • From ‘dumb iron’ to smart machines: Why data control is the real Industry 5.0
    On the modern factory floor, the phrase “industrial equipment” no longer tells the whole story. It conjures images of steel, hydraulics, conveyor belts and machinery built to perform the same task with unwavering precision day after day. Physical engineering remains fundamental, of course, but it’s no longer the sole measure of a machine’s value. The next generation of machines have capabilities that depend on far more than the factory floor, continuously exchanging infor
     

From ‘dumb iron’ to smart machines: Why data control is the real Industry 5.0

11 de Agosto de 2026, 06:00

On the modern factory floor, the phrase “industrial equipment” no longer tells the whole story. It conjures images of steel, hydraulics, conveyor belts and machinery built to perform the same task with unwavering precision day after day. Physical engineering remains fundamental, of course, but it’s no longer the sole measure of a machine’s value. The next generation of machines have capabilities that depend on far more than the factory floor, continuously exchanging information with cloud platforms, data centers and AI systems that allow them to act autonomously and “self-improve” long after they’ve been deployed. A robotic arm isn’t simply running a predefined script anymore – it’s generating a constant stream of operational intelligence that reveals how it is performing, when and whether it needs attention, and how production can independently improve itself and become faster, safer and more efficient tomorrow than it is today.

This new functionality is redrawing the concept of ownership for manufacturers. Increasingly, the asset is not just the machine itself, but the flow of data that supports it and reveals clues about its functionality. Every production cycle enriches digital models, refines predictive algorithms and deepens operational understanding, turning what was once a static piece of equipment into something that continuously improves over time. The term “phygital” has emerged to describe this convergence of physical infrastructure and digital intelligence, but whatever terminology ultimately sticks, the outcome will be the same. As manufacturing enters an era where competitive advantage is increasingly shaped by software, analytics and real-time AI inference, CIOs are having to think very carefully not just about who owns the machine on the factory floor, but who controls the data that turns that machine from “dumb iron” into something that can “think” intelligently.

Manufacturing has entered its software-defined era

The physical engineering on display on factory floors is already impressive. Autonomous haul trucks can navigate vast mining sites without drivers, robotic arms can self-adjust their movements in relation to contextual cues, and in the case of so-called “dark factories,” entire production lines can operate 24/7 for a long time without a single person on the factory floor. Every movement, vibration, temperature change and production cycle becomes part of a data-driven feedback loop that allows software to refine performance, anticipate failures and adapt operations contextually in ways that simply weren’t possible when industrial equipment functioned in siloes.

According to Deloitte’s 2025 Smart Manufacturing and Operations Survey, 92% of manufacturers believe smart manufacturing will be the primary driver of competitiveness over the next three years, while 78% are allocating more than a fifth of their improvement budgets to smart manufacturing initiatives. Those figures bring home the fact that industrial performance is no longer determined solely by what happens inside a machine, but by how effectively the data ecosystem it lives in functions as a whole.

Every smart factory runs on an invisible supply chain

Every intelligent machine exists within a much broader ecosystem that stretches far beyond the walls of a factory, connecting equipment manufacturers, cloud platforms, systems integrators, AI providers and operational teams through a constant flow of data. It’s easy to think of a production line as a collection of individual assets working side by side, but the reality is far more interconnected. Each machine is both producing and consuming information throughout the working day, allowing decisions made in one environment to influence outcomes somewhere else. A software update developed by an equipment manufacturer, for example, might be informed by performance data gathered from thousands of identical machines operating around the world, with improvements delivered back to the factory almost as quickly as they’re identified.

From that perspective, data begins to resemble a supply chain in its own right. Manufacturers have spent decades refining the movement of raw materials because every unnecessary delay carries a measurable operational cost, and that same principle now applies to information. The data flowing from production equipment, the analytics returning from cloud platforms, and the insights generated by AI have become just as vulnerable to delay as the components arriving at the loading dock.  According to the International Federation of Robotics, more than 4.8 million industrial robots are now operating in factories worldwide, and each one contributes to a growing stream of operational data that has become inseparable from the manufacturing process itself. The challenge for CIOs used to be, “How do we connect these environments?”, but now it’s “How do we ensure the data moving between them arrives with the speed, visibility and control needed to keep pace with modern manufacturing?”

The importance of network architecture

The value of data used to be measured solely by its accuracy, but now it depends on how reliably it can move between the organizations that create it, analyze it and act upon it. A predictive maintenance platform can’t identify an emerging fault if telemetry arrives too late, just like a digital twin is only as useful as the information it receives. As we bridge from Industry 4.0 to Industry 5.0, the network itself is becoming an active participant in the production process, prompting CIOs to think differently about connectivity.

Modern manufacturing depends on a growing ecosystem that needs to exchange data in near real time. Rather than relying on unpredictable routes across the public Internet, many organizations are turning to direct interconnection in the form of internet, cloud and AI exchanges, which act as neutral meeting points where enterprises and their suppliers, as well as network operators, cloud providers and digital or AI service providers, can establish direct, private connections with one another. By shortening the path data has to travel and avoiding unnecessary “hops” and congestion, these platforms reduce latency, improve resilience and give organizations far greater visibility and control over how production-critical information moves.

Every revolution in manufacturing has been defined by the emergence of a resource that reshaped how value was created, whether that was steam, electricity or silicon. Industry 5.0 is introducing another, albeit one that can’t be stored in a warehouse or delivered on a truck. Data has become the factory’s most valuable raw material, and controlling its movement is every bit as important as controlling the movement of physical goods. The term “industrial equipment” may continue to describe what’s happening on the factory floor, but it no longer captures where competitive advantage is really being created. Increasingly, the intelligence surrounding a machine is becoming just as valuable as the machine itself, and the networks carrying that intelligence are becoming part of the production process in their own right.

  • ✇Security | CIO
  • Why meta agents must become the economic intelligence layer of the agentic enterprise
    In “Micro and macro agents: The emerging architecture of the agentic enterprise,” I proposed a three-layer architecture for enterprise AI. Micro agents execute specialized tasks. Macro agents orchestrate end-to-end business processes. Meta agents provide governance through monitoring, compliance, security, and human oversight. As enterprises begin deploying thousands — and eventually tens of thousands — of autonomous agents, token costs have become a major co
     

Why meta agents must become the economic intelligence layer of the agentic enterprise

4 de Agosto de 2026, 09:00

In “Micro and macro agents: The emerging architecture of the agentic enterprise,” I proposed a three-layer architecture for enterprise AI.

  1. Micro agents execute specialized tasks.
  2. Macro agents orchestrate end-to-end business processes.
  3. Meta agents provide governance through monitoring, compliance, security, and human oversight.

As enterprises begin deploying thousands — and eventually tens of thousands — of autonomous agents, token costs have become a major concern. According to Gartner, rising token-driven AI spend is straining budgets and challenging cost justification.

To track this economic concern, meta agents should do more than simply being the governance agents.

They should become the economic intelligence layer of the enterprise.

Their responsibility is not only ensuring AI behaves responsibly.

It is ensuring AI creates measurable business value.

The missing economic model for AI

Every major technology revolution eventually develops its own economic framework:

  • Manufacturing measured productivity.
  • Cloud computing measured infrastructure utilization.
  • Digital businesses measured customer acquisition costs and lifetime value.

The agentic enterprise now requires its own financial discipline. Every AI prompt. Every reasoning cycle. Every interaction between agents. Every autonomous workflow.

Tokens have quietly become the operational currency of enterprise AI. Tokenomics is now a foundational part of enterprise AI architecture.

Yet today, most organizations measure only one thing: Cost. How many tokens were consumed? Which models cost the most? What was the monthly inference bill?

These are useful operational metrics.

They are not strategic business metrics. Boards rarely ask how much electricity a factory consumed. They ask how much value the factory produced.

Enterprise AI deserves the same conversation.

This is where I was thinking about the laws of physics.  Based on physics laws,  energy cannot be created or destroyed. It is transformed into another form. Electricity becomes light. Chemical energy becomes motion. Solar energy becomes electricity.

Enterprise AI offers a similar management lesson.

Intelligence must be transformed into value

Tokens are not valuable because they are consumed. They become valuable only when they are transformed into business outcomes. A faster loan application decision. A fraud detection. A better customer experience. Higher software quality. Greater employee productivity. A new business opportunity.

This leads to what I call return on tokens (ROT).

ROT measures how effectively an organization converts token consumption into measurable business value.

Instead of asking, “How many tokens did we consume,” leaders should ask, “How much enterprise value did every million tokens create?”

The Second Law of Thermodynamics tells us something equally important: Every energy transformation introduces inefficiencies. Although total energy is conserved, some inevitably becomes less useful for doing work.

Enterprise AI behaves similarly.

The second law: Every AI transformation creates friction

Not every token creates value. Some tokens are spent on repeated reasoning. Some generate redundant conversations between agents. Some support oversized context windows. Some produce hallucinations requiring correction. Some route simple tasks to unnecessarily expensive models.

The tokens are not lost. But they create very little useful business work.

I refer to this as token entropy. Token entropy represents the portion of AI activity that consumes intelligence without producing proportional business outcomes.

Every agentic enterprise will experience token entropy. The organizations that win will be the ones that continuously identify and reduce it.

Beyond energy: The importance of exergy

Thermodynamics offers another concept that is even more relevant. It is called Exergy.

Unlike energy, exergy measures the amount of energy that can actually be converted into useful work. Two systems may contain the same amount of energy while producing dramatically different levels of useful output.

The same principle applies to enterprise AI. Two organizations may consume exactly the same number of tokens.

One generates meeting summaries.

The other transforms loan  processing, accelerates software development, detects fraud, improves customer retention, and creates new revenue streams.

Their token consumption is identical. Their business impact is not.

Borrowing it as a management analogy, not claiming that AI tokens literally obey the thermodynamic definition of exergy. I think of this as token exergy. It’s not that AI tokens literally obey the thermodynamic definition of exergy. 

Token exergy measures how much of an organization’s AI intelligence is converted into useful business work. It is not enough to consume tokens efficiently. Organizations must convert those tokens into outcomes that matter.

The meta agent evolves

This is where meta agents become transformational.

Today we think of them as governance agents. Tomorrow they become economic governors.

Meta agents continuously monitor every interaction across the enterprise and answer questions such as:

  • Which agents produce the highest ROT?
  • Where is token entropy increasing?
  • Which workflows generate the highest token exergy?
  • Which models deliver the greatest business value per token?
  • Which agents should use smaller models?
  • Which prompts should be optimized?
  • Which workflows require human intervention?
  • Which autonomous processes should be redesigned?

Meta agents no longer simply supervise AI. They optimize its economics.

The economic intelligence layer

The architecture now becomes complete.

  • Micro agents: Perform work.
  • Macro agents: Coordinate work.
  • Meta agents: OGovern, observe, optimize, and continuously improve the economics of intelligence.

Their objective is straightforward:

  • Maximize return on tokens.
  • Minimize token entropy.
  • Increase token exergy.

This represents a shift from AI governance to AI economics.

The executive dashboard of tomorrow

The executive dashboard of the future will not focus solely on infrastructure metrics. It will measure intelligence performance.

Imagine a boardroom dashboard displaying:

  • Return on tokens (ROT)
  • Token entropy index
  • Token exergy score
  • Business value per million tokens
  • Agent productivity index
  • Cost per autonomous decision
  • AI value by business unit
  • Human escalation rate
  • Model effectiveness score

These metrics move AI discussions beyond engineering. They make AI accountable for business outcomes.

A new responsibility for CIOs

The next generation of CIOs will not simply deploy AI. They will manage an economy of intelligence.

Their role will resemble that of a portfolio manager — allocating AI capacity where it creates the greatest enterprise value, reducing waste, and continuously improving the productivity of every autonomous workflow.

That responsibility cannot be fulfilled by dashboards alone. It requires an intelligent layer capable of observing, learning, and optimizing the entire agent ecosystem.

That is the emerging role of the meta agent.

The next competitive advantage

Every technological revolution rewards organizations that learn to measure what others overlook.

Factories measured productivity — not fuel consumption.

Digital businesses measured customer engagement — not server utilization.

The agentic enterprise will reward organizations that measure intelligence itself.

The winners will not be those deploying the largest models. Nor the most agents. Nor consuming the fewest tokens.

They will be the organizations that continuously maximize return on tokens, relentlessly reduce token entropy, and increase token exergy.

I believe this is the next evolution of the agentic enterprise.

Not simply governed intelligence, but economically optimized intelligence.

The AI adoption spending spree is over. Time to focus on value.

And in that future, meta agents will serve not only as the guardians of AI — but as the stewards of enterprise intelligence economics.

Through this framework I strongly believe that executives can easily remember the key measures for economic intelligence. 

  • ROT (return on tokens): How much value did AI create?
  • Token entropy: Where are we wasting AI intelligence?
  • Token exergy: How effectively are we converting AI intelligence into useful business work?

This article is published as part of the Foundry Expert Contributor Network.
Want to join?

  • ✇Security | CIO
  • Why AI infrastructure needs a new operating model
    The next AI infrastructure crisis may come from unmanaged inference capacity. For the past several years, the AI infrastructure conversation centered on one question: how do we get more compute? That made sense. Enterprises needed GPUs, cloud capacity, foundation models and room to experiment. Compute became shorthand for AI readiness. Production AI changes the operating discussion. Utilization, routing, latency, throughput, cost control, policy, privacy and governan
     

Why AI infrastructure needs a new operating model

4 de Agosto de 2026, 07:00

The next AI infrastructure crisis may come from unmanaged inference capacity. For the past several years, the AI infrastructure conversation centered on one question: how do we get more compute?

That made sense. Enterprises needed GPUs, cloud capacity, foundation models and room to experiment. Compute became shorthand for AI readiness.

Production AI changes the operating discussion. Utilization, routing, latency, throughput, cost control, policy, privacy and governance now need to be managed together. A GPU that sits idle creates no business value. A model endpoint with unpredictable latency frustrates users. An inference stack that cannot be measured end-to-end becomes difficult to defend when usage grows and finance asks where the money is going.

CIOs need governed capacity.

Governed capacity means operating AI infrastructure as a production system rather than a collection of disconnected resources. They need to know how much useful output their infrastructure produces, where that output runs, why it runs there, what it costs, how it performs, what policy applies and whether the system can be controlled as demand changes.

Enterprises buy AI infrastructure to deliver answers, summaries, recommendations, software code, customer interactions, analysis, automation and agent workflows. Those outputs need to be reliable, measurable and affordable enough to keep running.

The pilot-era stack is reaching its limit

The first wave of enterprise AI rewarded speed. Teams bought GPUs, reserved cloud capacity, tested APIs, adopted open-source models and assembled whatever stack helped them move.

Infrastructure inefficiency then becomes a business issue.

The symptoms are familiar: more systems to manage, more vendors to coordinate, more integration work and less visibility into what drives cost and performance.

That creates friction across the organization. IT teams support AI workloads that behave differently from traditional enterprise applications. AI teams need speed, but often lack the infrastructure control to tune cost, latency, utilization and performance together. Finance teams want predictable unit economics, but the stack was assembled under pressure and is hard to measure end to end.

Most teams can now get access to models and compute. Fewer can show how each workload is performing, where it runs and what it costs.

Capacity needs control

Extra capacity can still leave teams with idle infrastructure, uneven latency and unclear unit costs.

The useful questions are operational. Can the organization see utilization across teams, tenants, models and infrastructure pools? Can it route workloads based on cost, latency, privacy, availability and service objectives? Can it measure cost per token, cost per inference, cost per user interaction or cost per business workflow?

Inference behavior changes constantly. Demand fluctuates. Longer contexts increase cost. Model choice affects latency and output quality. Utilization varies across workloads. A customer-facing assistant may prioritize response time. A batch workflow may prioritize throughput and cost.

A procurement-led AI strategy cannot manage that complexity on its own. CIOs need an operating model for production inference.

Enterprise Linux offers a useful analogy. Linux gave companies flexibility and attractive economics, but enterprises needed a trusted operating layer and support model before using it for business-critical systems. AI infrastructure is reaching a similar stage. The models, hardware and software components already exist. Many organizations now need a way to operate them consistently and economically in production.

Token economics is becoming a management discipline

The useful output of many AI systems is delivered through tokens. That makes token economics a practical operating metric.

Token volume needs context. A token that helps complete a task, answer a question or resolve a customer issue creates value. A token generated through poor routing, excess latency or an unnecessarily expensive model adds cost without improving the outcome.

How much useful output are we getting per dollar? How much per watt? How much per GPU? How much per workload? How much per unit of latency? How much per business outcome?

Manufacturing leaders do not only ask how many machines they own. They ask what those machines produce, how often they sit idle, how much waste they create, how much energy they consume and how efficiently raw materials become finished goods.

AI infrastructure needs the same operating discipline: utilization, throughput, reliability, cost control and visibility into what the infrastructure is producing.

Enterprises need usability and control

Serverless AI APIs are fast to start and easy for developers. They work well for many use cases. As usage grows, economics can become harder to control and visibility into infrastructure behavior is limited.

Self-managed infrastructure gives teams more control and can improve long-term economics for persistent workloads. It also adds operational burden. Teams have to manage deployment, scaling, routing, model serving, monitoring, reliability, performance tuning, security, isolation and utilization.

Enterprises want the simplicity of managed services without giving up visibility and control. Developers should be able to access AI services without managing the underlying stack. Infrastructure, security and finance teams still need to see placement, cost, latency, utilization, tenant policy, service levels and risk.

That is the role of an inference operating layer: turning fragmented infrastructure into governed, measurable capacity that teams can manage as demand changes.

Beyond procurement

The more successful an AI application becomes, the more inference it consumes. As inference grows, cost, latency, utilization and governance determine whether the application can scale.

AI can repeat the cloud-cost pattern many CIOs already know. A service begins as an innovation accelerator, usage expands across teams and the bill grows faster than governance. By the time the organization tries to regain control, the architecture, workflows and vendor dependencies are difficult to unwind.

GPUs remain essential. Models remain essential. Data remains essential. Production AI also needs an operating layer around those assets.

The next generation of AI leaders will ask a harder question:

How much useful intelligence can we produce from our infrastructure, at what cost, with what reliability, under what policy and under whose control?

The answer will determine whether AI becomes a controlled production capability or another expensive system the business struggles to explain.

This article is published as part of the Foundry Expert Contributor Network.
Want to join?

  • ✇Security | CIO
  • Forward-deployed engineering in the age of agentic AI: From vibe coding to governed autonomy
    Forward-deployed engineering (FDE), has moved from being a niche delivery model to becoming one of the most important operating patterns for enterprise artificial intelligence. In traditional software programs, organizations could usually separate product engineering, implementation consulting, operations and governance into different teams. Agentic AI changes that separation. An agent does not merely answer a question; it may plan, call tools, read enterprise data, update
     

Forward-deployed engineering in the age of agentic AI: From vibe coding to governed autonomy

29 de Julho de 2026, 08:00

Forward-deployed engineering (FDE), has moved from being a niche delivery model to becoming one of the most important operating patterns for enterprise artificial intelligence. In traditional software programs, organizations could usually separate product engineering, implementation consulting, operations and governance into different teams. Agentic AI changes that separation. An agent does not merely answer a question; it may plan, call tools, read enterprise data, update systems, create artifacts, trigger approvals and continue across several steps. Because of this, production success depends less on a model demo and more on the careful engineering of business context, workflow boundaries, controls, observability and human accountability.

FDE addresses this gap by embedding engineering capability close to the business problem. A forward-deployed engineer works with product teams, domain experts, security teams, platform owners and end users to convert an AI idea into a working, governed, measurable system. The role combines software engineering, data engineering, cloud architecture, model evaluation, security design, user research and operational ownership. In the Agentic AI world, this blend is not optional. It is the difference between a clever prototype and a dependable production workflow.

Why FDE is needed in agentic AI

Agentic AI deployment is rarely a simple matter of selecting a foundation model and connecting it to a user interface. Enterprise agents operate inside business processes that already contain policies, data quality issues, exception paths, audit requirements, identity controls and legacy systems. A sales operations agent, for example, may need to read CRM records, interpret account notes, generate a renewal recommendation, check discount eligibility, route an approval and update the opportunity. Each of those steps introduces risk. The agent must know what it is allowed to do, what it should never do, when it must ask a human and how its decisions can be traced later.

This is where FDE becomes valuable. FDE teams do not treat Agentic AI as a packaged tool to be installed. They treat it as a socio-technical system that must be shaped around a real business workflow. They discover the actual process, map data dependencies, identify integration points, define controls, implement the orchestration, create evaluation suites and help the client team learn how to operate the system after the initial deployment. The forward-deployed model is therefore especially suited to the last mile of AI adoption, where most enterprise AI initiatives struggle. This view is consistent with Gartner’s guidance that enterprise AI agent success depends on proportional governance aligned to autonomy level and scope of access.

The FDE operating model

A mature FDE operating model normally follows a compressed but disciplined cycle. The team first clarifies the business outcome rather than accepting the initial solution request at face value. Next, it decomposes the workflow into tasks, decisions, systems, data sources and approval points. It then builds a thin production slice rather than a detached proof of concept. This slice includes real authentication, realistic data, monitored tool calls, repeatable tests and rollback options. Once the system is usable, the FDE team iterates with business users, tunes the agent behavior, improves the prompts or policies, hardens the integration layer and transfers operating knowledge to the internal team.

Figure 1: A matured FDE operating model
Figure 1: A matured FDE operating model.

Magesh Kasthuri

The important distinction is that FDE is not conventional staff augmentation. It is also not advisory consulting that ends with a roadmap. FDE is outcome-oriented engineering in the field. The engineer is close enough to the customer environment to see real constraints, yet technical enough to change the system directly. In Agentic AI, that proximity matters because small details can decide whether a workflow is trusted: a missing approval step, an overly broad tool permission, an unlogged data access, a weak retry policy or an untested edge case can undermine the entire deployment. Forrester similarly positions agentic AI as a competitive frontier that requires leaders to redesign workflows, governance and engagement models rather than simply automate existing tasks.

Operationalizing multi-step agentic workflows

Operationalization begins by turning an agent idea into an explicit workflow. Instead of saying, “build an agent that handles vendor onboarding,” an FDE team defines the stages: collect supplier information, validate tax details, screen sanctions lists, check contract thresholds, request procurement approval, create a supplier record and notify stakeholders. Each stage is assigned to a deterministic function, an AI agent, a human approver or a hybrid step. This decomposition reduces ambiguity and makes it easier to govern the process.

Figure 2: Operationalizing an agent workflow with FDE.
Figure 2: Operationalizing an agent workflow with FDE.

Magesh Kasthuri

In production, an FDE must design for state, retries, failures, idempotency and observability. Agentic workflows may run for minutes, hours or days. They may wait for an external API, pause for approval, recover from a system outage or resume after a user changes input. A robust implementation therefore needs checkpoints, durable state, structured event logs, trace IDs, policy-aware tool execution and dashboards that show where the workflow is stuck. Without these engineering controls, an agent may appear intelligent in a demo but become fragile in live operations. IDC’s perspective on agent adoption also highlights that the scale of enterprise agents will create major demands around orchestration, token efficiency, governance and cost containment.

Governing agentic AI through FDE

Governance in Agentic AI cannot be added as a final compliance checklist. It must be embedded into the workflow design. FDE teams help by translating policy into executable controls. For example, they can define which tools an agent may call, which data classes it can access, what confidence thresholds require escalation, which outputs need review and how exceptions are recorded. They also help establish evaluation datasets that reflect real operating scenarios rather than sanitized prompts.

A practical governance model usually contains five layers. The first is intent governance, which verifies that the agent is solving an approved business problem. The second is data governance, which controls source quality, access, retention and lineage. The third is tool governance, which restricts actions such as writing to systems, sending emails, executing code or changing financial records. The fourth is decision governance, which determines where humans must approve or override agent recommendations. The fifth is runtime governance, which monitors drift, failure patterns, cost, latency and policy violations. Gartner’s 2026 guidance on AI agent governance reinforces this need for differentiated controls, warning that uniform governance across all agents can create both over-restriction and under-restriction risks.

Securing agentic AI workflows

Security for Agentic AI is broader than prompt safety. Agents can act and every action surface must be secured. FDE teams typically implement least-privilege access, scoped credentials, tool allowlists, secrets isolation, input validation, output filtering, data loss prevention checks and protected execution environments. They also design approval gates for high-impact operations. A procurement agent may be allowed to draft a purchase order, but it should not submit the order above a threshold without explicit authorization.

Another important responsibility is defending against indirect prompt injection and tool misuse. If an agent reads an email, document, ticket or web page, that content may contain instructions that attempt to override the system policy. FDE engineers reduce this risk by separating instructions from data, sanitizing retrieved content, validating tool arguments, limiting write permissions and logging every external action. In regulated environments, they also design audit trails that show not only the final answer but the path the agent took to reach it. Forrester’s AEGIS framework similarly argues that agentic AI security must move beyond traditional infrastructure-centric controls toward intent-aware guardrails, observability, accountability and least-agency principles.

FDE and vibe coding: Relationship and tension

Vibe coding refers to a style of AI-assisted software development where a person expresses intent in natural language and an AI system generates much of the code. It can be extremely useful for prototyping, exploration, internal tools and rapid experimentation. In the context of FDE, vibe coding can accelerate the early build cycle because forward deployed engineers can quickly sketch integrations, generate boilerplate, create test harnesses and explore workflow alternatives with AI coding assistants.

However, FDE also provides the discipline that vibe coding alone lacks. An enterprise agent cannot rely on generated code that nobody has reviewed, tested or secured. The FDE approach turns intent-driven development into responsible engineering. The engineer may use AI to generate code, but then verifies it through code review, unit tests, integration tests, security checks, policy validation and operational monitoring. In short, vibe coding helps move faster; FDE ensures that speed does not come at the cost of reliability, maintainability or accountability. This aligns with Forrester’s caution that agentic systems can become harmful when they are misaligned, poorly governed or deployed without adequate experimentation and control.

Example: LangGraph with FDE orchestration

LangGraph is well suited for FDE-led Agentic AI implementations because it models workflows as graphs with state, nodes, edges, persistence and human-in-the-loop control. An FDE team can use it to build long-running, auditable workflows where every step is explicit. Consider a customer support escalation workflow. The graph may begin with ticket intake, move to classification, retrieve policy documents, ask a diagnostic agent to propose a resolution, route uncertain cases to a human reviewer and finally update the ticketing system.

In this model, the FDE defines the state schema, selects which nodes use LLM reasoning, separates deterministic validation from agentic reasoning and adds checkpoints so the workflow can resume after interruption. Human review is not an afterthought; it becomes a graph transition. If the confidence score is low or a policy exception appears, the workflow pauses for approval. This is an example of FDE orchestration: the framework supplies the runtime primitives, while the forward-deployed engineer shapes those primitives into a secure business process.

A simplified LangGraph-style pattern may include nodes such as intake_agent, retrieval_node, policy_checker, resolution_agent, human_approval and ticket_update. The FDE ensures that ticket_update can only run after validation and that sensitive customer data is masked before being sent to the model. The result is not merely an autonomous assistant; it is a controlled workflow that can be inspected, resumed, tested and improved.

Example: Microsoft AutoGen and Microsoft Agent Framework with FDE orchestration

Microsoft AutoGen popularized the idea of multi-agent conversations where agents with different roles collaborate to solve a task. In newer enterprise settings, Microsoft Agent Framework provides production-oriented patterns for agents and workflows, including sequential, concurrent, handoff, group chat and manager-led orchestration. An FDE can use these patterns to design a governed multi-agent system rather than a free-form conversation among bots.

For example, imagine an enterprise architecture review assistant. One agent reads the solution brief, another checks cloud security requirements, a third evaluates cost and FinOps implications and a fourth prepares a decision summary. A manager or orchestrator coordinates the agents, decides when to ask for missing information and routes the final recommendation to an architect for approval. The FDE defines agent roles, tool permissions, routing logic, approval-required actions and telemetry. If the security agent recommends a design exception, the workflow can pause until an authorized reviewer approves it.

In this case, FDE orchestration prevents the system from becoming an uncontrolled debate among agents. It introduces structure: which agent speaks when, which tools each agent may use, which outputs must be machine-readable and what evidence is required before a recommendation is accepted. This is especially important in Microsoft-centric enterprises where identity, audit, data boundaries and cloud governance must align with existing platforms.

Example: CrewAI with FDE orchestration

CrewAI is useful when the solution naturally maps to a team of role-based agents. It supports crews, tasks, processes, tools, memory and flows. An FDE team can use CrewAI to model collaborative work where specialized agents perform defined responsibilities. Consider a market intelligence workflow for a product team. A research agent gathers public signals, a competitor analyst compares positioning, a financial analyst estimates market impact and an editor agent prepares the final brief.

The FDE’s role is to make this collaboration production-ready. The engineer defines task boundaries, expected outputs, data sources, tool limits, escalation rules and quality checks. If the market intelligence brief is used for executive decision-making, the FDE may require citations, confidence notes, evidence tables and human approval before publication. CrewAI Flows can then be used to orchestrate event-driven execution, manage shared state and resume longer workflows where human feedback or external triggers are involved.

A practical CrewAI FDE pattern is to separate creative agent work from controlled workflow steps. Agents may draft, analyze and summarize, but deterministic validators check schema, sensitive content, data completeness and approval status. This hybrid design gives the enterprise the benefit of agent collaboration without surrendering control of the process.

Reference architecture for FDE-led agentic AI

A typical FDE-led Agentic AI architecture includes six layers. The experience layer contains chat, workflow, API or embedded user interfaces. The orchestration layer manages graphs, crews, workflows, handoffs, retries, checkpoints and human approvals. The agent layer contains specialized agents with defined roles, instructions, memory and tool access. The tool and integration layer connects to enterprise systems such as CRM, ERP, ticketing, document repositories, email, messaging, databases and APIs. The governance and security layer enforces identity, policy, secrets, monitoring, evaluation, logging and audit controls. The operations layer provides deployment automation, dashboards, incident handling, cost tracking and continuous improvement. Everest Group’s 2025 AI and Generative AI Services PEAK Matrix also notes that enterprises are moving beyond pilots toward production-grade AI initiatives, with emphasis on scalable architectures, responsible AI, security, compliance and outcome-based partnerships.

The FDE connects these layers into one operating system for AI adoption. The value is not only in the code. It is in the ability to make the code work inside the customer’s real environment, with the customer’s data, controls, users and accountability model.

Best practices for FDE in agentic AI programs

  • Start with a business workflow, not with a model capability.
  • Define measurable outcomes such as cycle time reduction, error reduction, risk reduction or user productivity improvement.
  • Use explicit orchestration for multi-step processes rather than relying on open-ended agent behavior.
  • Separate deterministic logic from probabilistic reasoning wherever possible.
  • Apply least-privilege access to every agent, tool, connector and data source.
  • Build human-in-the-loop controls for high-risk, low-confidence or irreversible actions.
  • Create evaluation suites using real examples, edge cases, policy scenarios and adversarial prompts.
  • Instrument every workflow with traces, logs, metrics, cost visibility and audit records.
  • Use AI-assisted coding to accelerate delivery, but review, test and secure generated code before release.
  • Design for transfer of ownership so the client team can operate and extend the system after deployment.

Conclusion

Forward-deployed engineering is becoming central to the Agentic AI era because it solves the problem that models alone cannot solve: making AI work safely, reliably and measurably inside real organizations. Agentic systems introduce autonomy, but autonomy without orchestration becomes risk. They introduce speed, but speed without governance becomes fragility. They introduce new coding possibilities, but AI-generated code without engineering ownership becomes technical debt.

FDE provides the missing bridge. It brings engineering to the field, policy into the runtime, security into the workflow and operational discipline into agent design. Whether the implementation uses LangGraph, Microsoft AutoGen, Microsoft Agent Framework, CrewAI or another orchestration stack, the core principle remains the same: enterprise Agentic AI must be co-designed with the business, governed by architecture, secured by default and operated as a living system. That is the practical promise of forward-deployed engineering. Everest Group’s Innovation Watch on Agentic AI Products further reinforces this direction by describing agentic AI as the next stage of automation, where autonomy, adaptability and decision-making are embedded into systems to improve efficiency and responsiveness.

This article was made possible by our partnership with the IASA Chief Architect Forum. The CAF’s purpose is to test, challenge and support the art and science of Business Technology Architecture and its evolution over time as well as grow the influence and leadership of chief architects both inside and outside the profession. The CAF is a leadership community of the IASA, the leading non-profit professional association for business technology architects.

This article is published as part of the Foundry Expert Contributor Network.
Want to join?

  • ✇Security | CIO
  • When it comes to AI, bigger isn’t always better
    There is growing concern about trust in AI as the technology is adopted by more people. Large language models (LLMs) continue to face persistent challenges with hallucinations and inaccurate outputs. LLMs are probabilistic systems trained to give answers, even when the correct answer is unclear or unknowable with the context provided. Humans are more likely to admit that they do not know an answer, especially when there is a financial or reputational consequence at stak
     

When it comes to AI, bigger isn’t always better

28 de Julho de 2026, 10:00

There is growing concern about trust in AI as the technology is adopted by more people. Large language models (LLMs) continue to face persistent challenges with hallucinations and inaccurate outputs.

LLMs are probabilistic systems trained to give answers, even when the correct answer is unclear or unknowable with the context provided. Humans are more likely to admit that they do not know an answer, especially when there is a financial or reputational consequence at stake, while on the other hand, LLMs are designed to act confidently, no matter what.

Most leading models fall within a 20 to 27 percent range of hallucination rate, making it a persistent and unresolved challenge across current AI systems, because it’s not just an architectural problem, it’s a contextual one.

Enterprise AI can’t afford hallucinations

While consumer AI is typically optimized for scale and creativity, enterprise AI must optimize for consistency and precision.

There are many high-stakes industries where getting the answer wrong can have detrimental and long-lasting effects. For example, in healthcare, legal and finance, the margin for error is zero, and one single hallucination can lead to serious consequences. A wrong supplier name, a misread total, or a compliance misstep isn’t a quirky model behaviour, it’s a liability. Typically, in business, it’s not just the one issue that is the concern, it’s the compounding of issues and scale.

In enterprise AI tools, just bolting a general-purpose LLM onto a workflow and hoping for accuracy is a dangerous gamble. Frontier models can change overnight, resulting in a workflow that was 92 percent accurate on Monday but, by Tuesday, produces entirely different results, may be under export controls, or may refuse to process some items. When your product has dependencies on something not designed for the job, you may not get the accuracy you need or the cost you expect because you’re effectively renting your house, and the cost of the rent can change at any time. It’s fast to build with frontier AI models, but you can usually tell when there is no accuracy claim: ‘AI can make mistakes, we may or may not train on your data…’

Before generative AI, we lived in a world of deterministic code – there were bugs, but you could reason over the system. As we move into a world of generative AI and purely probabilistic systems, things are going to behave differently. Looking forward, enterprises are aiming to achieve a balance between the two. A blend of deterministic logic, specialised models, frontier systems and the correct grounding context with human supervision. Building and maintaining this orchestrated symphony at scale—while ensuring absolute trust—is a non-trivial challenge.

Enter SLMs: Faster, cheaper and more accurate

We need to move away from a one-size-fits-all approach to AI, or even a one-model system, and this is where small language models (SLMs) come into play. SLMs are specialised, domain-specific language models designed for a purpose.

By training on narrow, high-quality datasets, these models operate in a world focused on accuracy. Because they leverage highly targeted, niche datasets, compared to the internet and world knowledge which LLMs are trained on, SLMs are inherently leaner, faster and cheaper to run. Consequently, their logic is easier to reason about, allowing them to guarantee a much higher degree of accuracy in their output. Unlike LLMs, they aren’t trying to be clever; they’re trying to be correct.

A report from Gartner found that LLM response accuracy declines when tasks require specific business context. As a result, they predict that by 2027, smaller, context-specific models will see usage volumes at least three times greater than those of general-purpose LLMs. It’s critical to state that the difference lies in the focus and the distribution of data these models have seen.

Early adoption of AI was driven by experimentation and productivity gains. The next step is for AI systems to influence operational decisions, and that shifts the standards required for trust. For CIOs and technology leaders, we need to build systems that operate consistently under real-world conditions – ones that maintain performance over time and will withstand regulatory and customer scrutiny.

Building trusted systems

The most effective enterprise AI tools will combine both LLM and SLM models – using LLMs for orchestration and SLMs for deterministic verification. By using multiple specialist models built for precision, organizations can build AI systems with accuracy they can stand behind and that can run efficiently from a cost perspective.

1. Getting work done right requires AI and humans to work together

AI adoption isn’t a binary choice between models; true operational resilience comes from utilizing the best elements of different architectures. When large and small models are combined, large thinking models can dispatch and reason while smaller models can act and verify, and the whole system can adapt to improve itself.

Models only get better when there are feedback loops through signals and context. On the surface things should look simple and feel like magic, but under the covers, the complexity is often many layers deep.

In financial document processing, for example, large thinking models can handle complex reasoning and user/company/accounting preferences, while the smaller models handle values, tax and line items. The key is that it is not a one-size-fits-all, and it’s not just the models; it’s the embeddings of similarity for processing and the context and guidance that shapes the outcome.

2. Creating best in class

To build a best-in-class system, you need data, compute and a feedback loop. We achieve this by training in-house models on our billions of documents for fast and accurate extraction, then leverage frontier-level models for reasoning. Relying solely on frontier-level models for extraction would tank our accuracy and skyrocket costs. This balance of accuracy comes at the price of generability and reasoning, but leveraging reasoning post-processing provides a top-tier solution, especially when that reasoning is looking to mimic the user’s preferences.

3. The human feedback loop

Human and AI oversight must be built into frontier systems from the beginning, not treated as a fallback when something goes wrong. There are always edge cases that have “it depends” answers – sometimes the answer is we do not know, but without any auditing or sampling, it’s impossible to know how well you did. AI and complex systems can drift, and accuracy is heavily dependent on the distribution of data, so continuous sampling and monitoring are critical components to these systems.

Human users and reviewers in the accounting world are those who are accountable for outcomes, and they provide the signal of feedback to the models and weights. On our systems, we sample over 100,000 documents every month to ensure accuracy is a measurable metric, not a hope.

In our system, when a bookkeeper accepts, edits, or rejects AI suggestions, that data is fed into an evaluation loop too. The result is a continuously improving system that understands the unique financial context of each of the 700,000 SMBs in our system. Through these feedback loops and rigorous evaluations, outputs remain accurate and dependable over time.

The AI hype

The market is flooded with AI tools and demos that trivialise complex systems. Yet many of these crumble when they encounter the complexity and nuances of real-world workflows, leading to a new tagline: “AI can make mistakes”.

This era of AI will not be defined by who is building the biggest model or who has the best demos. It will be defined by those delivering the biggest sustainable impact – the true overnight success stories, 15 years in the making.

This article is published as part of the Foundry Expert Contributor Network.
Want to join?

❌
❌