What Is Actually Working in Agentic AI
Real enterprise deployments, measurable results, and the implementation realities behind vendor case studies
The Agentic AI Economy | Article 4 of 8
Horizon SPI Executive Intelligence Series
By Steven Kiss, MBA | July 2026
Executive Summary
Agentic AI is moving from concept to deployment, but the evidence is still uneven. Some enterprise use cases are producing measurable gains, while others remain closer to experimentation than operational transformation.
This article examines what is actually working in agentic AI today. It looks at real deployment patterns, vendor case studies, reported productivity gains, and the operational conditions that appear to separate successful implementations from superficial adoption.
The central argument is simple: measurable value does not come from deploying agents alone. It comes from selecting the right use cases, connecting them to real workflows, governing their actions, and ensuring the organization is ready to absorb the change.
For executive leaders, the question is no longer whether agentic AI can produce results. The more important question is where those results are credible, repeatable, and governed well enough to scale.
Real deployments. Measurable results. And the implementation realities no vendor will tell you about.
This is the fourth article in the Horizon SPI Executive Intelligence Series: The Agentic AI Economy. Articles 1 through 3 examined why the shift from generative to agentic AI is structurally significant, how agentic systems are built and governed, and which platforms are shaping the competitive infrastructure of the emerging agentic enterprise. This article turns to the operational evidence: what is working, what is failing, and why those outcomes are more predictable than many executives currently recognise.
The conversation about agentic AI has shifted. Twelve months ago, the dominant question in enterprise boardrooms was whether autonomous AI agents were ready for production deployment. That question is now being answered by operational evidence across industries, functions, and organisational sizes.
The more useful question is which deployments are actually delivering, which are quietly failing, and what separates organisations generating measurable returns from those spending budget on agentic AI theatre.
The answer requires looking past vendor case studies. It requires examining the operational evidence, including the failures, with the same clarity executives apply to any significant capital investment.
The evidence now points in two directions at once. Google's 2025 ROI of AI Report found that 74% of executives deploying AI agents in production report achieving ROI within the first year, and 39% report having deployed more than ten agents in production, the cohort reporting the highest productivity returns. PwC research indicates that 79% of organisations now use AI agents in some form, with 66% reporting measurable productivity improvements. At the same time, Gartner predicts that more than 40% of agentic AI initiatives will be cancelled by the end of 2027 because of escalating costs, unclear business value, and inadequate risk controls.
Those findings are not contradictory. They describe a market moving quickly, but unevenly: value is appearing where the operating conditions are strong, while weaker deployments are beginning to expose the cost of poor preparation.

Diagram 1: The Horizon SPI Agentic AI Deployment Reality Map shows where enterprise agentic AI is working, where it remains experimental, and what conditions support credible scaling.
Where Agentic AI Is Generating Measurable Value
Five domains are showing the clearest enterprise returns. They are not necessarily the most glamorous use cases. They are the ones where the work has enough volume, structure, and measurable outcomes for agents to perform reliably. When those conditions are present, the gains tend to appear in speed, cost control, and process consistency.
1. Financial Services: Autonomous Operations at Scale
Financial services has become the leading enterprise environment for agentic AI deployment. The results from major institutions are operationally significant rather than experimental.
JPMorgan Chase is running more than 450 distinct AI use cases in daily production across trading, risk, compliance, fraud detection, and client operations. The bank's AI infrastructure is credited with generating approximately $1.5 billion in value across fraud prevention, personalisation, trading, and operational efficiencies, with a reported 20% gross sales lift in targeted client segments. With a technology budget exceeding $18 billion in 2025 and rising to $19.8 billion in 2026, JPMorgan is not treating AI as an isolated productivity tool. At that scale, agentic capability is becoming part of the bank's operating infrastructure.
Goldman Sachs, working with Anthropic, is deploying autonomous agents for trade accounting, transaction reconciliation, and client onboarding. These are use cases where structured, high-volume decisioning creates direct operational value. CEO David Solomon has publicly stated that he wishes the bank could invest even more in AI.
Citigroup has made its AI tools available to approximately 180,000 employees across more than 80 countries, with CEO Jane Fraser reporting adoption above 70%. The bank's automated code review agents save approximately 100,000 developer hours per week, equivalent to the output of 2,500 developers working a standard 40-hour week. In September 2025, Citi launched an agentic AI pilot for 5,000 colleagues, enabling complex multi-step tasks from a single prompt, with a broader rollout planned. The scale matters, but so does the nature of the work: code review, development support, and multi-step task execution sit close to the bank's operating capacity.
Wells Fargo processed more than 245 million autonomous AI interactions in 2024, an 11x increase from the prior year, with engineers reporting productivity gains of up to 35% through AI-assisted development. That volume matters because it shows agents handling routine, repeatable work at a scale that changes the operating baseline, even when each individual interaction is not strategically complex.
- Executive observation: Across these institutions, the highest-value deployments are not replacing entire functions. They are automating high-volume, structured, rules-intensive workflows within functions while directing human expertise toward judgement-intensive exceptions.
2. Customer Operations: The Most Deployed and Most Instructive Use Case
Customer operations is the most widely deployed agentic AI environment in the enterprise. It has produced some of the strongest successes and some of the most useful failures.
Salesforce's own internal deployment of Agentforce is resolving 85% of customer queries without human involvement across its global operations. One European logistics platform reduced customer support response time from two hours to under 90 seconds. That kind of improvement is not only a cost story. It can change the customer's experience of the company.
Few cases have shaped the enterprise AI conversation as much as Klarna. In early 2024, Klarna deployed an AI customer service agent it claimed was doing the work of 700 full-time employees, coinciding with a reduction in its workforce from 5,500 to approximately 3,400. The numbers were large, but they did not capture the whole consequence. By mid-2025, customer satisfaction had declined, service quality had become erratic, and the company began rehiring human agents. By Q3 2025, the system was credited with the equivalent output of 853 employees and savings of $60 million, but at the cost of institutional knowledge that proved difficult to rebuild and a customer experience Forrester described as 'the poster child for bad AI deployment.
CEO Sebastian Siemiatkowski acknowledged publicly what the numbers obscured: "Cost unfortunately became too large a factor when making these decisions, and you end up having lower quality." Klarna subsequently shifted to a blended model, where AI handles high-volume simple interactions and human agents address complex, emotionally sensitive, or high-value customer situations.
- Executive observation: The Klarna case is not best understood as a failure of agentic AI technology. It is a failure of deployment philosophy. Organisations that deploy agents mainly to minimize cost tend to produce weaker outcomes than those that design for customer value first and treat efficiency as a consequence.
3. Supply Chain: From Reactive Scheduling to Autonomous Optimisation
Supply chain management has become one of the clearest environments for agentic AI value. It combines high complexity, real-time data requirements, and decisions that must be made continuously across interdependent variables.
Supply chain is less visible than customer service, but the deployment logic is straightforward. Agentic systems are being used for demand forecasting, inventory replenishment, production balancing, and supplier disruption response. Earlier AI models produced recommendations for human review. More advanced agentic deployments execute decisions directly within predefined limits, adjusting production schedules, triggering sourcing actions, and rerouting logistics without waiting for human authorization at each step.
The market data points in the same direction. Grand View Research projects the global AI-powered supply chain planning market will reach $51.1 billion by 2030, growing from $13.9 billion in 2026. A January 2026 Gartner poll found that 19% of organisations had made significant investments in agentic supply chain AI, with an additional 42% making conservative investments. The more important point is operational: supply chain decisions repeat constantly, depend on live data, and become expensive when delayed.
- Executive observation: The strongest supply chain results appear where organisations identify specific decision nodes, such as replenishment triggers, capacity allocation, and supplier switching, where agent speed and consistency create operational advantage. Human oversight then sits at the boundary conditions where judgment still matters.
4. Legal, Risk, and Compliance: High-Value, Low-Profile Deployments
Legal, risk, and compliance functions are among the least publicized but highest-value deployment environments in large organisations. High document volume, structured regulatory logic, and the cost of manual expert time create conditions where autonomous agents can deliver disproportionate returns.
Six use case categories are seeing significant enterprise traction: autonomous regulatory monitoring, contract lifecycle management, eDiscovery at litigation scale, compliance training personalization, sanctions screening, and litigation risk analysis. In each category, agents operate continuously across document volumes that no human team could process at equivalent speed or cost.
For regulated industries such as financial services, healthcare, pharmaceuticals, and energy, the strategic significance is material. Regulatory compliance is not a function where speed alone creates competitive advantage. It is a function where failure creates serious institutional risk. Agentic systems that reduce the probability of missed regulatory changes or misclassified compliance events are not simply productivity tools. They are risk management infrastructures.
- Executive observation: Legal and compliance create a strong ROI case for agentic AI but also some of the highest governance requirements. Successful deployments treat agents as participants in a governed compliance function, with audit trails, escalation protocols, and human review thresholds built into the architecture from day one.
5. Software Development: The Fastest-Scaling Enterprise Deployment
Autonomous coding agents have achieved the fastest enterprise adoption curve of any agentic AI category. The productivity data from organisations that have deployed them at scale is consistent across industries and organisational sizes.
Wells Fargo engineers report up to 35% productivity gains through AI-assisted development. Morgan Stanley's DevGen AI tool, built on OpenAI models and unveiled in early 2025, saved developers more than 280,000 hours, equivalent to approximately 11,600 working days, by mid-2025. Anthropic's Claude Code moved into a leading position among VS Code AI extensions in early 2026 and is reported to serve more than 300,000 business customers, including several of the largest global enterprises. The common thread is the feedback loop: developers can see quickly whether an agent has helped or created rework, which makes this category easier to measure and improve than many enterprise AI deployments.
The deeper implication is what this technology does to enterprise software economics. If software development productivity increases by 30-40% through agentic assistance, the effective cost of building and maintaining custom enterprise software declines materially. For organisations with large custom software portfolios, that creates a compounding advantage in their capacity to build and deploy agentic capabilities across their own operations.
- Executive observation: Developer productivity is often the fastest-returning agentic AI use case because the feedback loop is short. Code either works or it does not. For organisations where technology is a strategic function, this is one of the highest-confidence starting points for enterprise agentic investment.
What Vendors Will Not Tell You: The Implementation Realities
Most Agentic AI Failures Are Not Technology Failures
Analysis of enterprise agentic AI failures in 2025 points to a consistent root cause: capable architecture built on a brittle operational foundation. Agents failed not because the models were inadequate but because of latency spikes under production load, stale data contradicting live inventory, governance gaps that allowed agents to access data they should not have seen, and organisational ambiguity about accountability when something went wrong.
Gartner identified three failure modes that account for the majority of cancelled projects: escalating integration costs with legacy systems, unclear business value, and inadequate risk controls. The last point is especially important. A large share of organisations have deployed or are scaling AI agents, but far fewer have comprehensive security controls designed for agentic systems.
The Governance Gap Is the Defining Enterprise Risk
Deloitte's 2026 State of AI in the Enterprise report, based on a survey of 3,235 IT and business leaders across 24 countries, finds that only 21% of enterprises have mature governance in place to manage agentic AI risk, making this one of the most important findings in enterprise agentic AI this year.
Approximately 80% of organisations currently lack clear agent decision boundaries, real-time monitoring systems that flag anomalies, and audit trails that capture the full chain of agent actions. The consequence is not hypothetical. Poorly governed agents can access sensitive customer data, make decisions on behalf of employees, and execute actions across the enterprise technology stack with limited oversight.
Deloitte puts the issue directly: "Rushing to deploy AI agents widely before establishing these governance foundations could expose organisations to significant and potentially costly risks—while also likely negating the competitive advantage that AI agents otherwise could afford."
Agent Washing Is Distorting Enterprise Investment Decisions
Gartner estimates that, of the thousands of vendors claiming agentic AI capabilities, only approximately 130 are delivering genuine agentic solutions. The rest are engaging in what practitioners now call agent washing: rebranding existing chatbots, RPA tools, and AI assistants without substantive new capabilities.
The diagnostic is straightforward. If a system can only respond to queries but cannot autonomously plan, execute, and adapt across multiple steps, it is not an agent. When those systems disappoint, executives may blame agentic AI rather than the vendor that oversold them. Organisations making vendor selections based on marketing claims rather than production deployment evidence carry the highest exposure to this risk.
Payback Timelines Vary Significantly by Deployment Approach
Deployment approach is one of the largest variables in determining investment return. Buy-and-configure deployments, using platforms such as Salesforce Agentforce, Microsoft Copilot Studio, or AWS Bedrock, typically return investment in 8 to 18 months. Custom-built agentic systems often take 18 to 36 months to reach payback.
Organisations that start with configured platforms in high-volume structured workflows tend to generate faster returns and build the organisational capability required to justify custom development later. Reversing that sequence, by building custom before establishing production discipline, remains one of the more expensive mistakes in enterprise agentic deployment.
What the Successful 60% Do Differently
The organisations achieving durable returns from agentic AI share five operational characteristics. None are primarily about having the largest budget or the most advanced model.
They choose use cases by decision structure, not by ambition. The strongest deployments focus on complex, dynamic, multi-step processes where agent capabilities create distinctive value, rather than on simple tasks that conventional automation can handle.
They build guardrails before they scale. Human-in-the-loop checkpoints for high-stakes decisions, comprehensive audit trails, and role-based access controls are architectural requirements from day one, not governance additions bolted on after deployment.
They measure outcomes that matter. The useful metrics are not only cost-per-transaction. They include time-to-resolution, decision accuracy, process completion rate, and customer outcome quality.
They design for human-agent collaboration, not human replacement. The organisations with the strongest long-term results, including JPMorgan, Goldman Sachs, and Citigroup, are not using agents to pretend human judgment is no longer needed. They are using agents to remove repetitive work from expert teams, so people can spend more time on exceptions, interpretation, and decisions that still require accountability.
They treat agentic AI as an organizational capability, not technology installation. The most important investments are not only in models or platforms. They are in data quality, integration architecture, governance structures, and ownership models that determine whether agents can operate reliably in production.
The trust gap remains the defining constraint on enterprise agentic adoption. According to Workato / Harvard Business Review Analytic Services research, 86% of organisations plan to increase agentic AI investment, but only 6% trust AI agents to autonomously handle core end-to-end business processes. That gap between investment intent and operational trust will determine which organisations scale from pilot to production and which remain stuck in the experimental stage.
The Executive Perspective
The pattern is becoming clear. Agentic AI can produce material value when organisations deploy it with operational discipline. It also creates expensive lessons when teams move with vendor enthusiasm but without governance maturity.
By the end of 2026, Gartner projects that 40% of enterprise applications will embed task-specific AI agents, up from less than 5% at the start of the year. That is not a gradual adoption curve. It is a rapid restructuring of the enterprise software stack within a single fiscal year. The organisations that lead this transition will not necessarily be those that moved fastest. They will be those that moved with the most discipline.
The question is not whether agentic AI is useful. It is whether your organization has built the foundation that determines deployment quality: data readiness, governance architecture, human-agent collaboration design, outcome measurement, and clear ownership.
In which part of your organization does the gap between AI investment intent and operational governance readiness pose the greatest strategic risk, and who is responsible for closing it?
______________________________________________________________________________________________________________
Selected references
• Google Cloud --- ROI of AI Report 2025: Enterprise Agent Deployment and Return Benchmarks
• PwC --- AI Agent Survey 2026: Enterprise Adoption, Productivity, and Trust Data
• Gartner --- Predicts 2026: Over 40% of Agentic AI Projects Will Be Cancelled by End of 2027
• McKinsey Global Institute --- Enterprise Agentic AI Deployment and Productivity Analysis 2026
• JPMorgan Chase --- 2025 Annual Report and Technology Budget Disclosure (2026)
• Reuters --- JPMorgan Says AI Helped Boost Sales, Add Clients in Market Turmoil (May 2025)
• Citigroup --- Q4 2025 and Q1 2026 Earnings Calls: AI Adoption and Productivity Disclosure
• Business Insider --- Morgan Stanley's DevGen AI Tool Saves Developers 280,000+ Hours (July 2025)
• Forbes --- When AI Customer Service Goes Wrong --- And How to Get It Right (April 2026)
• Fast Company --- Klarna Tried to Replace Its Workforce with AI (January 2026)
• Forrester Research --- Klarna: Poster Child for Bad AI Deployment (November 2025)
• Swift Water Co. --- Agentic AI in Legal: 6 Real Examples for In-House Teams (2026)
• CrewAI --- 2026 Enterprise Agentic AI Survey: 500 Executives on Production Deployment and Expansion
© 2026 Horizon SPI. All rights reserved.
Executive Intelligence Series | horizonspi.com
Article 5 will examine the governance imperative: how leading organisations are building the oversight frameworks, risk controls, and accountability structures required to scale agentic AI with confidence.
