Business+AI Blog

The ROI of AI Agents: Hard Numbers from Real Deployments

July 18, 2026
AI Consulting
The ROI of AI Agents: Hard Numbers from Real Deployments
AI agents are delivering measurable ROI—but only for companies that deploy them right. See the real numbers from Klarna, JPMorgan, General Mills, and more.

Table Of Contents

  1. The Paradox That Won't Go Away
  2. What the Aggregate Data Actually Shows
  3. Hard Numbers by Industry: Real Deployments, Real Returns
  4. Why the Bottom Quartile Still Fails
  5. The Four Conditions That Separate Winners from the Rest
  6. What This Means for Your Business

The ROI of AI Agents: Hard Numbers from Real Deployments

Every boardroom in Asia and beyond has heard the pitch: AI agents will transform your operations, compress costs, and unlock revenue streams you couldn't access before. Most executives have sat through the decks. Many have approved the pilots. And a significant number are still waiting for the numbers to appear in their P&L.

The frustrating truth is that the numbers are appearing—just not evenly. A small cohort of organizations is recording returns that are hard to argue with: Klarna's customer service AI agent handled the equivalent workload of 853 full-time employees and saved $60 million. General Mills deployed an AI-driven supply chain system that assessed over 5,000 daily shipments and generated more than $20 million in savings since fiscal 2024. A healthcare system in New Jersey saved its clinicians 66 minutes per day, per doctor. These are not projections or aspirational case studies. They are documented production outcomes.

But those wins sit alongside a sobering counterweight: MIT research found that 95% of generative AI pilots fail to deliver rapid, measurable revenue acceleration. S&P Global data shows that 42% of companies scrapped most of their AI initiatives in 2025. The gap between the winners and everyone else is not about access to technology—it is about how they deploy it.

This article cuts through the noise to examine what the hard numbers from real AI agent deployments actually say, which industries are generating the clearest returns, and what specifically separates the organizations that are winning from those stuck in pilot purgatory.

Business+AI Research Brief

The ROI of AI Agents:
Hard Numbers from Real Deployments

AI agents are delivering measurable ROI — but only for companies that deploy them right. Here are the real numbers.

Aggregate ROI Data

Avg. Agentic ROI
171%
U.S. enterprises reach ~192%
Forrester Study Avg.
540%
ROI within 18 months across 287 deployments
Median Payback
7.3mo
74% achieved ROI within year 1
Time Saved / Week
7.2hrs
Per knowledge worker using AI agents

Real Deployments, Real Returns

Financial Services
Klarna Customer Service AI
$60M in savings
  • 853 FTE equivalent workload
  • 2.3M conversations handled
  • 11 min → <2 min resolution
Supply Chain
General Mills Supply Chain AI
$20M+ in savings
  • 5,000+ daily shipments assessed
  • Autonomous routing optimization
  • Up to 95% forecast accuracy
Healthcare
AtlantiCare Clinical Documentation
66 min saved/doctor/day
  • 80% provider adoption achieved
  • 42% reduction in doc time
  • $3.20 return per $1 invested
Banking / Dev
JPMorgan + Morgan Stanley AI
280K dev hours saved
  • 450+ AI use cases in production
  • 9M lines of legacy code reviewed
  • >50% effort reduction in dev

The Paradox: Wide Adoption, Uneven Returns

Companies using generative AI (at least one function) 78%
Companies with no material P&L contribution from AI 80%
AI pilots failing to deliver rapid revenue acceleration (MIT) 95%
Top-quartile deployments exceeding 800% ROI 25%

Why the Bottom Quartile Still Fails

📊No Baseline Metrics
ROI that can't be attributed can't be reported to boards or used to justify scaling.
🧪Curated Pilot Data
Clean test data doesn't survive production volume — performance collapses when real variability arrives.
💸Misaligned Budgets
>50% of gen AI budgets go to sales tools. Biggest ROI is in back-office and operational workflows.
🏗️No Change Management
Only 37% of organizations invest significantly in change management alongside AI (Deloitte).
🔨Building vs. Buying
Vendor tools succeed ~67% of the time vs. ~33% for internal builds (MIT research).
📉Poor Data Quality
Gartner: 85% of all AI projects fail due to poor data quality — the #1 post-launch failure cause.

4 Conditions That Separate Winners

1
P&L-Linked KPIs Defined First
Every high-performing deployment defines measurable business metrics before a single line of agent code runs.
2
Scope Narrowly, Go Deep
Solve one specific painful process with high precision. 60–90 day scoped pilots, then scale based on real data.
3
Data Infrastructure as Prerequisite
Data readiness is not an afterthought — it's the single most common source of post-launch shortfall.
4
Genuine Stakeholder Buy-In
Teams whose workflows change must participate in design — not just be informed. Culture precedes technology.
The Competitive Gap Is Already Here
more likely to scale AI
(AI high performers vs. peers)
higher revenue growth
(AI-native vs. laggards)
productivity growth uplift
(most AI-exposed industries)

These are not future projections — they are current performance gaps between companies at different stages of the same transition. (Sources: McKinsey, BCG, PwC)

Business+AI
Singapore's AI ROI Ecosystem for Executives

The Paradox That Won't Go Away {#the-paradox}

By any measure of adoption, enterprise AI should be delivering widespread economic results by now. More than 78% of companies are using generative AI in at least one business function, according to McKinsey's Global Survey on AI—a figure that nearly doubled from 55% the year prior. Investment is accelerating sharply: global enterprise AI spending reached roughly $37 billion in 2025, more than triple the 2024 figure. And yet, more than 80% of companies still report no material contribution to earnings from their AI initiatives.

This is what researchers have started calling the "gen AI paradox," and it has a specific structural cause. The vast majority of enterprise AI deployments to date have been horizontal—company-wide copilots, internal chatbots, productivity assistants. These tools are easy to activate and require minimal workflow redesign, which is exactly why they get deployed first. The problem is that their benefits are diffuse: a few minutes saved per employee, spread thinly across thousands of people, rarely aggregates into anything that shows up on a balance sheet.

The real opportunity has always been in vertical deployments—AI embedded into specific, high-value business processes like credit risk assessment, claims processing, supply chain orchestration, or clinical documentation. These use cases carry direct P&L exposure. But they are also harder to build, require deeper integration, and demand genuine process redesign. McKinsey's research found that fewer than 10% of vertical use cases ever make it past the pilot stage. The technology is not the bottleneck. The organization is.

AI agents represent the clearest path out of this paradox. Unlike passive chatbots or copilots that respond only when prompted, agents can understand goals, plan multi-step tasks, interact with enterprise systems, and execute actions autonomously—with minimal human intervention. They bring memory, orchestration, and real-time adaptability to workflows that previously required constant human coordination. And when they are deployed correctly, the ROI is not marginal. It is transformational.


What the Aggregate Data Actually Shows {#aggregate-data}

The macro-level numbers on AI agent ROI are striking, though they come with an important caveat about distribution.

Companies report an average ROI of 171% from agentic AI deployments, with U.S. enterprises specifically achieving around 192%—exceeding traditional automation ROI by approximately three times. Google Cloud's 2025 ROI of AI Report, surveying executives with production AI agent deployments, found that 74% achieved ROI within the first year, and 39% saw productivity at least double in their organizations. A Forrester study analyzing 287 enterprise AI agent deployments across 14 industries found an average ROI of 540% within 18 months of production deployment, with a median payback period of 7.3 months.

Those headline numbers, however, mask a wide distribution. The top quartile of deployments exceeded 800% returns, while the bottom quartile saw returns below 200%. That spread is not random—it reflects specific, predictable differences in how organizations structure and execute their deployments. McKinsey's 2025 State of AI research found that AI high performers are three times more likely to be scaling agents across multiple business functions compared to their peers. BCG research quantifies the downstream consequence: AI-native companies achieve five times higher revenue growth and three times greater cost reductions than competitors slower to integrate AI into their core operations.

The aggregate data also reveals a meaningful shift in time-to-value. Across major datasets from McKinsey, Salesforce, and Slack, knowledge workers using AI agents are saving between 5.9 and 7.2 hours per week. That is not a rounding error—it represents roughly 15–18% of a standard working week returned to higher-value work, at scale, across an entire organization.


Hard Numbers by Industry: Real Deployments, Real Returns {#hard-numbers}

Aggregate statistics tell part of the story. Specific deployments tell the rest.

Financial Services {#financial-services}

Financial services leads enterprise AI agent adoption, and the returns reflect that commitment. JPMorgan Chase runs more than 450 AI use cases in production daily, allocating $18 billion annually to technology—a figure that includes AI agents generating investment banking presentations in 30 seconds, compared to the hours junior analysts previously spent on the same work. Morgan Stanley's DevGen.AI code review agent reviewed over 9 million lines of legacy code, saving approximately 280,000 developer hours across a team of 15,000.

Klarna's deployment is perhaps the most cited case in enterprise AI. In early 2024, its customer service AI agent handled roughly two-thirds of all incoming support chats in its first month, managing 2.3 million conversations. Average resolution time dropped from approximately 11 minutes to under 2 minutes. By Q3 2025, the cumulative impact equated to the workload of 853 full-time agents and $60 million in savings. The company also reported a 40% reduction in cost per transaction since Q1 2023 and a 25% drop in repeat inquiries—a customer experience metric, not just an efficiency one.

In credit risk, McKinsey's work with a retail bank demonstrated what is possible when agents are applied to knowledge-intensive financial processes. Relationship managers had been spending weeks manually extracting data from over ten sources to write credit-risk memos. An agentic workflow replaced the manual extraction, drafted memo sections with confidence scores, and shifted the analyst's role to oversight and exception handling. The result: a potential 20 to 60% increase in productivity, including a 30% improvement in credit turnaround time.

Customer Service & Operations {#customer-service}

Customer service is consistently the fastest category to demonstrate AI agent ROI, for a simple reason: the baseline metrics already exist. Average handle time, first-contact resolution rate, cost per interaction, and CSAT scores are all pre-existing measures that make an agent's impact visible within weeks rather than quarters.

ServiceNow's internal deployments reported deflection rates as high as 54% on common service requests, a 14% increase in employee self-service, and annualized savings of roughly $5.5 million from case and incident avoidance. Salesforce Agentforce users reported ROI in as little as two weeks. When AI agents handle customer service at the third level of process reinvention—proactively detecting issues, initiating resolution steps, and communicating directly with customers—McKinsey's analysis suggests up to 80% of common incidents can be resolved autonomously, with a 60 to 90% reduction in time to resolution.

In insurance claims processing, a mid-to-large insurer deployed a seven-agent system covering the full claims lifecycle: intake, policy verification, fraud detection, damage assessment, reserve setting, and settlement communication. The measured outcomes showed a 30% reduction in operational costs and a material compression in claims cycle time—with secondary effects on customer retention.

Supply Chain & Manufacturing {#supply-chain}

Supply chain is where AI agents move from efficiency into resilience. General Mills deployed an AI-driven supply chain optimization system that autonomously assesses more than 5,000 daily shipments, evaluating routing, timing, and vendor performance, and flagging exceptions for human review. The system has produced over $20 million in savings since fiscal 2024.

Across enterprise supply chain deployments in demand forecasting, measured outcomes show 20 to 40% reductions in forecast error and average inventory reductions of 31%, with some implementations reaching 95% forecast accuracy. Autonomous procurement agents—deployed on platforms including IBM Watsonx and SAP Business AI—monitoring supplier performance and detecting risk signals show measured outcomes of 15% reductions in direct procurement costs and 20% improvements in supply chain resilience scores. In manufacturing, 61% of executives report decreased costs as a direct result of AI in supply chain operations.

Healthcare {#healthcare}

Healthcare organizations report a $3.20 return for every $1 invested in AI within 14 months—a figure that reflects how well-defined the productivity constraints are in clinical environments. AtlantiCare, a regional healthcare system in New Jersey, deployed an AI documentation agent that listens to consultations, generates structured clinical notes, and pre-populates electronic health record fields. The measured outcomes were specific: 80% provider adoption within the first months, a 42% reduction in documentation time, and 66 minutes saved per clinician per day. That time went back to direct patient care.

What distinguished this as a genuinely agentic deployment—rather than a simple transcription tool—was the agent's ability to structure notes according to clinical coding requirements, flag missing information, and surface relevant prior visit data. It closed a documentation loop that previously required physician attention at every step.

Software Development {#software-dev}

In software, the McKinsey case study of a large bank modernizing a legacy core system—consisting of 400 pieces of software with an original budget of over $600 million—demonstrates the compound effect of deploying agent squads in place of manual coding teams. Human workers moved into supervisory roles overseeing agents that documented legacy systems, wrote new code, reviewed each other's output, and integrated features for testing. The result was a more than 50% reduction in time and effort among early adopter teams. Microsoft's small and mid-sized business data shows up to 353% ROI from Copilot-based workflows in knowledge work more broadly.


Why the Bottom Quartile Still Fails {#why-fail}

With returns of this magnitude available, why do so many deployments still underperform? The data is consistent across multiple research bodies, and the answer is not the technology.

MIT's NANDA research found that 95% of generative AI pilots fail to deliver rapid revenue acceleration. S&P Global data shows 42% of companies scrapped most of their AI initiatives in 2025—more than double the abandonment rate from the prior year. Gartner predicts that over 40% of agentic AI projects will be cancelled by 2027 due to escalating costs, unclear business value, and inadequate risk controls. RAND Corporation research found that 80.3% of enterprise AI projects fail to deliver promised business value, with only 19.7% fully delivering on their business case.

The recurring causes are organizational, not technical:

  • No baseline metrics before deployment. Without a documented baseline, there is no way to prove what the agent changed. ROI that cannot be attributed cannot be reported to boards or used to justify scaling.
  • Piloting on curated data that doesn't exist in production. Pilots often run on clean, hand-picked datasets. When they move to production volume with real-world variability, performance collapses—and with it, executive confidence.
  • Misaligned resource allocation. MIT found that more than half of generative AI budgets are devoted to sales and marketing tools, yet the biggest ROI concentrates in back-office automation and operational workflows.
  • Organizational resistance without change management. Google Cloud's DORA 2025 report found that 70% of AI transformation value comes from people, organizations, and processes. Yet Deloitte's 2026 State of AI survey found only 37% of organizations have invested significantly in change management alongside AI deployments.
  • Building instead of buying. MIT's research found that purchasing AI tools from specialized vendors succeeded approximately 67% of the time, while internal builds succeeded only one-third as often.

The companies stuck in pilot purgatory typically launched AI initiatives as technology experiments rather than business transformations. They optimized for demo performance, not production performance. And they never built the measurement layer that makes ROI attributable rather than estimated.

If you're working through these challenges inside your own organization, Business+AI's consulting practice partners with executives to structure AI deployments around measurable business outcomes—not pilot metrics.


The Four Conditions That Separate Winners from the Rest {#four-conditions}

Across the deployments that have consistently delivered strong returns, four conditions appear reliably in the organizations that succeed.

1. They define P&L-linked KPIs before a line of agent code runs. Every high-performing deployment in the research literature shares this characteristic. Customer service pilots track handle time and resolution rates. Supply chain pilots track inventory levels, logistics costs, and service levels. Credit risk pilots track turnaround time and analyst productivity. The KPI exists before the agent does. Organizations that deploy first and measure later rarely recover from the confusion that follows.

2. They scope narrowly and deeply, not broadly and shallowly. The most profitable deployments solve a single, specific, painful process problem with high precision. When agents are forced to navigate complex, multi-layered workflows from day one, failure rates escalate sharply. The winning pattern is a scoped pilot on one workflow or one business unit, rigorous measurement over 60 to 90 days, and then a scaling decision based on real data. This is also why industry-specific vertical agents consistently outperform general-purpose horizontal tools—the specificity of the task is what makes the ROI attributable.

3. They treat data infrastructure as a prerequisite, not an afterthought. Gartner's research found that 85% of all AI projects fail due to poor data quality. The supply chain deployments at General Mills and the insurance claims system described above were both built on data infrastructure that matured over years before autonomous operation was trusted at full scale. Underestimating the data preparation phase is the single most common source of post-launch performance shortfall.

4. They build genuine stakeholder buy-in, not just executive sign-off. The teams whose workflows change need to understand the rationale, participate in the design, and see early results. Deployments where affected teams were informed rather than involved consistently experience higher early resistance and longer adoption timelines. This is a cultural and change management challenge before it is ever a technical one.

Business+AI's workshops and masterclasses are designed to build exactly this kind of organizational readiness—equipping teams at every level to design, evaluate, and adopt AI agent deployments that deliver real results, not just impressive demos.


What This Means for Your Business {#what-this-means}

The AI agent ROI data presents a clear and uncomfortable message for executives: the returns are real, they are significant, and they are not evenly distributed. The gap between organizations actively scaling agents in production and those still cycling through pilots is widening—and it is starting to show up in competitive performance.

McKinsey's 2025 State of AI research found that AI high performers are three times more likely to be scaling agents across multiple business functions. BCG research shows that AI-native companies achieve five times higher revenue growth and three times greater cost reductions compared to peers. PwC's 2025 AI Jobs Barometer found that since 2022, productivity growth in industries most exposed to AI has nearly quadrupled. These are not future projections—they are current performance gaps between companies at different stages of the same transition.

For business leaders in Asia, the opportunity is particularly pointed. Singapore, India, and Japan are among the leading APAC markets in AI agent experimentation, particularly in e-commerce and customer support. The regulatory environment is evolving to support responsible deployment—Singapore's Model AI Governance Framework for Agentic AI Systems provides a practical foundation that many organizations can build on.

The question is no longer whether AI agents deliver ROI. The documented evidence is too extensive to dispute. The question is whether your organization has the deployment discipline—the defined metrics, the scoped use cases, the data infrastructure, and the change management investment—to join the cohort that is extracting that ROI today.

Business+AI's annual Business+AI Forum brings together executives, consultants, and solution vendors who have navigated exactly this transition. If you want to understand how organizations in your region are moving from pilot to production—with the specific numbers and deployment conditions that made it work—this is where those conversations happen.

The Numbers Have Landed. The Question Is Execution.

The era of debating whether AI agents can deliver meaningful business value is effectively over. Klarna's $60 million in savings, General Mills' $20 million in supply chain gains, AtlantiCare's 66 minutes saved per clinician per day, and the 540% average ROI across 287 enterprise deployments in Forrester's research are not projections. They are the documented outputs of production systems running at scale.

What the data also makes clear is that these results are not a consequence of deploying the most sophisticated technology. They are a consequence of deploying the right technology against the right problem, with the right measurement framework, on a foundation of quality data and genuine organizational commitment. The 95% of pilots that fail to deliver P&L impact are not failing because agents don't work. They are failing because the conditions for success were never built.

For executives who want to move past the paradox and into the group that is actually realizing these returns, the path is not more experimentation. It is more structured deployment—narrower scope, clearer baselines, deeper process integration, and the organizational infrastructure to sustain performance past go-live.

That transition is exactly what Business+AI is built to support.


Ready to Move from Pilot to Production?

Business+AI is Singapore's leading ecosystem for executives who want to turn AI ambition into measurable business performance. Whether you need strategic consulting to structure your first production deployment, hands-on workshops to build team capability, or access to a peer network of leaders who have already navigated this transition, we have the resources to accelerate your path to ROI.

Join the Business+AI membership community today →

Get access to expert-led masterclasses, live workshops, our annual Business+AI Forum, and a network of executives, consultants, and solution vendors who are deploying AI agents that actually deliver.