Select Page

How IBM Saved $3.5 Billion With AI Agents in Two Years

By Chris Linus

July 8, 2026

The image tells how IBM saved $3.5 billion

On January 30, 2025, IBM’s Chief Financial Officer James Kavanaugh was on a call with Wall Street analysts discussing the company’s fourth-quarter results. The numbers were fine. Consulting revenue was under pressure. Software was growing. Infrastructure was cycling through a product transition. Standard earnings-call fare.

Then he said something that wasn’t standard at all.

The Register covered it the same day: “We exited 2024 at $3.5 billion of annual run rate savings,” Kavanaugh told analysts, “and we continue to see these efforts play out in our margin performance this quarter.”

That was it. One sentence, buried in the back half of an earnings call, describing what IBM had achieved by deploying AI agents across its own operations. No press release fanfare. No product launch event. Just a CFO telling investors, under the legal obligations of a public earnings disclosure, that the company had saved $3.5 billion a year by building and running AI agents on itself.

What followed that sentence is the detail of how IBM actually did it, which functions, which agents, which results put together; then becoming one of the clearest and most verifiable roadmaps the enterprise world currently has for deploying AI at scale.

So, What Did IBM Actually Do?

IBM deployed AI agents across more than 70 of its own business functions — HR, IT support, supply chain, finance, sales, and software development — in a programme it calls “Client Zero.” The idea is simple: before IBM sells AI transformation to clients, it has to prove it on itself. By the time CFO Kavanaugh made that earnings-call statement in January 2025, IBM had reached an annualised run-rate productivity saving of $3.5 billion against a $2 billion target it had set two years earlier. IBM Korea CTO Ji-eun Lee confirmed the same figure at an April 2025 press conference in Seoul, adding the detail that AI had been applied across more than 70 business areas.

One important framing note before the detail: this is a productivity savings figure stated on a public earnings call, not independently audited in the same way revenue is. What makes it credible and worth examining closely is that it reconciles with IBM’s segment margins and free cash flow across consecutive earnings reports, and is supported by specific, named, measurable results at the business-function level that IBM has published in detail on its case study pages. We’ll get into those specifics now.

Why IBM Is a Meaningful Test Case

IBM is not a startup. It is not a venture-backed AI company with a financial incentive to make AI look impressive. At the time this programme was running, it employed over 270,000 people globally, operated in more than 170 countries, and processed the kind of transaction volume and operational complexity that breaks pilot programmes which looked good at smaller scale.

Most enterprise AI experiments run in controlled sandboxes. IBM ran these agents in production, on live business processes, serving its own workforce and supply chain. The HR agent was handling real employee payroll queries. The IT agent was resolving real support tickets. The supply chain agents were managing real supplier commitments across 2,000+ suppliers. When the CFO reports the savings in an earnings call, the margin performance he says they’re driving is real margin — visible in the same financial statements IBM’s investors, analysts, and auditors review.

That context is worth establishing before the numbers, because it changes what the numbers mean.

The Three Agents That Did the Heavy Lifting

IBM deployed agents across a wide span of functions, but three areas produced the most documented, measurable results.

AskHR — Human Resources

IBM’s AskHR agent began in 2017 as a basic HR chatbot and has been systematically expanded since, most recently upgraded to a fully agentic system on IBM’s watsonx Orchestrate platform in 2025. It now automates more than 80 HR tasks — vacation requests, payroll queries, employee transfers, compensation changes, organisational updates — and handles them across multiple languages for IBM’s global workforce.

The numbers IBM has published on AskHR are specific and internally consistent across multiple sources. In 2024, AskHR handled more than 11.5 million employee interactions, containing 94% of them within the platform — meaning only 6% needed to be escalated to a human HR advisor. Since the system launched, IBM’s HR team has seen a 75% reduction in support tickets compared to historical levels. The HR operating budget has fallen by 40% over four years. Adoption among managers sits at 99%. The Net Promoter Score — an internal measure of how employees feel about the system — has gone from -35 at launch to +74, a bigger swing than most enterprise software products achieve in their entire lifetimes.

IBM’s own Think Insights piece from February 2026 directly credits AskHR’s contribution to the $3.5 billion result: “Our HR team is proud that our Client Zero work contributed to the USD 3.5 billion in productivity savings (against a USD 2 billion target) that IBM realized in 2024.”

AskIT — IT Support

Launched in May 2023, IBM’s AskIT agent was built to reduce the volume of IT support calls, chats, and tickets that IBM’s workforce generates — an estimated 785,000 tickets annually before the agent launched.

IBM’s published results for AskIT: support tickets fell 56% from 2022 to 2024; calls and chats dropped 74% since launch; 86% of queries are now resolved by AI agents. The system delivered an initial $18 million cost reduction and now yields ongoing annual cost avoidance. Employee satisfaction with IT support sits at 91.6% CSAT, up 11.6 points since launch.

The comparison to AskHR’s early NPS (-35 at launch) is worth noting: AskIT also had a difficult early period before workflows were refined. The lesson IBM draws from both is that initial user experience scores are not the right signal to watch — containment rates and cost trajectory are.

Supply Chain

IBM’s supply chain spans more than 2,000 suppliers and operations in over 170 countries. AI agents were deployed to handle complex, multi-step tasks: identifying risks, validating supplier commitments, initiating corrective actions, managing invoice review, and automating procurement workflows.

IBM’s Client Zero page reports: approximately 30% logistics cost reduction over three years since 2022; 26,000 hours saved annually in procurement; 15.6% productivity savings in supply chain hardware costs. IBM’s agentic AI and supply chain piece puts the total at $361 million in supply chain savings over three years. In finance, AI agents now achieve over 95% first-pass invoice review accuracy across 18 of 19 invoice formats tested.

FURTHER READING

➤ Top 10: AI Adopters (Hyperscalers) in the World

How IBM Decided What to Automate First

The result that surprises most people when they look at IBM’s programme is not the scale but the discipline. IBM did not start by asking “what can AI do?” It started by asking a harder question: does this process even need to exist?

The principle IBM’s HR team describes as their guiding framework is: eliminate first, simplify second, automate third. Before AskHR touched a single HR process, the team reviewed whether that process should exist at all. A practical example is that which IBM’s VP of HR Technology has spoken about publicly: IBM had a specific leave policy for employees competing in the Olympic Games — a detailed, multi-step policy for a scenario affecting perhaps a handful of employees every four years. It was eliminated entirely. IBM also consolidated more than 25 distinct types of leave policies down to one.

Only after that elimination and simplification pass did automation begin. The reason this matters is practical: automating a broken or unnecessary process just makes the broken process faster. IBM’s large savings came not from AI making complex old processes efficient, but from AI being applied to processes that had been redesigned to be simple enough to automate reliably.

This is the core discipline behind the “Client Zero” label. IBM is not just a large company that deployed AI — it is a company that used deploying AI on itself as the mechanism to force a long-overdue simplification of how it operates. The agents came second. The process redesign came first.

The Numbers Have Only Gone Up Since

The $3.5 billion figure is now the baseline, not the ceiling.

By Q2 2025, IBM’s CFO had raised the internal target: the company now expected to achieve approximately $4.5 billion in annual run-rate savings by the end of 2025, which was achieved. By June 2026, analysts reported that IBM had banked $4.5 billion in internal productivity savings since 2023 and was targeting a further $1 billion that year. IBM’s own published data for 2024 shows that AI and automation saved IBM employees more than 3.9 million hours in a single calendar year — the equivalent of roughly 2,000 person-years of work.

Meanwhile, IBM’s AI programme expanded into software development. IBM’s Q3 2025 SEC earnings remarks describe “Project Bob,” an AI-powered software development assistant now used by more than 8,000 IBM developers, with those developers reporting average productivity gains of 45%. The same filing reports IBM’s generative AI book of business had reached $9.5 billion inception-to-date as of Q3 2025, with more than 1,000 Client Zero engagements underway with external clients.

The internal programme has, in the most direct sense possible, become a sales asset.

What IBM Did With the Savings

This is the part of the story that gets overlooked in the headline number, and it’s arguably more important for business buyers thinking about their own AI strategy.

IBM CEO Arvind Krishna confirmed publicly in May 2025 speaking in a conversation with The Wall Street Journal said that AI had replaced the work of several hundred HR employees at IBM. He also confirmed total headcount went up. The capital freed by those automations was reinvested into engineers, salespeople, and client-facing roles including, directly, a team of former HR professionals who now advise IBM clients on how to replicate what IBM did.

“Our total employment has actually gone up,” Krishna told the Journal, “because what [AI] does is it gives you more investment to put into other areas.”

That pattern — automate the transactional work, reinvest the freed capacity in higher-value work — is not just a story IBM tells to make the numbers look better. It is what the financial performance shows. IBM’s software margins improved. Its AI services revenue grew. The savings from internal efficiency became the investment that funded the revenue-generating side of the business.

The CFO called it a “flywheel.” It is, accurately, a flywheel: AI saves cost → savings fund reinvestment → reinvestment produces revenue growth → revenue funds more AI. The savings were not extracted and distributed. They were compounded.

What This Means If You’re Not IBM

IBM’s scale makes it easy to dismiss as a special case. It is not. Four things from IBM’s programme transfer directly to organisations of almost any size.

1. The sequence matters more than the tool. IBM’s biggest gains came from eliminating and simplifying before automating — not from picking the best AI model. The equivalent question for any organisation: which of your current processes would you rebuild from scratch if you were starting today? Those are the candidates for automation. Everything else is just speeding up a legacy workflow.

2. Start where humans are doing high-volume, rule-based work with clear outcomes. AskHR was not IBM’s most complex system — it was its most transactional. Vacation requests, payslip queries, policy clarifications: clear inputs, known outputs, high volume, and an easy measurement of whether the agent got it right. That’s the right starting point. Complex, judgment-heavy work comes later, as trust builds.

3. Measure containment rate before cost savings. IBM’s most useful early metric was containment rate — what percentage of interactions the agent resolved without human escalation. That metric tells you whether the agent is actually working before you count the savings. A low containment rate means the agent is creating work, not reducing it.

4. The Client Zero logic applies at every scale. IBM proved its AI products on itself before selling them. The principle behind that — don’t speculate about what’s possible, prove it in your own operations — applies equally to a 50-person company deciding whether to automate its onboarding process. IBM’s advantage was scale and proprietary tooling. The discipline is available to anyone.

A fifth point worth raising honestly: IBM’s programme took years, not months. AskHR launched in 2017 and reached its current results by 2024. AskIT launched in 2023. The $3.5 billion run-rate figure represents the accumulated output of a long, methodical build — not a deployment sprint. If a vendor is promising enterprise-grade AI agent results in weeks, the IBM timeline is a useful reality check.

IBM AI Agents FAQs

Is the $3.5 billion figure independently verified?

It is not independently audited in the way revenue is, but it is a publicly stated figure made by IBM’s CFO on a Q4 2024 earnings call under the legal obligations of public company disclosure. It reconciles with IBM’s segment margin improvements across consecutive earnings reports and is supported by specific, named metrics IBM has published at the business-function level.

Which AI tools did IBM use to build these agents?

IBM built its agents on its own watsonx platform — including watsonx Orchestrate for agent coordination and watsonx.ai for the underlying models (IBM’s own Granite-class large language models). That is worth noting when benchmarking IBM’s results: the company built on tools its own engineers developed and understood deeply, which is an advantage most organisations deploying third-party AI tools will not have.

Did IBM replace its workforce with AI?

Hundreds of HR roles were replaced by AskHR, as IBM CEO Arvind Krishna confirmed in May 2025. However, IBM’s total headcount increased over the same period, because the capital freed by those automations was reinvested in engineering, sales, and client-facing roles. IBM’s model is redeployment-first: HR employees who were displaced moved into client advisory roles, teaching other companies how to replicate what IBM built.

How long did it take IBM to achieve these results?

Longer than most AI deployment timelines suggest. AskHR launched in 2017 and spent years being refined before it reached the containment rates and cost reduction figures IBM reports today. AskIT launched in 2023. The $3.5 billion run-rate figure reflects the accumulated result of a multi-year, multi-function programme — not a rapid deployment.

Can smaller companies replicate IBM’s approach?

The specific numbers cannot be replicated at a smaller scale, but the approach can be. The core principles — eliminate before you automate, start with high-volume transactional work, measure containment rate before cost savings, prove it internally before scaling — are tool-agnostic and size-agnostic. IBM’s advantage was scale and proprietary tooling. The discipline is available to anyone.

How does IBM’s experience relate to the broader AI agent adoption picture?

IBM is one of the clearest production-scale examples of what enterprise AI agents actually deliver but it is not representative of where most companies are. Across the enterprise broadly, estimates suggest only around 11% of companies that have “adopted” AI agents are running them in production at meaningful scale. IBM is the successful end of a distribution that still has a long tail of pilot projects that haven’t converted to real impact. We cover that gap in detail in our post on the agentic AI production gap.

The Bottom Line

IBM set a $2 billion target for AI-driven productivity savings. It hit $3.5 billion, further raised the target to $4.5 billion. It hit that too, and is now targeting more. The number keeps moving because the flywheel keeps turning: every efficiency generates the capital to fund the next round of investment, which generates more efficiency.

That is not how most enterprise AI stories end. Most end at the pilot phase, with promising results that never quite converted into a measurable line on the P&L. IBM’s story ended differently because it started differently, with a deliberate decision to prove the technology on itself, at full operational scale, before selling any of it to anyone else.

The discipline of building and verifying before claiming is the lesson that transfers, regardless of your size, your industry, or which AI tools you’re evaluating.

At Doshby, it is also how we work. We do not pitch agentic AI deployments based on what the models can theoretically do. We build, run, and verify them and we can show you what that looks like in practice. Get in touch if you’d like to talk through what a Client Zero approach could look like for your business.

You May Also Like…

What Is OpenCode AI?

What Is OpenCode AI?

The Open-Source Coding Agent Taking On Claude Code and Cursor In early 2026, Anthropic enforced its terms of service...

What is Vibe coding?

What is Vibe coding?

A Precise Definition for 2026 In mid-2025, a GitHub README for a personal audio project uploaded by Linus Torvalds...