On February 27, 2024, Klarna and OpenAI held a joint press conference to announce something the tech industry had been waiting for: proof, at scale, that AI could replace a meaningful share of human customer service work.
The numbers from Klarna’s first month were hard to argue with: 2.3 million customer conversations handled. Two-thirds of all customer service chats automated. Resolution time cut from 11 minutes to under 2 minutes. Customer satisfaction scores on par with human agents. The equivalent work of 700 full-time agents. A projected $40 million in profit improvement for 2024.
CEO Sebastian Siemiatkowski was buoyant. In a February 2024 Bloomberg interview, he said: “I am of the opinion that AI can already do all of the jobs that we, as humans, do.” He told his own employees to use AI to fill the gaps left by the colleagues who were leaving. The company implemented a hiring freeze that lasted over 12 months. Klarna’s headcount dropped from 5,527 at the end of 2022 to 3,422 by the end of 2024, this happened mostly through attrition, rather than direct layoffs, though the company rarely foregrounded that distinction.
Fifteen months later, Siemiatkowski was back on Bloomberg — this time with a different message.
“We went too far,” he told the outlet in May 2025. “We focused too much on cost. The result was lower quality, and that’s not sustainable.” Klarna was hiring human customer service agents again.
The story generated enormous coverage, most of it framed as a morality tale about AI hubris. That framing is not wrong, exactly. But it misses the more useful lesson, which is both more specific and more transferable than “AI failed.”
So What Actually Happened?
Klarna’s AI customer service deployment did not fail. The automation continued working at scale and by Q3 2025 had grown: the AI was handling the equivalent work of 853 agents (up from 700) and delivering approximately $60 million in annual savings (up from the projected $40 million). What failed was the scope — deploying AI across every category of customer interaction, including complex disputes, emotional escalations, regulatory edge cases, and sensitive financial situations, where AI was never reliably the right tool, and where the gap between “resolved” and “resolved well” was largest and most consequential.
The course correction Klarna made in 2025 was not an abandonment of AI. It was a redesign of where AI sits in the workflow: routine, tier-1, high-volume queries handled by AI; escalation, complexity, and anything requiring empathy or judgment handled by humans. That hybrid model is now expanding, and it is the architecture most of the enterprise AI industry is converging on anyway. Klarna just got there the loud, expensive way.
What Klarna Claimed vs. What Was True
Before drawing lessons from this case, it is worth spending a moment on the three headline numbers that defined Klarna’s February 2024 announcement This is because each one is more complicated than it appeared, and the gap between the claim and the reality is itself one of the lessons.
“The equivalent work of 700 full-time agents”
This figure was Klarna’s own calculation, not an independent audit: 2.3 million conversations divided by average human agent throughput equals approximately 700 agent-equivalents. The calculation is mathematically defensible. What it did not mean was that Klarna had fired 700 people. The real workforce reduction was a hiring freeze, the company stopped backfilling roles it would otherwise have needed to fill during a growth phase. As The Pragmatic Engineer’s Gergely Orosz noted in his contemporaneous analysis of the launch, the chatbot was functioning primarily as an advanced L1 filter — handling the same straightforward questions that would previously have gone to a junior contact centre agent, while transferring anything more complex to a human. “Basically as a filter,” was how Orosz put it after testing it himself.
“$40 million in profit improvement”
This was a projection, not a measured result, and it was announced jointly by Klarna and OpenAI, the company supplying the model, unlike that made by Netflix in their $1 Billion save by AI, which was well measured and published. The figures were “self-reported PR, launched jointly with the vendor that supplied the model — useful as intent, not as audited fact.” as noted by Paul Okhrem’s sourced analysis. The $40M figure was based on the assumption that the AI would continue to perform at first-month levels across all interaction types. That assumption didn’t hold for the complex end of the interaction spectrum.
“On par with human agents in customer satisfaction”
This was true — for routine queries. Klarna’s original press release states the AI was “on par with human agents in regard to customer satisfaction score” and “more accurate in errand resolution, leading to a 25% drop in repeat inquiries.” Those aggregate metrics held up for high-volume, FAQ-style queries. What they masked was performance on the less frequent but higher-stakes interactions such as disputed transactions, financial hardship cases, regulatory complaints, where customers reported generic, insufficiently nuanced responses and where satisfaction deteriorated in ways the aggregate score smoothed over.
None of this makes the February 2024 announcement dishonest. It makes it incomplete — a first-month dashboard presented as a finished result. The gap between the two is where the problem lived.
FURTHER READING
➤ Will AI Replace Software Engineers?
What the Dashboard Showed vs. What Was Actually Breaking
According to Digital Applied’s March 2026 analysis of Klarna’s internal dynamics, the warning signs were in the data. They just weren’t in the metrics anyone was watching most closely.
The headline metrics looked good throughout 2024: overall chat volume handled, time-to-first-response, resolution rate — all of these stayed strong or improved. What deteriorated silently were two other signals: direct satisfaction scores on complex interaction types (not the aggregate CSAT, but the scores on specific escalation categories), and repeat contact rate — the number of customers who had to contact support multiple times for the same issue. A customer who contacts support twice for a refund dispute that should have been resolved on the first call is not a resolved case. They are a brand problem.
The Asisteclick operational analysis is direct about what happened: “The dashboard showed the average. Reality was hidden in the distribution.” The AI was performing excellently in the high-volume centre of the distribution — routine queries with standard answers. It was performing poorly at the tails — rare but high-stakes interactions where the customer most needed something the AI structurally couldn’t provide: flexibility, empathy, discretion, and the willingness to treat a complex situation as genuinely complex rather than as a pattern-matching problem.
By mid-2025, Klarna’s internal team had begun quietly hiring human agents again. No public announcement. No visible reversal. Just a silent course correction — until Siemiatkowski told Bloomberg in May 2025 that customers would always have the option of speaking to a human, that quality had suffered, and that he was “backtracking.”
The lesson is not that AI customer service fails. It is that volume-based metrics are not quality metrics, and optimising for the former can quietly degrade the latter for months before the signal becomes undeniable.
The Course Correction: What Klarna Actually Changed
Klarna’s 2025 pivot was smaller, and more precise, than the coverage suggested.
The AI assistant was not shut down. It was not demoted. By Q3 2025, Klarna reported the assistant was now doing the equivalent work of 853 agents which is up from 700 at launch, with approximately $60 million in annual savings and response times 82% faster than the pre-AI baseline. The technology held and expanded, right through the period the press was calling a “reversal.”
What changed was the model’s scope. Klarna introduced what it described as an “Uber-style” flexible workforce of remote human agents; primarily targeting students, parents with partial availability, and rural workers who are available for the cases the AI cannot handle well. The architecture is now explicitly hybrid: AI handles tier-1, high-volume, routine queries. A human is always reachable for escalation, disputes, emotionally complex situations, and anything where the stakes of a wrong answer are high enough to matter to the customer’s trust in Klarna.
This is the model that works in customer service at scale. It is also the model most enterprise AI deployments in customer service converge on when given enough time and honest data. The Klarna case is notable not because the outcome is unusual, but because the journey to it was more public than anyone planned.
Why the CEO’s Framing Made the Reversal Worse Than It Needed to Be
There is a separate lesson inside the Klarna story that has nothing to do with AI capability, it is about what happens when executive communication outpaces operational reality.
Siemiatkowski’s February 2024 statement “I am of the opinion that AI can already do all of the jobs that we, as humans, do” sits in the public record. It was not buried in an earnings call or a corporate filing. It was a headline. When the course correction came 15 months later, it was therefore not just a technical adjustment to a workflow. It was a visible, public reversal of a very public claim made by the company’s founder.
As Digital Applied’s analysis observed, “the Klarna case is now the canonical enterprise cautionary tale for 2026, executives evaluating AI workforce strategies are increasingly required to explain how their plan avoids the Klarna outcome.” That reputation was earned not by the technology failing, but by the gap between what the CEO said and what the technology was actually capable of.
The technical course correction Klarna made — moving from AI-only to AI-plus-human —exactly how the top 10 AI hyperscalers use AI, which was the right call. Klarna’s own spokesperson told Fortune in May 2025 that the company remains “very much still AI-first.” That framing is accurate and sensible. The reputational cost the company paid to get there was not the cost of the correction. It was the cost of the original overclaim.
For business leaders evaluating AI deployments, that distinction is worth separating: the technical decision and the communications decision are different risks. Getting the technical decision wrong is fixable. Getting the communications decision wrong is fixable too, but it is more expensive and more visible.
What This Means for Your AI Deployment
Klarna’s experience maps onto four concrete principles that apply at almost any company deploying AI in a customer-facing context.
1. Scope the AI to the interactions where it reliably performs well before you deploy it to the ones where it doesn’t.
Klarna’s AI was excellent at L1 support: routine queries, standard answers, high volume, clear resolution criteria. It was poor at L3 escalations: complex disputes, emotionally fraught conversations, edge cases requiring human judgment. The mistake was not distinguishing between these categories in the initial deployment. The right architecture classifies queries by risk band — informational, account-specific, high-stakes — and only fully automates the lower-risk tiers.
2. Measure quality, not just containment.
A query that the AI call agent “resolves” but that sends the customer back three days later is not a resolved query — it is a deferred cost with a satisfaction penalty attached. The metric to watch alongside containment rate is repeat contact rate, and the customer satisfaction score to watch is the one on complex interaction types, not the aggregate. Klarna’s aggregate CSAT held long after the quality problem had started.
3. Keep a human always reachable, especially in financial services.
Siemiatkowski’s post-correction framing says it clearly: “From a brand perspective, a company perspective, I just think it’s so critical that you are clear to your customer that there will always be a human if you want.” For companies operating in fintech, healthcare, legal, or any regulated industry, the option to reach a human is not a product feature — it is a trust signal, and in some jurisdictions a compliance requirement. AI can handle the volume. A human path has to exist.
4. Communicate what the AI actually does and not what you hope it will eventually do.
The $40M figure, the 700-agent equivalence, and the “AI can do all human jobs” statement were all, in different ways, optimistic framings of incomplete data. The reputational cost of walking back those claims was higher than the cost of having made more conservative statements at launch. [As the IBM case study demonstrates](link to IBM post), public savings claims that are conservative and well-sourced are more durable than headline projections that have to be revised. Set expectations at what you can demonstrate, not at what the first-month dashboard suggests.
For business leaders in non-customer-service contexts, the same principles apply with adjusted parameters. Find the boundary between what AI does reliably and what it does poorly. Watch quality metrics, not just throughput. Keep the human loop visible where it matters to users. And communicate what you have actually achieved, not what you expect to achieve.
The Bottom Line
The Klarna story is not really about AI failing at customer service. The AI is still running, still growing, and now saving more than it was projected to save when the triumphant press conference was held in February 2024. What the story is about is the gap between two prepositions: deploying AI instead of humans, versus deploying it alongside them.
As the operational analysis from Asisteclick puts it directly: “Klarna didn’t fail by using AI. It failed by using it instead of humans rather than alongside humans.” The hybrid model that Klarna arrived at in 2025 is not a consolation prize. It is the correct architecture — one that was available from the beginning, if the initial deployment had been scoped more carefully and communicated more honestly.
The Klarna case has become the enterprise reference point for what happens when AI deployment scope and executive communication outrun what the technology has actually demonstrated. Boards and investors now routinely ask companies to explain how their AI strategy avoids the same outcome.
At Doshby, we start with scope, then define precisely where AI adds reliable value and where the human loop is non-negotiable before we build anything. If you’re designing an AI deployment strategy and want to avoid building toward a course correction you’ll have to explain publicly, speak to us about it.



