🇹🇷 Türkçe: Bu yazının Türkçesini oku →

I have been watching operations chiefs get pitched some version of “AI will handle two thirds of your customer inquiries” for three years, and the pitch keeps landing the same way. A cost figure goes on a slide, a headcount reduction gets penciled in against it, a pilot gets spun up. Nine to eighteen months later the same conversation happens again, quieter and in a smaller room, with a different question: how do we get some of this quality back without admitting we cut too far.

Klarna has already run the round trip in public. February 2024: an OpenAI-powered assistant doing the work of seven hundred agents. May 2025: the CEO on Bloomberg saying quality had come down and human staff were being brought back. November 2025: the same company telling a quarterly earnings call the assistant was now doing the work of 853, and that customer service costs were up anyway. That last sentence is the one that matters for anybody sitting in front of the same pitch this month.

I have some sympathy for the pitch, because the productivity numbers are real. The problem is that the productivity got priced in as a saving before anybody knew what it did to the rest of the P&L. That is the pattern worth reading Klarna’s arc against.

Klarna is the cleanest public experiment in AI customer service anyone has run

Klarna is a Swedish buy-now-pay-later platform that listed in September 2025, roughly a year and a half into its most public bet on generative AI. Its scale is what makes the case useful: the February 27, 2024 press release put the assistant into 23 markets and 35+ languages from the first month. Not a departmental pilot, but a production system carrying real support load in front of tens of millions of shoppers, with a listed company’s disclosure obligations now attached to the numbers.

Every buyer I speak to on this ends up in one of two camps: Klarna proves the technology works and the reversal was a communications wobble, or the reversal proves the technology fails. Neither reading survives the filings. The technology works at the volumes claimed, and it creates second-order problems the original business case never priced. The organisations that benefit from watching this are the ones that adjust their business case rather than pick a side.

The before-state was a support operation buckling under BNPL scale

Klarna’s 2023 support load was a function of what buy-now-pay-later is: high-volume, low-value transactions split across four or six instalments, spanning weeks per order, with dispute and repayment questions layered on top of ordinary retail queries. Average human handling time before the assistant was eleven minutes per resolved inquiry. That is exactly the shape of work where an assistant should help — a long tail of similar questions, most answerable from account data and policy documents.

Klarna employed 5,527 people in 2022, the last full year before any of this, per the company’s own Q3 2025 investor presentation. Sebastian Siemiatkowski was already telling the market that AI would let Klarna reduce headcount, and a hiring freeze went in ahead of the assistant launch. Whatever you make of the strategy, the first move was operational, not technological. Klarna committed to a smaller headcount before the tooling had proved itself. That commitment shaped everything that came after, including the reversal.

The change was not “add an assistant” — it was “reset the size of the support function”

This is the part most write-ups skip. Klarna did not just deploy a chatbot. The assistant launched alongside a hiring freeze, an outsourcing exit, and an internal campaign about AI-driven productivity — each defensible on its own, and together a rewrite of what customer service was for. Any one of them is manageable. Together they closed the exits. Once support was shrinking on all fronts, the assistant had to work, because there was nothing behind it to catch what it did not.

Siemiatkowski framed it in the launch release as “superior experiences for our customers at better prices, more interesting challenges for our employees, and better returns for our investors.” Note the order, customer experience first and investor return last. Then note that what actually got measured, most closely and most publicly, was the cost line.

The measured result depends heavily on which quarter you ask

The February 2024 numbers were striking. The assistant handled 2.3 million conversations in its first month. Two-thirds of Klarna’s customer service chats went through it. Average resolution time dropped from eleven minutes to under two. Repeat inquiries fell twenty five percent. Customer satisfaction was reported “on par” with human agents. A $40 million profit improvement was projected for the calendar year.

Fifteen months later, the picture had shifted. Speaking to Bloomberg on May 8, 2025, Siemiatkowski told the wire, “As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality.” What Klarna reopened was not a conventional hiring round. It piloted what Siemiatkowski described to Bloomberg as an “Uber type of setup”: remote freelance agents choosing their own schedules from anywhere in Sweden, starting at 400 Swedish krona, and two of them at the point the story broke. The guarantee mattered more than the volume. In his words, “it’s so critical that you are clear to your customer that there will be always a human if you want.”

A month later, at a fireside chat at London SXSW on June 4, 2025, he sharpened it into a market position: “We think offering human customer service is always going to be a VIP thing.” Human contact, in the new framing, was not a fallback. It was a paid tier. Meanwhile the headline number kept moving in the other direction: speaking to CNBC on May 14, 2025, Siemiatkowski said AI had helped shrink the workforce by 40%, mostly through attrition under the freeze rather than layoffs.

By the Q3 2025 results, the assistant was doing the equivalent work of 853 full-time agents, up from 700, and Klarna put the annual saving at $58 million. The same presentation reports 28 million AI conversations between November 2023 and September 2025, with the assistant resolving 81% of customer service chats — a higher share than the two-thirds claimed at launch, not a lower one. And in the quarterly earnings release filed alongside it, customer service and operations costs for Q3 2025 were $50 million, against $42 million in Q3 2024. The AI is more productive than it was, it is carrying a larger share of contacts than it was, and the total support cost is 19% higher year on year, because volume has grown and human support has been added back. Siemiatkowski on the call: Klarna continues to see demonstrable value from the assistant and continues to invest.

The productivity figure is the one worth reading last, and it is the strongest number Klarna has. Average revenue per employee went from $344,000 in 2022 to $1,104,000 on a trailing twelve-month basis at Q3 2025 — a 3.2x improvement — against a headcount that fell from 5,527 to 2,907, or 47% fewer people. Staff costs fell from $696 million to $203 million. These are Klarna’s own disclosed figures, not an analyst’s model. They are not evidence that AI replaced the humans; they are evidence that AI, working alongside a smaller and differently-shaped team, moved more revenue per remaining person. The productivity story and the customer-quality story are the same story told from two ends.

What is still unsolved is the definition of quality itself

Here is the part I would flag to any buyer looking at a similar programme. Klarna has never published a quality metric that reconciles the “on par with human agents” claim of 2024 with the “lower quality” admission of 2025. Both are executive statements. Neither is a number a board can hold across both dates. The company appears to have discovered, in production, that some of its measured satisfaction was survivorship bias: the customers who could resolve their own issue quickly loved the assistant, and the ones who could not churned, complained on social channels, or escalated in ways the initial CSAT sample never caught.

This is not a Klarna-specific problem. Every AI support programme I have watched reach production has this shape: the first quarterly report is dominated by the happy path, because the unhappy path takes longer to surface. The honest question is not how good the assistant is, but how good your metrics are at catching the case where it fails quietly. That is a governance question, and it is the one Klarna has been rebuilding in public since May 2025.

The second unresolved piece is what the “VIP” framing means at scale. Selling human contact as a premium tier is coherent for a payments platform whose margins carry it. It is much harder if you are a mid-market retailer whose competitors offer human contact as standard. Read that line as an experiment Klarna is running, not a template to copy.

Inside a real WooCommerce store, this is a governance build, not a chatbot build

The same programme is being pitched to mid-market commerce operations every day, one size down. Instead of “replace 700 agents” it is “replace two agents and outsource the overflow.” The mechanics are identical. Here is what we would build for an organisation in that seat, before touching the model at all.

  • A contact ledger that survives the assistant. Every touchpoint — chatbot, email, phone, WhatsApp — writes to one record keyed to the WooCommerce customer id, carrying resolution status, escalation flag, and a churn signal if that customer stops ordering within ninety days. This is what catches the Klarna problem: the customers who left quietly. On WooCommerce, a custom post type plus a small ETL job from the ESP and the helpdesk. Two to three weeks for the base version.
  • A quality metric finance can hold across quarters. Not CSAT, which is survivor-weighted. A composite: first-contact resolution rate, escalation-to-human rate, and thirty-day repeat-contact rate for the same customer on the same issue — reported next to contact volume, so productivity and quality sit on one page.
  • An assistant with a scope, not a chatbot with a personality. Bind it to known intents — order status, refund status, payment method changes, address updates — and route everything else to a human on the first turn. The assistants that produce Klarna’s “lower quality” outcome are the ones asked to answer anything a shopper types.
  • A human tier that is real, not a marketing line. If you cannot fund named human contact for the top ten percent of customers by lifetime value, do not promise it. If you can, wire it into the account area as a routing rule, not a landing page. Klarna’s “VIP” positioning only works because someone actually answers.
  • A published rollback plan. The most useful thing Klarna’s arc gives an operations chief is permission to write one on day one. Ours reads: if first-contact resolution, escalation-to-human or thirty-day repeat rate moves more than fifteen percent against a locked baseline, human capacity comes back within one quarter and the assistant’s scope narrows. Nobody wants to write this. Everybody eventually needs it.

Total effort to build that governance layer around an assistant, before any model work: eight to twelve weeks with two engineers, an operations lead and a part-time analyst. A real six-figure spend at consultancy rates, and the spend the pitch decks leave out. The assistant is the cheap part. The apparatus that catches failure modes before they become a public reversal is what we would sequence first.

The lesson is not “do not use AI” — it is “do not book the saving before you know the quality”

I do not read Klarna’s arc as a cautionary tale about AI. The technology did what it was asked to do and then did more of it: two-thirds of a very large support load at launch, 81% of chats by Q3 2025, across dozens of markets and thirty-plus languages. The cautionary part is about how organisations translate a productivity gain into a P&L decision. Klarna booked the saving early, shaped the support function around it, and had to unpick that shape when the second-order effects arrived.

If you are being asked to approve a similar programme this quarter, four questions for the supplier or the internal team. What is our composite quality metric, and how does it differ from CSAT. Who owns rollback by name, and what trigger would they pull. What did we agree the saving line would be, and how do we hold it separately from the operations line that will grow. Which customers are we quietly de-prioritising, and how would we know if they left. Clean answers, the programme is probably worth running. Slides, it is not.

The boring half of this work is the half that holds: the ledger, the metric, the scope, the rollback. If you want a second read on a programme already in the room, or one you are about to commission, we would rather build that apparatus than watch it get rebuilt in public a year later. Come and talk to us.

Leave a Reply

Your email address will not be published. Required fields are marked *

Close Search Window
↑