Cut Too Deep: Klarna, Duolingo, and the Cost of Borrowed Conviction
In February 2024, Klarna announced its AI assistant was doing the work of 700 agents; fifteen months later its CEO told Bloomberg they'd 'cut too deep.' A year after that, Duolingo put AI usage into performance reviews, watched its most loyal users revolt, and quietly took it back out. Neither company was wrong about AI. Both were wrong about what their own principles required — and the distance between those two mistakes is where this whole series has been heading.

In February 2024, Klarna published the most quoted press release of the AI era. Its new AI assistant, built on OpenAI’s models, had handled 2.3 million customer service conversations in its first month — two-thirds of the company’s volume — and was, by Klarna’s own arithmetic, doing the work of 700 full-time agents. Average resolution time had dropped from eleven minutes to under two. The company projected a $40 million profit improvement for the year. Every number in that paragraph is Klarna’s own, self-reported, and worth exactly that caveat — but the numbers weren’t really the point. The equivalence was the point. “The work of 700 agents” is not a metric; it’s a headline with a body count implied, and the market read it exactly as designed. Klarna became the poster child for AI-replaces-headcount, cited in a thousand board decks by executives who had never used its product.
Fifteen months later, in May 2025, CEO Sebastian Siemiatkowski sat down with Bloomberg and said the thing poster children aren’t supposed to say: they had cut too deep. Service quality had fallen. Customers with messy, emotional, edge-case problems — the ones an eleven-minute human conversation existed to absorb — were hitting a wall. Klarna would be restoring a hybrid model, making sure a human was reachable again.
Here’s what makes this a story worth closing a series with, rather than just a news cycle: Klarna never actually fired 700 people for AI. The famous number was a workload-equivalence claim; the headcount reduction underneath it came from outsourced BPO contracts and a hiring freeze. And Klarna did not abandon its AI agent in 2025 — the assistant kept handling the bulk of the volume. Both the legend of the announcement and the legend of the retreat are wrong. What actually happened is quieter and more useful: a company optimized a borrowed narrative until it collided with its own product’s first principles.
What Klarna’s principles actually required
Strip away the AI framing and ask what Klarna is. It’s a payments company whose entire consumer promise is smoothness — buy now, pay later, and when something goes wrong with a refund or a disputed charge, get it resolved without pain. For a company like that, customer service isn’t a cost center that happens to be attached to the product. In the moments that matter most to a customer — money stuck, order wrong, stress high — customer service is the product.
Klarna’s own February 2024 numbers, read carefully, were genuinely impressive on exactly that axis: resolution time falling from eleven minutes to under two is a service-quality story, not just a cost story. The AI assistant was good at the high-volume, well-structured cases, and by all accounts it stayed good at them. The mistake wasn’t deploying it. The mistake was letting the borrowed narrative — AI replaces headcount, and the replacement ratio is the score — set the target, so that the system got optimized for the equivalence number rather than for the promise. “Work of 700 agents” measures substitution. It does not measure whether the customer with the genuinely weird problem, the one no flow anticipated, still ends the conversation feeling like Klarna kept its word. Those customers are a small fraction of volume and a large fraction of brand. That’s the edge-case territory Siemiatkowski’s reversal was actually about, and it’s why “Klarna gave up on AI” is as false as “AI fired 700 people.” They gave up on a scoreboard, not a technology.
If you’ve read the metrics post in the frameworks series, you’ll recognize the shape: a proxy metric, publicly committed to, drifting away from the outcome it was supposed to stand for. The unusual part here is only that the proxy was borrowed. “Agents replaced” wasn’t derived from Klarna’s tree at all. It was imported, fully formed, from the story the entire industry was telling itself in early 2024 — and an imported metric carries someone else’s assumptions about what your product is for.
Duolingo borrowed a memo
Fourteen months after Klarna’s announcement, on April 28, 2025, Duolingo’s Luis von Ahn published an “AI-first” memo to the company, explicitly in the mold of Shopify’s a few weeks earlier. Among its provisions: AI usage would become part of performance reviews. Constrained headcount growth. Contractors phased out where AI could do the work. It was, almost line by line, the era’s standard-issue conviction — the same conviction, from the same shared narrative, that Klarna had ridden up and was at that very moment publicly climbing down from.
The revolt came from an unexpected direction: not employees, but users. And not casual users — the streak holders. Duolingo is a company whose growth machine, as its own former CPO has described in detail, runs on the streak and the emotional relationship users have with a pushy green owl. By its own shareholder reporting, more than ten million users hold streaks of a year or longer. Those are people who had structured a small piece of their daily identity around the product’s persona — and the persona’s whole register is warmth. Anthropomorphic, teasing, slightly unhinged warmth, but warmth. When the company behind the owl started talking like an efficiency memo, users with multi-year streaks began publicly deleting the app, posting the streak-loss confirmation screens as protest. The asset revolted against the strategy.
Von Ahn walked it back in stages. By June 2025 he was conceding, in his own words, that he “didn’t do that well” in how the message was communicated, clarifying that AI was meant to augment employees, not replace them. The walk-back continued for months, and by April 2026, AI usage was out of performance evaluations entirely. One detail deserves precision, because the legend-vs-record gap is the whole theme here: Duolingo’s daily active user growth decelerated through this period. It did not decline. The company kept growing; what it lost was some of the effortless momentum — and, less measurably, some of the accumulated trust that made the owl’s pushiness feel like affection instead of extraction.
Two reversals, one cause
Put the stories side by side and the surface details diverge everywhere — fintech versus edtech, customers versus users, a Bloomberg interview versus a LinkedIn memo. But the underlying mechanism is identical, and it’s worth naming exactly, because I think it’s the most common strategic failure of the last few years and it will be the most common one of the next few.
Neither company was wrong about AI. Klarna’s assistant genuinely worked; Duolingo has been shipping AI-generated course content and AI conversation features to real effect. Both companies were wrong about what their own principles required of them. Klarna’s founding promise was that dealing with money through Klarna would feel humane and frictionless; an AI strategy derived from that principle would have automated the routine ruthlessly and protected the human escape hatch as sacred, from day one, on purpose — instead of rediscovering it by public apology. Duolingo’s principle was that motivation is the product and the owl’s warmth is the delivery mechanism; an AI strategy derived from that would have been framed, internally and externally, as “more owl, faster” — never as a performance-review compliance line, which is the least warm sentence a company can utter.
What both companies actually did was skip the derivation. They adopted the era’s narrative whole — the same way teams once adopted a squad model its own authors said they weren’t using — and let someone else’s conviction stand in for their own. I’ve started calling this borrowed conviction, and its signature is that it arrives with the confidence of a conclusion but none of the reasoning that would tell you where it stops applying. Conviction you derived yourself comes with boundary conditions attached: you know which customer, which metric, which promise it must not violate, because you walked past them on the way to the conclusion. Conviction you borrowed has no edges. It runs until it hits something — and what it hits, both times here, was the company’s own product telling it no. Klarna’s edge-case customers and Duolingo’s streak holders performed the same function the sprint demo performed in the Sentinel story: the reality check that borrowed narratives never schedule for themselves.
The record, versus the legend
Every case study I’ve told this way has had the same skeleton under different skin. A legend that’s clean, quotable, and load-bearing for someone’s strategy deck. A record that’s messier, better documented, and almost always more instructive. Blockbuster didn’t laugh Netflix out of the room; it built a counterattack and then abandoned it. Kodak didn’t ignore digital; it led digital and couldn’t cascade. The famous growth number varies by telling, and the person who shipped it said out loud he couldn’t prove causation. Klarna didn’t fire 700 people for AI and didn’t abandon AI in remorse. Duolingo’s users didn’t flee; growth decelerated while the company renegotiated its own tone.
In every case, the legend eventually loses to reality — the only variable is who’s standing on it when it does. And the people standing on legends are, overwhelmingly, people running on borrowed conviction: the framework adopted because a famous company blogged it, the metric adopted because a famous investor tweeted it, the memo adopted because a peer CEO published one three weeks earlier. The closing argument of the frameworks series was that tools are scaffolding for principles, and the goal is to internalize the principle and drop the scaffold. This has been the same argument run against stories instead of tools. A famous case study is a framework with a plot. It compresses someone else’s context into a portable conclusion, and the compression is lossy in exactly the same place: the edges, where your context and theirs diverge.
The only durable move — the one thing every case above agrees on — is to reason from your own evidence. Your own customers’ behavior, your own product’s promise, your own numbers with their own caveats attached. Borrow the stories for what they’re genuinely good for: a catalog of failure modes, a vocabulary, a cheap way to feel expensive pain secondhand. Then do the derivation yourself. Klarna and Duolingo both, eventually, did — that’s what the reversals were. The reversal isn’t the embarrassing part of either story. It’s the part where each company stopped borrowing and started reasoning, and the only real indictment is how publicly they had to do it.
Put it to work
- Before adopting any industry-wide narrative, write the derivation from your own principles. One page: here is our product’s core promise, here is what this strategy (AI-first or otherwise) would look like derived from that promise, and here is where the imported version conflicts with it. If you can’t write the derivation, you don’t have a strategy yet — you have a citation.
- Name the constituency your borrowed metric ignores, and instrument them first. Klarna’s equivalence number was silent about edge-case customers; Duolingo’s memo was silent about streak holders. Whatever headline number you’re about to commit to publicly, ask who matters to your product that this number cannot see — then put a measure on them before you announce anything.
- Pre-write the walk-back. Draft the paragraph you’d have to publish if this bet is 30% wrong — what you’d restore, what you’d concede, what it would cost. Siemiatkowski and von Ahn both had to compose theirs live, in public. If drafting yours in advance feels unbearable, that discomfort is your first honest read on how much conviction is actually yours.
Further reading
- Klarna’s February 2024 press release on its AI assistant — the primary artifact, worth reading in full precisely because every number in it is self-reported; a masterclass in how an equivalence claim becomes a headline becomes an industry narrative.
- Bloomberg’s May 2025 interview with Sebastian Siemiatkowski (Bloomberg, May 2025) — the “cut too deep” conversation, and notably not an abandonment of AI; the nuance the aggregators dropped.
- Luis von Ahn’s April 2025 AI-first memo and his subsequent public clarifications — read the original and the walk-back side by side as a study in tone colliding with brand.
- Jorge Mazal’s account of Duolingo’s growth model on Lenny’s Newsletter (2023) — the practitioner’s view of the streak machinery that explains exactly why the memo’s tone was the expensive part.
Product Stories
2 parts in this series.
A two-part pair on discovery done right and wrong — Segment and Superhuman's evidence-driven pivots, set against Quibi, Juicero, and Humane's false validation, closing with Klarna and Duolingo's AI-first reversals — told from the record rather than the legend.
- 01Watch What They Do: Segment, Superhuman, and Discovery That Workedprevious
- 02Cut Too Deep: Klarna, Duolingo, and the Cost of Borrowed Conviction← you are here

What did you take away?
Thoughts, pushback, or a story of your own? Drop a reply below — I read every one.
Comments are powered by Disqus. By posting you agree to theirterms.