What each case actually proves, and what gets wrongly read into it
Alan opens its codebase to non-engineers. PayFit structures its adoption around an AI Ops role and 29 champions. Klarna automates its customer service at scale before reinvesting in humans. Doctolib moves 600 engineers to an agentic-first mode. Taken separately, any one of these cases can be used to defend almost any thesis about the future of work.
That's precisely the problem.
These four companies aren't demonstrating four stages of the same AI revolution. They show four organisational responses to four different problems. Cherry-picking begins the moment you forget the starting problem, the nature of the metric, the guardrails put in place, and what the company itself hasn't yet demonstrated.
These cases, documented among thirty companies observed between January and July 2026, call for two precautions: the set was selected for its organisational signals and makes no claim to being representative; a single company can fall under several organisational models at once.
Company case studies have become the raw material of the discourse on AI.
One company announces a hiring freeze: AI supposedly lets you grow without hiring. Another shrinks a team: organisations are supposedly getting mechanically smaller. A third opens up code to non-technical roles: professional boundaries are supposedly dissolving. A fourth reports 100% adoption: it must have found the recipe for productivity.
Each shortcut sometimes contains a grain of truth. None is sufficient proof.
An adoption metric doesn't necessarily measure productivity. A production metric doesn't necessarily measure quality. A headcount decline doesn't let you isolate AI's causal role. A new organisational structure doesn't prove it would be better in a different context.
That's why the useful question isn't "which company should we copy?". It's: "what problem was this company trying to solve, what human-AI working arrangement did it put in place, and which part of the result is actually documented?"
| Case | Documented change | Tempting reading | More robust reading |
|---|---|---|---|
| Alan | Non-engineers contribute directly to the product with AI. | "Everyone codes, so engineering matters less." | Production capacity is distributed, but control, rules and review become more important. |
| PayFit | A cross-cutting AI Ops role + sponsors + 29 champions + builders. | "A network of champions is enough to transform the company." | Distributed adoption requires orchestration, allocated time, accountability and capitalisation. |
| Klarna | Massive automation of support, a headcount reduction, then reinforced access to a human. | "AI replaces agents", or conversely, "AI failed." | AI can absorb a large share of the volume while hitting a quality ceiling that forces a hybrid model. |
| Doctolib | 600 engineers switched to agentic-first. | "100% adoption = massively multiplied productivity." | The organisation of development genuinely changes; the net productivity gain still has to be measured. |
Alan's transformation is probably one of the easiest to over-interpret.
With Everyone Can Build, the company wants designers, PMs and other non-engineers to be able to change the product directly. In an account published in late 2025, Alan said 283 pull requests from non-engineers had shipped to production after two quarters. The company describes a standardised environment, shared rules for coding agents, and support designed to make contribution accessible.
The account gathered by TPC goes further: an initially limited scope, an engineering buddy, a complexity assessment, review batches, and the merge decision kept on the engineering side. The same account also notes that Alan doesn't yet have a robust measure to attribute an overall ROI to this transformation.
The misreading: "If a PM can code, we need fewer engineers"
That's possible in some setups. But it's not what the Alan case demonstrates.
What it actually shows is the removal of a handoff: a small change no longer necessarily has to pass from PM to designer to engineer before it exists in the product.
The distinction matters. The engineering role doesn't disappear from the system: part of its value shifts from exclusive production toward the environment that makes distributed production safe. Review, architecture, developer experience, tests and contribution rules become the conditions for opening things up.
Alan, in fact, already had a culture of distributed accountability and written documentation before this wave of AI. Its product organisation has historically rested on cross-functional teams and strong autonomy. AI is therefore landing on especially favourable ground; it isn't creating this culture from scratch.
What this proves: some barriers between Product, Design and Engineering can be lowered when the review system and the technical environment allow it.
What this doesn't prove: that every company can open up its codebase the same way, that expert depth becomes less necessary, or that the number of contributions alone is a measure of value.
Alan's real transformation, then, is probably not "everyone becomes a developer". It's rather: more people can produce, while expertise shifts toward creating and protecting the frame that makes that production possible.
PayFit tells a different story.
The Senior AI Ops role was designed to own adoption, governance, team support, tool tracking and running an internal community across the whole organisation. The job listing PayFit published placed it within Product Ops, working with Product, Engineering and the operational functions.
The field account gathered by TPC then documents 29 champions spread across nine major departments, with one AI Ops role for more than 700 employees. Sponsors set the ambition and provide the resources; champions identify and spread use cases; builders implement and measure them. A champion with no allocated time produces little effect.
The misreading: "Just create an AI Ops role and a network of champions"
The number of champions is an organisational fact, not an impact metric.
A network of 29 people can be extremely active, or it can become a collection of ambassadors with no real mandate. PayFit itself treats it as adoption infrastructure: the sponsor, the time actually available, the use cases shipped to production, and measured results matter more than the number of volunteers.
This is where the case gets interesting. It doesn't really settle the question of centralisation versus decentralisation. It combines both.
The centre doesn't try to produce every use case itself. It supplies method, consistency, governance and dissemination. The business units keep the knowledge of the problem, and gradually take back ownership of the use cases.
The move now goes beyond internal adoption. In July 2026, PayFit announced it had reached profitability and was shifting its positioning from software vendor to an AI-augmented payroll and HR service, combining its platform, PayFit AI, and around 200 payroll experts. The company states that clients keep validation checkpoints at key steps and that human experts step in when needed. These points come from PayFit's own communications: they document a strategic direction, not its future economic impact.
What this proves: a company can move from scattered experimentation to an explicit adoption programme without necessarily building a large central AI department.
What this doesn't prove: that AI Ops is the optimal model everywhere, that 29 champions produce a measurable ROI, or that this setup would work without sponsors who genuinely have the ability to free up time.
The lesson is less "create an AI Ops role" than: "give an owner to the cross-cutting problems local teams can't solve alone, then push ownership of the use cases back as close to the business as possible."
Klarna is probably the best stress test of our own intellectual discipline.
In February 2024, the company announced its AI assistant had handled 2.3 million conversations in one month — two-thirds of its customer-service chats — for a workload presented as equivalent to 700 full-time agents. Klarna also reported resolution time falling from 11 minutes to under 2, a 25% drop in repeat enquiries, and satisfaction comparable to human agents. All these figures were reported directly by Klarna and should be read as such.
A year later, the company was still highlighting major efficiency gains. In its Q1 2025 results, Klarna reported having cut headcount by around 40% since 2022, sharply raised revenue per employee, and lowered its cost to serve per transaction. Here too, these are figures and attributions published by the company itself.
Then the narrative changed.
In May 2025, Sebastian Siemiatkowski acknowledged that cost had become too dominant a criterion in how support was organised, and that the result could be lower quality. Klarna then began building a system to strengthen access to human agents, with the explicit ambition that a customer could always reach a person if they wanted to.
First misreading: "Klarna replaced 700 employees with AI"
That isn't what the original announcement said. Klarna was talking about a workload equivalent to 700 FTEs, not 700 lay-offs directly and causally attributed to the assistant. The headcount reduction sat within a hiring freeze, attrition, and a broader company restructuring.
Second misreading: "Klarna tried AI, it failed, it's hiring humans back"
This reading is just as fragile. The company hasn't abandoned its AI. At the very moment it was strengthening human support, a Klarna spokesperson made clear the move was meant to complement AI rather than signal a general retreat.
The case therefore documents something more useful than a success or a failure: an automation boundary was moved after running up against a variable other than cost — perceived quality, and the ability to get a human interaction.
This is precisely what automation dashboards can hide. An organisation can improve response time, cut repeat contacts, and lower average cost, while discovering that some interactions carry a relational value or a complexity that justifies different handling.
The organisational choice then stops being "human or AI". It becomes: which requests can be automated, which have to be escalated, what authority does the human have over exceptions, and which indicators stop a local cost optimisation from degrading the overall experience?
Doctolib supplies the most spectacular case on the engineering side, but also one where methodological caution is most instructive.
Since January 2026, Doctolib's 600 development engineers have been using AI agents in a set-up described as agentic-first. Some run up to eight agents in parallel. In parallel, the organisation has strengthened its specifications, built an internal platform, set up shared services, and trialled agents on ticket handling, some technical changes, and review.
The misreading: "600 engineers at 100% with agents equals a massive productivity gain"
Doctolib doesn't say that. On the contrary, Julien Tanay, Staff SRE, explained in spring 2026 that it was still too early to draw conclusions about the transformation's impact. Teams code faster, but also spend more time on review. The company was still tracking quality and the duration of operations before determining whether agentic-first actually shipped more features to production.
This is a major point: adoption is an input variable. Net productivity is an output variable.
Between the two sit the specs, the quality of the context, the volume generated, review, tests, errors, rework, the merge queue, and ultimately the value actually delivered.
This is also why the Doctolib case is more interesting as a transformation of the production system than as a productivity benchmark. When hundreds of agents need to understand a codebase, the company has to improve its documentation. When production accelerates, acceptance criteria and review matter more. When tasks become automatable, you have to distinguish the ones that are repeatable enough from the ones that still require human judgement.
The same parallel exists in other knowledge professions: the more AI accelerates production in consulting, the more accountability tends to move up to whoever checks and signs off. This analogy helps in thinking through the Doctolib case, but it obviously doesn't count as additional proof about engineering.
What Doctolib proves: an organisation of several hundred engineers can roll out a deeply agentic mode of development across the board, and adapt its infrastructure and practices accordingly.
What it doesn't yet prove: the size of the productivity gain, whether team sizes can be reduced, or the long-term sustainability of these new workflows.
This is probably the most important conclusion from the comparison.
Alan, PayFit, Klarna and Doctolib don't sit on a line running from "little AI" to "lots of AI". They moved different boundaries.
At Alan, the boundary moves between who can produce and who controls production. At PayFit, it moves between central accountability and local ownership. At Klarna, it moves between automated handling of volume and human intervention on interactions where relational quality matters more. At Doctolib, it moves between code production, specification, orchestration and review.
Even the categories used in this article, then, have to be handled with care: Alan, PayFit and Doctolib all fall under distributed adoption, but PayFit also fits the "human on exceptions" model and Doctolib the agentic-development one. The models overlap because they're answering different problems.
AI transformation looks less like choosing a target org chart and more like renegotiating a series of working arrangements: who produces, who controls, who owns the context, who can decide, and who carries final accountability.
They don't demonstrate that organisations will mechanically shrink their headcount. They don't demonstrate that high adoption produces high productivity. They don't demonstrate that specialist roles are disappearing in favour of generalists. Nor do they demonstrate that every deep transformation requires changing the organisation.
Two counter-examples from the same set of cases are particularly useful.
Lucca has integrated AI features while keeping its six-to-eight-person product teams and a doctrine in which AI proposes while the HR professional, recruiter or manager remains the decision-maker. No public data lets you attribute a productivity gain or a headcount reduction to AI.
Pennylane, for its part, continues to hire heavily, brings former chartered accountants directly into its Product functions, and shows no general AI-driven organisational overhaul in the documented material.
These counter-examples don't refute the four preceding cases. They refute the broader idea that using AI intensively would necessarily force you to adopt a smaller, flatter or more agentic organisation.
The most robust organisational signal isn't that one specific model is winning out. It's that the most advanced companies are gradually moving away from thinking in terms of "tasks handed to AI" and instead redrawing complete loops of production, control and decision-making.
This distinction changes a great deal.
Buying a copilot changes a task. Letting a designer push a PR changes a workflow. Creating prompt owners and champions changes accountability. Making an agent the first reviewer changes quality control. Reserving certain cases for a human changes the division of labour. Redefining what's reversible or irreversible changes decision rights.
These are the shifts a leader should look for when studying an outside case. Not the number of licences. Not the number of agents. Not even the percentage of code or tickets handled by AI, taken on its own.
The real indicator of a transformation is that the human work left over is no longer organised the way it was before.
The first issue is telling apart diffusion from transformation. An organisation can have hundreds of active users without having changed its handoffs, its accountability, or its decisions. Conversely, a targeted automation can deeply transform one function without the whole company using AI.
The second is never steering a programme by a single metric. Automated volume has to be read alongside quality, the exception rate, rework and satisfaction. The number of PRs generated has to be read alongside review time and defects. The number of champions has to be read alongside the use cases actually shipped to production and their results.
The third is watching where the new bottleneck appears. At Alan it moves to review. At PayFit, to orchestration and capitalisation. At Klarna, to the quality of complex interactions. At Doctolib, to specs, context and review. AI doesn't necessarily remove the constraint: it moves it. This is one of the most solid cross-cutting lessons from this body of cases.
Finally, you have to accept that a good programme can be corrected. Klarna is interesting precisely because the organisation didn't stay a prisoner of its first narrative. In a field where technical capability moves very fast, the ability to revisit a human-AI boundary is probably more useful than a rigid "AI-first" doctrine.
If several answers are missing, the case is still interesting. But it should be used as a hypothesis, not as proof.
There is no "AI case study" you can copy without losing what matters most.
Alan is not proof that every role should learn to code. PayFit is not proof that you need to hire an AI Ops role. Klarna is neither proof that AI replaces support, nor that it can't. Doctolib is not proof that 100% adoption automatically produces a more productive organisation.
These companies are useful because they let you study concrete organisational decisions made under real constraints.
The question worth importing is therefore never "how do we do what they did?". It's: "what problem did they solve, what organisational price did they pay, and do we have the same conditions here for that choice to make sense?"
These positions shape the way we approach IT recruitment in the age of AI. Let's talk about your context.