Our stance on AI · Scale-ups & startups

From producer to supervisor: how AI reshapes roles once it absorbs the standard work

What supervision actually involves — and why the answer changes by role

When AI takes over the first draft of code, the simple reply to a support ticket, the review of hundreds of files, or matching a job spec against CVs, one phrase keeps coming up: the human becomes a supervisor.

That describes part of what's genuinely happening. But it's too simple to guide an organisation. "Supervising" can cover very different realities: checking an output, orchestrating several agents, resolving cases the machine couldn't handle, defining criteria upfront, or keeping the final decision. These activities require different skills, different organisation, and different metrics.

The real change lies elsewhere: AI doesn't just change how much human work there is. It redraws the role between production, control, exceptions and decisions. And this reshuffling can also create new problems: review overload, a build-up of the hardest cases, lost learning, or greater accountability with no extra resources.

Summary

In brief

The role doesn't shift toward a single supervisor function

A role can be broken down, very schematically, into five categories of activity: frame what needs doing → produce → check → handle exceptions → decide and be accountable.

Historically, many office roles spent a large share of their time on the second step: producing. Writing the code. Answering the ticket. Reviewing the file. Reading applications. Preparing the analysis.

AI attacks this zone first, wherever it's repetitive and formalisable enough. In some workflows, it's also starting to take on a first layer of checking. Human work doesn't necessarily disappear, then: its centre of gravity shifts.

Role Standard work absorbed by AI Human work that gains weight New risk
Engineering Generating code, tests, documentation, first-pass review. Specs, architecture, review, evals, trade-off calls. Turning review into a bottleneck.
Support Simple questions, information lookup, standard replies. Complex, sensitive or ambiguous cases. Leaving humans with nothing but the hard situations.
Ops Bulk checks, reconciliation, sorting, prep work. Diagnosing anomalies, resolving exceptions. Losing sight of the full flow.
Recruitment Initial screening, matching, application summaries. Defining criteria, interviewing, judgment, decision. Automating poorly defined or biased criteria.

This reading lines up with the "human on exceptions" model identified in several of the cases observed, with a condition that's often forgotten: without measuring automated volume, exceptions, errors and rework, there's no way to know whether the work has genuinely improved or the load has simply moved.

An analysis of the consulting sector describes a comparable shift: less producing and assembling, more contextualising, orchestrating and deciding. The analogy is useful for thinking about startup and scale-up roles, but it remains a transposition: a consulting engagement, a support workflow and a software production chain obviously don't carry the same risks or the same constraints.

Engineering: when producing code gets easier, checking what should be built matters more

Engineering is where the shift from producer to supervisor is best documented today.

What's documented

At Anthropic, CPO Mike Krieger has said that 90 to 95% of Claude Code's own code is produced by Claude Code itself. Engineers then spend more time steering, reviewing and integrating. The same case reveals another shift: as pull-request volume rises, the merge queue and decision-making capacity become new bottlenecks. This case is still rated medium confidence, though: Anthropic is an AI-native lab, not directly comparable to a typical scale-up.

Doctolib offers a different kind of signal. After a 2025 pilot phase, the company moved its 600 engineers to a mode described as "agentic-first" in early 2026. The documented change isn't only about tooling: the decision framework distinguishes reversible decisions, which builders can make without waiting on a product call, from irreversible decisions, which stay under the PM's control.

These examples converge on one idea: if producing becomes cheaper, knowing how to point that production in the right direction becomes more valuable.

But there's a hidden tax. Google Cloud's software-development research programme talks in 2026 about a genuine verification tax. Its qualitative analysis of 1,110 responses from Google engineers shows the time saved on writing is frequently reinvested in auditing and verification. More AI adoption is associated with more delivery throughput, but also with more instability.

Speeding up the producer without resizing the control system can simply move the bottleneck.

An engineer able to generate three times more changes doesn't necessarily improve the system if another engineer then has to spend three times more attention checking what was produced.

What this doesn't prove

These cases don't prove developers will stop writing code, or that every team should turn its engineers into "agent managers". The most striking situations involve organisations with high technical maturity. There, AI is described as an amplifier: a team with robust tests, good internal tooling and clear workflows can extract more value from it; a fragile organisation mainly risks accelerating its production of debt.

Dryve's read

The engineering role, then, doesn't simply shift from code to review. It shifts toward a broader set: understanding the problem, specifying it well enough, choosing what can be delegated, building validation criteria, examining edge cases, and owning production integration.

The management question is no longer just "how much code can our team produce?". It becomes: "how much production can we reasonably specify, control and absorb?"

Support: humans on the exceptions is credible, but it isn't automatically a better job

Support probably offers the most intuitive picture of the new division of labour. The machine answers the simple questions, the human takes the particular situations. On paper, the model seems obvious. On the ground, it's more ambivalent.

What's documented

At Alan, 40% of member requests are said to be resolved by AI, and up to 70% of the simple ones, according to a first-hand account. The member-relations team, around 100 people, is said to have stayed stable while membership grew 34%. The "Agent Care" role there is described as evolving toward more validation and oversight of generated responses. Here again, this is a company account, not an independent measure of ROI.

Older academic work offers an interesting counterpoint. In a study of 5,179 support agents, Erik Brynjolfsson, Danielle Li and Lindsey Raymond observed a productivity gain close to 14% with generative assistance, with the gains especially high among less experienced agents. But here humans kept doing the work: AI assisted them more than it replaced them with an exception-only workflow.

Two trajectories are therefore possible: an AI that augments the agent within the flow, or an AI that takes over the standard flow and surfaces only the exceptions.

The Klarna case is a reminder not to confuse the two. After heavily automating its customer service and reporting the equivalent of hundreds of agents' work handled by its assistant, the company reintroduced more explicit human capacity for complex situations. The CEO acknowledged that the drive to cut costs had outweighed quality by too much. AI wasn't abandoned: the model was recalibrated toward a more hybrid setup, keeping humans for certain sensitive cases such as identity theft.

What this doesn't prove

Klarna doesn't prove that "AI in support doesn't work". Alan doesn't prove that a support function can grow its volume indefinitely with the same headcount either. And the 5,179-agent study doesn't prove an assistant will deliver 14% productivity everywhere: it concerns one company, one technology and one specific environment. Above all, the three cases show that the boundary between machine and human is a variable to be designed, not a technological truth.

Dryve's read

The danger of the "human on exceptions" model is baked into its own name. If AI gradually strips away all the easy requests, the human work left over is no longer a representative sample of support. It becomes a concentrated stream of frustration, ambiguity, conflict, sensitive cases and problems the system failed to resolve. The number of cases drops, but their average cognitive and emotional intensity can rise.

The People question, then, can't just be "what share of tickets can we automate?". It also has to become: "what job are we actually building for the people who stay in the loop?"

Ops: the deepest change is sometimes not automating the processing, but automating the search for problems

The Roundtable case is especially instructive because it shows an important distinction between processing and sorting.

What's documented

Roundtable manages, among other things, investment SPVs across several European countries. According to an account gathered by TPC, a process previously checked file by file was automated with Claude Code: out of around 1,000 SPVs reviewed, only 15 to 20 problem cases are now said to reach a human — moving ops from mass processing to managing the 1-2% of exceptions.

Fleet offers a second signal. According to an account gathered by TPC, the company is said to have tripled its revenue since 2023 while keeping its supply and support headcount constant; two people, notably, are said to still be running the supply function thanks to the automation of a growing share of the work. Here again, the figures come from a company account, not an independent study.

What this actually proves

These two cases show that a sufficiently repetitive operational workflow can be restructured around a powerful principle: the machine examines the volume, the human examines the anomalies.

This principle can change the role far more profoundly than a simple copilot would. The operator no longer runs the same check a hundred times: they have to understand why certain files fall outside the norm, determine whether the anomaly is real, decide on the action, and possibly improve the rule that will let a similar case be handled automatically tomorrow. The work then starts to look like diagnosis.

What this doesn't prove

Neither Roundtable nor Fleet proves the model generalises to just any Ops function. These are relatively compact organisations, with a strong technical culture and leaders directly involved in building the automations. Nor does Roundtable prove automated checking reaches, unsupervised, the same quality as human checking — a limitation explicitly noted.

Dryve's read

The trap is treating the exception rate as a simple efficiency metric. It has to become an organisational metric. If 2% of volume now absorbs the entire human workload, you need to know who owns that 2%, how much time it takes, what risk it concentrates, which profiles know how to resolve it, and how the lessons feed back into the system.

Without that loop, the human simply becomes the dumping ground for the cases automation didn't understand.

Recruiting: automating the screening doesn't mean delegating the judgment

Recruiting shows a different variant of this transformation. Here, the most instructive case may be the one where the organisation specifically refuses to turn the recruiter into a mere algorithm supervisor.

What's documented

Lucca uses AI to match a job spec's criteria against the information in a CV. But the company explicitly states the recruiter stays the decision-maker at every step and can override the criteria or recommendations coming out of the analysis. Its own guide on CV matching flags several limitations: quality depends heavily on the input criteria, atypical CVs can be disadvantaged, and the gaps between automated ranking and human decisions need to be tracked to understand the system's errors.

Lucca is precisely one of the counter-examples to structural transformation: teams stay organised in a comparable way, and AI slots into their workflows rather than triggering a visible overhaul of roles.

What this doesn't prove

No data lets us claim Lucca reduces its need for recruiters thanks to this automation, or precisely measure the impact on time-to-hire or hire quality. The case shows a choice about how responsibilities are shared, not an organisational ROI.

Dryve's read

When screening gets easier, the quality of a hire depends even more on what happens before and after the screening. Before: defining what you're genuinely looking for. After: understanding a background, testing a capability, running an interview, weighing different opinions, and making a decision.

AI can therefore cut the time spent reading CVs while raising the value of a far less automatable skill: the ability to set good criteria, and to recognise when an interesting candidate specifically doesn't fit them.

Here, "supervising the AI" would almost be a poor description of the job. The recruiter remains, first and foremost, responsible for a human decision, with a machine helping prepare that decision.

Three assumptions the evidence contradicts

Not every role becomes a supervisory role

The shift of work toward supervision is documented at several companies. But the opposite is just as well documented: AI can be integrated with no significant restructuring of the role. Lucca keeps teams of six to eight people organised as before. Pennylane automates part of its accounting work while continuing to grow its teams and bringing chartered accountants into product design. Joko develops AI use cases while continuing to grow its headcount.

The claim that "AI necessarily turns producers into supervisors" is therefore too strong. A more defensible formulation: when AI genuinely takes over a significant share of a standard workflow, human work tends to concentrate on specification, review, orchestration, exceptions and decisions — but how much reshuffling actually happens depends on the role, the real level of automation, and organisational choices.

Supervising isn't necessarily higher-value work

The idea that AI would mechanically strip away the grunt work and leave humans with more interesting activities is appealing. It isn't demonstrated.

In engineering, faster production can create a verification overload, with identified risks to learning when AI lets people bypass, too early, the effort needed to build deep expertise. In support, humans can end up inheriting almost exclusively dissatisfied customers or situations impossible to standardise. In Ops, they can lose daily contact with the simple cases that used to give them an intuitive feel for how the system works. In recruiting, they can end up treating the algorithm's shortlist as objective reality when it actually reflects, first and foremost, the criteria it was given.

The shift toward supervision sometimes raises the value of the role. It can also raise its cognitive load, its accountability, or its monotony in a different way.

That's a matter of organisational design, not an automatic consequence of the technology.

Less human production doesn't mechanically mean fewer humans

This is a particularly tempting extrapolation. It isn't backed by the cases observed: Anthropic keeps hiring despite highly agentic code production, Lucca has kept hiring, Pennylane is growing its headcount sharply, bsport says it wants to keep hiring juniors despite automating part of its development and support work.

Productivity can be used to cut headcount. It can just as well be used to absorb more volume, improve quality, shorten lead times, launch new products, or move people to other activities.

The available data, then, supports talking about a redrawing of work with far more confidence than a general reduction in human work.

There isn't one supervisor role — there are four

The notion of supervision becomes genuinely useful once you break it down.

Form of supervision Human mission Dominant example Skill that gains value
Reviewer Check what the AI produced. Engineering Expertise, quality criteria, error detection.
Orchestrator Break down and distribute work between humans and agents. Engineering / Ops Workflow architecture, prioritisation, specification.
Exception specialist Resolve what the standard system can't handle. Support / Ops Diagnosis, judgment, handling ambiguity.
Decision owner Use AI without delegating the final decision to it. Recruiting / HR Criteria, accountability, contextualisation.

This distinction has an important consequence for managers and People teams.

You can't simply strip 30% of a role's tasks and assume the new role will appear on its own.

A reviewer needs the time it takes to review. An exception specialist needs training on the hard cases. An orchestrator needs decision rights that genuinely allow them to make trade-off calls. A recruiter accountable for the decision shouldn't be evaluated solely on the number of applications processed.

The redesign of the role starts precisely here: when responsibilities, skills and metrics are redefined around whatever work is left.

What this changes for managers

The first change is to stop conflating output with contribution. Someone who writes fewer replies, less code, or fewer files can be more decisive than before if they've become the control point for hundreds of automatically generated outputs.

The second is to explicitly manage review capacity. In many organisations, this used to be a secondary activity: everyone mostly checked work produced by a limited number of colleagues. With AI, production capacity can grow far faster than human attention capacity.

The third concerns decision rights. The framework documented at Doctolib is interesting precisely because it goes beyond "using agents": it tries to formalise which decisions can be made locally and which still require a higher-level call.

Finally, the manager has to protect learning. The risk is especially significant for less experienced employees. If all the standard tasks disappear too fast, you can strip away, at the same time, some of the repetition that used to build professional instincts. It's recommended to preserve deep-learning situations and to organise the review of generated decisions with experienced people — a topic that goes far beyond engineering.

An organisation can become more efficient today while weakening the mechanism that builds its experts of tomorrow.

Redrawing a role before automating it further

Before chasing a new automation rate, this checklist helps check whether the human role left behind has actually been designed.

An organisation that can't answer these questions probably hasn't redesigned the role yet. It has mainly automated part of its old workflow.

Conclusion

The conviction

The shift from producer to supervisor is a real signal. But taken literally, it hides the most important part of the transformation.

What AI automates is relatively easy to identify. What it leaves to humans takes far more design work.

In engineering, that remaining work can be specification, review and architecture. In support, it's often the sensitive and ambiguous situations. In Ops, diagnosis and anomalies. In recruiting, defining criteria and exercising judgment.

The problem, then, isn't only how far to automate. It's deciding what human role you build around the automation.

A good transformation isn't measured by the fact that humans now only touch 5% of the flow. It's also measured by the quality of that 5%, its workload, how clear the accountability is, the capacity to keep learning, and whether people can genuinely exercise judgment.

AI can absorb the standard work. The organisation still has to design the work that's left.

← Back to Scale-ups & startups

Want to talk it through with Dryve?

These positions shape the way we approach IT recruitment in the age of AI. Let's talk about your context.