Our stance on AI · Scale-ups & startups

When AI handles the simple, humans inherit the hard: organising exceptions without wearing teams out

The right metric isn't the automation rate — it's the quality of the system that handles what it can't do

Automating simple requests seems to produce an obvious equation: less volume for teams, so less load. But that reading forgets what's left over. When AI absorbs the predictable cases, humans mechanically end up with a higher proportion of ambiguous, contentious, unprecedented or high-stakes situations.

The "human on exceptions" model can be very effective. It's already observable at several organisations. But it only holds up durably if the exception is designed as an organisational object in its own right: with escalation rules, decision rights, quality metrics and a learning loop. Otherwise, the company risks simply shifting the load, from volume to complexity.

The right indicator isn't the percentage of tasks automated. It's the quality of the system that handles what automation can't do.

Summary

In brief

The trap: managing the automation rate instead of the work left over

In many automation projects, the first indicator people look at is immediately understandable: 30%, 50%, 70% of requests now handled by AI.

That figure says something about the volume avoided. It says almost nothing about the organisation left behind it.

At PayFit, according to an account gathered by TPC, rolling out its Copilot came with a 30 to 35% drop in requests reaching support. But the remaining tickets are more complex. That's precisely the shift to watch: the team no longer has the same job. Before, a day could mix easy questions, medium cases and a few hard files. After automation, human activity can end up made up of far more uncertainty, interpretation and judgment calls.

The same mechanism shows up, in a much more advanced form, at Roundtable. The company says it automated checks on 1,000 SPVs so that only 15 to 20 problem cases now reach operations — around 1 to 2% of volume. The case is especially interesting because it involves a regulated activity and the company keeps a significant legal team. But this account doesn't prove that automated sorting reaches human quality unsupervised.

That's the difference between two questions:

How many files does AI no longer pass on to humans?

Among the files it no longer passes on, how many should it have?

The second question is far harder to measure. And far more important.

Automation can reduce the load: the shift toward exceptions isn't inevitable

It would be tempting to infer a new organisational law from this: the more AI handles the simple cases, the more it makes human work gruelling. The data doesn't support that claim.

Brynjolfsson, Li and Raymond's study of 5,179 support agents is a useful counter-example. There, AI assistance raises average productivity by 14%, with a 34% gain for novice or lower-performing agents, while its effect is far smaller among already highly experienced staff. The researchers also find positive signals on customer behaviour and new-hire retention — while stressing the study doesn't let us infer aggregate employment effects.

AI, then, can make a job less difficult when it supports the human instead of simply stripping away the simple cases.

A systematic review published in 2026 on the cognitive load of healthcare professionals using various forms of AI also finds mixed effects: certain documentation tools are associated with reduced load, while decision-support or imaging systems give more mixed results, sometimes with increased load. The medical context obviously doesn't transpose directly to support or operations, but it's a reminder to avoid a simplistic equation between automation and wellbeing.

What matters is how human work and automated work are assembled together.

Klarna: neither "AI replaced support" nor "Klarna backed off"

The Klarna case is especially instructive because it's been used to back two opposing narratives: the first has AI massively replacing human support; the second has Klarna trying automation, finding it failed, then reversing course. The available facts tell a more interesting story.

In 2025, CEO Sebastian Siemiatkowski publicly acknowledged that the drive to cut costs had outweighed things too much, and stressed the need to be able to offer quality human support. Klarna is, then, one of the important counter-examples to an "all-AI" model.

But the company's more recent regulatory filings show support automation wasn't abandoned in the process. Klarna states its assistant handled 80% of customer-service conversations in 2025. The company also states, based on its own logs and internal surveys, that satisfaction didn't decline, that repeat requests fell 25% after launch, and that average resolution time was two minutes for the assistant versus twelve for human agents in 2024. This data is produced by Klarna itself and so isn't an independent evaluation.

The reasonable conclusion, then, isn't that "the human won against the AI". It's that automating volume and keeping a human available for the moments that matter may both need to coexist.

That's exactly what Klarna now describes as a two-track approach: keep automating broadly while still leaving customers the option to reach a human.

This case directly challenges this article's thesis: the human-on-exceptions model can work at very large scale without necessarily producing more repeat contacts or degrading self-reported satisfaction metrics. But it reinforces the thesis on one point: the boundary between what gets automated and what's handed to a human is a service-design decision, not the leftover of a cost-cutting target.

The real risk: turning the human into automation's after-sales service

The problem starts when humans receive three types of work all at once.

The genuinely difficult cases: a new situation, an ambiguous rule, an emotionally charged request, or a decision with major consequences.

Automation's normal limits: missing information, low confidence, context impossible to interpret.

Errors produced by the automation itself: wrong classification, incorrect answer, poor execution, or a file closed too early.

These three categories are often lumped together into the same "escalations" queue. Yet they have nothing to do with each other. In the first case, the human brings exactly the skill you want to keep. In the second, they complete the system. In the third, they're doing rework: fixing a piece of work the organisation had counted, too early, as automated.

This is where the automation rate becomes misleading. An organisation can report 70% of files automated while a significant share comes back later as a reopening, a complaint, a correction, or an escalation.

What Dryve proposes measuring: useful automation

Useful automation = files resolved automatically with no human correction or reopening within the relevant quality window / total volume.

This isn't an existing standard: it's a Dryve reading we're proposing. It changes the conversation. A 55% automation rate with very few errors can outperform an 80% rate that generates a lot of rework.

When checking becomes a job in itself

The second risk concerns oversight. Writing "the human checks" into a process is easy. But checking a plausible-looking answer produced by AI isn't necessarily simpler than producing the right answer yourself.

An interesting phenomenon shows up in knowledge work: agentic tools can deliver a first version that's already largely structured, which shifts the moment when critical thinking gets applied. Instead of reasoning step by step during production, the professional now has to inspect an output that arrives all at once and spot what's wrong with it.

The human-factors literature suggests taking this shift seriously. A systematic review of work on automation bias concludes that errors linked to trusting the system are notably associated with verification complexity and cognitive load. Interventions that simply remind users to stay alert seem to have limited effect.

More broadly, research on supervised systems has long described an "out-of-the-loop" problem: when a human only steps in rarely, they can find it harder to quickly reconstruct the situation when an unusual incident occurs. Experiments in supervisory settings, in particular, show reduced engagement can slow reactions when automation suddenly requires intervention.

These results come notably from fields like aviation or industrial control: they don't prove a support agent will suffer the same effects. They still offer a useful warning.

Reserving humans for rare events doesn't guarantee they'll stay able to handle them properly.

"Human-in-the-loop" isn't enough: the human needs authority

Another anti-pattern shows up when the organisation nominally keeps a human in the loop, but strips away the means to actually exercise judgment.

An employee then receives an AI recommendation, has to "validate" it, but has little time, doesn't know the sources used, doesn't understand why the case was routed to them, and can barely change the decision. They stay legally or managerially accountable without holding real authority. That's a bad allocation of responsibility.

Recent research on Human-in-the-Loop systems is, in fact, a reminder that implementing them isn't only about a human's presence: cognitive load, trust calibration, and where the interaction sits within the workflow remain significant problems.

A genuine human-on-exceptions model requires at least four powers: seeing — accessing the context and reasoning available; contradicting — being able to reject or change the recommendation; stopping — when a system degrades quality or presents a risk; and teaching — turning a correction into information the system can actually use.

Without this, the human isn't in the loop. They're at the end of the chain.

Not every exception should go into the same queue

A financial dispute, an emotionally sensitive question, a rare technical error and a case that's simply missing information can all need human intervention. Yet they require different skills, different urgency, and different levels of authority.

As the automation rate rises, exception segmentation therefore needs to get finer.

Exception type Why AI hands off Organisational response
Missing information Insufficient context. Additional data collection or a quick contact.
Low confidence Uncertain result. Validation by a trained operator.
Complex business case Multiple rules or trade-offs. Domain expert.
Sensitive / high-stakes case Financial, legal, human or reputational impact. A decision-maker with explicit authority.
New case A situation absent from the rules or the data. Expert + person owning the learning loop.
Automation error Wrong decision or poor execution. Correction + root-cause analysis + eval updates.

This distinction also changes how skills are managed. A new case isn't just a ticket to close: it's potentially information about what the company can't yet automate.

The learning loop is what turns the exception into an advantage

In a poorly run organisation, the same type of exception comes back every week and someone handles it manually. In a learning organisation, that exception becomes data.

The employee first fixes the case. But they also flag why the automation failed: missing data, an absent rule, wrong classification, outdated knowledge, a poorly calibrated confidence threshold, an action that couldn't be executed.

That information can then update a rule, a documentation source, a prompt, a workflow, or an eval set. Once the fix is tested, some cases gradually rejoin the automated flow.

This is precisely the condition the "human on exceptions" model sets: a correction-and-learning loop has to exist, and volume, exception, error and rework metrics have to be tracked. Without them, there's no way to tell whether the company is genuinely automating or simply shifting the load.

The organisational stake here is significant: someone has to own this loop.

If operators fix the cases but nobody turns their corrections into system improvements, the AI doesn't learn from the company. The company simply learns to live with the AI's errors.

The dashboard to track is no longer the classic support one

A human-on-exceptions model can't be steered with just "number of tickets" and "average handling time". The composition of the work has changed.

Indicator What it actually measures Warning signal
Automated share Volume that no longer reaches a human up front. Rises with no improvement in the other indicators.
Exception rate Share passed to humans. Unexpected rise or very high variability.
Escaped errors Cases handled automatically that should have been escalated. This is the least visible risk.
Rework Human corrections of work already produced by AI. The apparent automation gain is overstated.
Reopening / repeat contact A resolution that didn't actually solve the problem. Quality decline masked by the automation rate.
Time per exception Effort required by the cases left to humans. Rises despite falling volume.
Perceived cognitive load Experienced difficulty, focus and fatigue. Worsens even as volumes fall.
Learning loop Share of recurring exceptions that led to a system fix. The same causes keep coming back.

The number a COO should care about, then, is no longer just cost per ticket. It's also the total cost of the exception: human time, delay, rework, the level of expertise deployed, the risk created, and how often the same problem repeats.

What the data lets us claim, and what it doesn't

What the sources let us claim: there are organisations where AI absorbs a significant share of standard work; human roles can shift toward supervision and exceptions; the tickets left to humans can become more complex; AI assistance can also improve performance and experience for some employees; and, finally, a hybrid model can keep automating heavily while still maintaining human access.

What they don't let us claim: that the "human on exceptions" model systematically cuts headcount; that it systematically increases fatigue or burnout; that a high automation rate means better quality; that the figures a company publishes are enough to isolate AI's causal effect; or that an organisation observed at a fintech, a healthtech, or a regulated SME can be copied as-is elsewhere.

The body of cases observed is, in fact, explicit on this point: its thirty companies were selected for the organisational signals they let us study, with no claim to being representative.

Checklist: does your organisation genuinely handle exceptions?

Conclusion

The conviction

Automating the simple part is only half the job.

The second half is organising what's left.

A company can automate 70 or 80% of a flow and build an excellent system. It can also reach the same figure while leaving behind humans who fix errors, absorb conflict, and make hard decisions under pressure.

The percentage is identical. The organisation isn't.

The "human on exceptions" model becomes genuinely valuable when exceptions are visible, measured, routed to the right people, and used to learn. In this model, the human isn't AI's spare wheel. They remain the part of the system the organisation entrusts with whatever it still believes needs context, judgment and accountability.

← Back to Scale-ups & startups

Want to talk it through with Dryve?

These positions shape the way we approach IT recruitment in the age of AI. Let's talk about your context.