Our stance on AI · Scale-ups & startups

Managing humans and agents: who decides, who checks, who's accountable?

In the age of agents, management becomes an architecture of authority

As AI agents move from suggesting to acting, the manager's problem changes in nature. It's no longer just about splitting work among people, but about defining what a machine can decide on its own, what needs checking, when a human needs to step back in, and who stays accountable for the outcome.

The available cases don't sketch out one single model. They do reveal a useful dividing line, though: the more reversible, observable and bounded a decision is, the more its execution can be delegated; the harder it is to undo, the more ambiguous it is, or the more significant its consequences, the more human authority needs to stay explicit.

This is probably where one of management's main shifts lies in the age of agents. The manager shouldn't approve everything, at the risk of becoming the new bottleneck, nor "let AI do its thing" in the name of autonomy. Their job increasingly consists of designing the decision-making system humans and agents operate within.

Summary

In brief

The manager no longer just hands out tasks: they hand out decision rights

With a classic copilot, the boundary seemed simple: the machine suggested, the human did the work.

With an agent able to look up information, change a database, write code, close a ticket, prepare a customer reply, or trigger a workflow, that boundary becomes insufficient. Two agents using the same model can have radically different autonomy levels. One drafts something nobody will see before approval. The other directly changes a production system.

The right question, then, isn't "is our agent autonomous?", but: "on which decisions does it hold the authority to act, within what limits, and under what control?"

This distinction already shows up in several cases observed. At Stripe, according to an account gathered by TPC, an agent can act as the first code reviewer, ahead of the human pass. At Alan, non-engineers can produce changes, but the merge decision stays on the engineering side. At Finary, agents produce components, but the developer keeps the final say.

These organisations don't all apply the same method. They do converge on one principle, though: production, checking, and the final decision can be distributed across several different actors.

This is a significant change for management. Yesterday, delegating often meant giving a person a goal and assigning them accountability. Tomorrow, the manager also needs to define:

Autonomy becomes a property of the workflow, not a general feature of the tool.

Doctolib: the most interesting distinction may not be human versus AI, but reversible versus irreversible

Among the organisations studied, Doctolib offers one of the most directly usable frameworks for a manager.

The company redefined "who makes the decision" within its product organisation. The PM keeps decisions classed as irreversible — priorities between projects, certain legal or structural choices — while product builders can decide reversible matters such as handling an edge case or certain interface trade-offs. The rule reported is pragmatic: when a decision can be undone with a pull-request revert, there's no need to systematically wait on the PM.

What this case actually shows: an organisation can increase its contributors' autonomy, whether human or agent-augmented, by pushing certain decisions down when their reversal cost is low.

What it doesn't show: that every technically reversible decision has no consequences, or that the framework transposes as-is to recruiting, finance, or customer support. Nor is there a measurement letting us attribute a precise productivity gain to this decision rule alone.

A decision becomes a better candidate for agentic autonomy when it combines several properties.

Question Easier autonomy Reinforced human authority
Can the action be easily undone? Yes No, or with difficulty
Will an error be detected quickly? Yes No
Is the affected scope limited? Yes Wide
Is there an objective success criterion? Yes Ambiguous judgment
Are there enough precedents? Yes New situation
Does the action heavily commit a customer, an employee, or the company? Not much Heavily
Is there a specific regulatory, contractual or ethical requirement? Low High

Reversibility, then, isn't a sufficient rule on its own. It's the first axis of a delegation matrix.

"Human in the loop" doesn't mean putting a human in front of every button

A common reaction to agent risk is adding human validation everywhere. It's reassuring. But it can also recreate exactly the problem automation was supposed to solve.

At several organisations where agents produce more, the bottleneck shifts toward review, evals, and decisions. Anthropic is presented as a case where code production rises enough to shift the constraint toward decisions and the merge queue. Alan likewise accepts that review becomes a central control point.

An analysis of the consulting sector frames the same issue from a different angle: when AI-assisted production speeds up, the risk doesn't disappear; it concentrates more heavily on whoever has to defend or sign off on the result. It therefore recommends making assumptions, areas of uncertainty, controls and accountability explicit, rather than confusing production speed with reliability.

The managerial consequence matters: the manager shouldn't become the universal reviewer of agents.

Their role is instead to decide which type of control matches which risk. A control can take several forms:

NIST specifically recommends documenting human-oversight roles, organisational responsibilities, output evaluation, incident procedures, and the mechanisms for disabling a system when that becomes necessary.

The managerial question, then, becomes less "who reviews it?" than: "what assurance do we need before letting this action take effect?"

The five autonomy levels a manager can actually manage

Rather than a binary choice between "human" and "autonomous", a team can assign an autonomy level to each workflow.

Level 1 — The agent informs

The agent searches, synthesises, or flags. It makes no operational decision. Example control: spot-checking source quality and coverage.

Level 2 — The agent proposes

The agent prepares a reply, a diagnosis, a ranking, or an action. A human always decides whether to execute it. This is the model Lucca explicitly claims for several HR use cases: AI can pre-analyse or propose, while the recruiter or manager stays the decision-maker.

Level 3 — The agent executes after validation

The system prepares the whole piece of work, but human authorisation is still needed before the action takes effect. This level fits especially well when production can be automated but authority shouldn't be.

Level 4 — The agent executes alone within a bounded scope

The agent can act with no prior validation once the defined conditions are met. The action then needs to be observable enough, and the limits precise enough: type of operation, scope, threshold, accessible systems, escalation criteria. Human control shifts from every single operation toward the system's overall functioning.

Level 5 — The agent orchestrates a workflow, the human steps in on exceptions

The agent chains several operations together and only calls on a human in case of uncertainty, an anomaly, or a threshold being exceeded. Roundtable offers a telling example of this logic: according to an account gathered by TPC, a process covering around 1,000 SPVs surfaces only 15 to 20 problem cases to a human. The case's interest lies less in the figure, which comes from an account rather than independent validation, than in the division of labour: the machine sorts the volume, the human focuses their time on the exceptions.

None of these levels is "better" than the others. The right level depends on the workflow's risk.

A company can perfectly well have a highly autonomous agent on a routine internal operation while keeping an "AI proposes, the human decides" mode for a far more sensitive decision.

What Klarna reminds us: more autonomy isn't always more performance

Klarna's recent history offers a useful counterpoint to a linear view of automation.

The company had communicated very aggressively about automating its support and cut its headcount, in a context combining a hiring freeze, attrition, and AI's rise. It then moved back toward a hybrid model, with the CEO publicly acknowledging the drive to cut costs had outweighed things too much and that quality had degraded on certain complex or high-value cases. The company didn't abandon automation: it brought back human capacity for the moments it judged that presence necessary.

This case doesn't prove agents are worse than humans at support — the data Klarna published also reported improvements on certain indicators, and headcount or quality changes can't simply be attributed to AI alone. It shows something else:

optimising a workflow for its automation rate isn't the same as optimising it for its overall outcome.

A manager should therefore be wary of a KPI like "80% of decisions are made with no human". Taken in isolation, it says almost nothing. It needs to be paired with, at minimum:

These are exactly the metrics the "human on exceptions" model considers necessary for telling a genuine gain apart from a simple shift in load.

Who's accountable when the agent gets it wrong?

This is probably the most uncomfortable question, and the most important one. Two notions need distinguishing first.

Operational accountability answers the question: within our organisation, who needs to make sure this workflow works, and react when it fails?

Legal liability depends on the context, the use case, the contracts, and the applicable law. It can't be reduced to a general rule where "the manager is legally liable for everything the agent does".

The EU AI Act, for example, allocates obligations across different parties in the value chain and requires human-oversight mechanisms for certain high-risk categories. The text stresses AI that can genuinely be controlled and supervised by people; these obligations, however, don't apply indiscriminately to every assistant or agent a company uses.

For the manager, the organisational rule can be simpler:

An agent can be given autonomy. It shouldn't become the abstract owner of an outcome.

Every agentic workflow should have an identifiable human owner. That doesn't mean this person has to check every operation. It means they're responsible for designing or evolving:

One trap would be officially keeping a human accountable while, in practice, stripping away the means to exercise that accountability: no visibility into the actions, no logs, too many operations to check, no way to stop the system.

Oversight that's purely nominal isn't oversight.

The hybrid manager has four new responsibilities

Define the decision territory

The manager needs to separate what can be delegated, what can be delegated under conditions, and what needs to stay human. This map shouldn't be set once and for all: autonomy can grow after observing performance, or shrink after an incident.

Design the boundaries, not just the instructions

With an employee, much of the constraint stays implicit: experience, shared culture, judgment. An agent needs far more explicit limits. The accessible systems, the authorised data, the relevant financial or operational thresholds, the escalation criteria, the forbidden situations, and how to interpret uncertainty all need defining.

Organise quality assurance

Several cases observed show a shift of work toward review and evals. The manager's job isn't necessarily to run these checks themselves, but to make sure they exist, are sized correctly, and that someone holds the authority needed when the criteria aren't met.

Turn errors into system improvements

In a classic organisation, an error can lead to individual feedback. With an agent, you also need to ask: was the context provided insufficient? Was a rule missing? Should the eval have caught the error? Was the escalation threshold too high? Did we grant this autonomy level too soon? Did the agent have access to an action it shouldn't have been able to perform?

The post-mortem, then, no longer focuses only on who made the mistake, but on why the control architecture let the error through.

Watch for the second-order effect: humans inherit the hardest decisions

A heavily automated organisation can give the impression the human role becomes more interesting: fewer repetitive tasks, more judgment. That's possible. But it isn't automatic.

The "human on exceptions" model carries a paradox: as the agent gradually eliminates the standard cases, the human's day can turn into an almost unbroken string of ambiguous, emotionally charged, or risky cases.

Relegating the human to exceptions with no authority and no learning loop is a recognised anti-pattern. For the manager, this introduces a new capacity problem.

Before, ten people could handle a hundred fairly uniform files. Tomorrow, the agent might handle ninety of them. But the remaining ten don't necessarily represent 10% of the effort. They can concentrate most of the stress, the judgment, and the accountability.

Team sizing, then, can no longer rest solely on the leftover volume. It also needs to track:

A hybrid team's manager, then, manages two different capacities: automated production capacity, and human capacity for absorbing uncertainty.

What the data does not yet let us claim

Dryve's read: the real issue isn't the agent's autonomy, it's the architecture of authority

The mistake would be treating agents like new employees to whom you'd simply apply existing delegation methods.

A human holds tacit judgment, generally understands an action's social consequences, and can be asked about their intent. An agent works differently. Part of management's implicit rules, then, needs replacing with an explicit architecture of authority.

In the age of agents, the quality of management will be measured less by how many decisions the manager keeps for themselves than by their ability to place each decision at the right level: automatic when it's bounded enough, delegated when it can be, escalated when it becomes uncertain, and owned by a human when it genuinely commits the organisation.

This also changes the notion of human autonomy. Doctolib is interesting precisely because the same logic that lets you work with agents can also give employees more authority back: if a decision is reversible, why systematically wait on the manager?

The paradox, then, is interesting. The agentic organisation could produce more centralisation, because every action seems to need checking. Or, conversely, it could force the company to make its decision rights more explicit and allow more decentralisation.

The tool doesn't settle which way it goes. Managerial design does.

Playbook: putting an agent under accountability in ten questions

Before an agent moves from an experiment to a real workflow, the accountable manager should be able to clearly answer the following questions.

A delegation sheet to use for every agent

Element To document
Human owner The person accountable for how the workflow runs.
The agent's mission Expected outcome.
Accessible data Authorised sources and information.
Authorised actions What the agent can actually do.
Forbidden actions Boundaries that can't be crossed.
Reversibility level Simple / costly / difficult / impossible.
Autonomy thresholds Volume, amount, scope, or uncertainty level.
Validation Before, after, by sampling, or automated.
Escalation When, and to whom.
Evals / KPIs Quality, errors, exceptions, rework, lead time, cost.
Emergency stop Who can suspend the agent, and how.
Learning loop How errors change the system.
Conclusion

The conviction

Hybrid management probably won't be a job spent watching agents work all day.

It will be more a job of designing accountability.

The point isn't choosing once and for all between "human decision" and "AI decision". It's building an organisation able to gradually grant autonomy where the risks are manageable, while keeping real human authority wherever uncertainty, impact, or irreversibility demand it.

The best managers, then, won't necessarily be the ones who check the most. They'll be the ones who can answer, for every workflow, three simple questions:

Who can decide? Who needs to check? Who needs to be able to say stop, and own what happens next?

← Back to Scale-ups & startups

Want to talk it through with Dryve?

These positions shape the way we approach IT recruitment in the age of AI. Let's talk about your context.