Companies Candidates

Site Reliability Engineer (SRE)

They decide what level of reliability a service actually has to guarantee, which production risks the organisation accepts, and when stability must take precedence over speed of change.

The role at a glance

The role is not about watching dashboards, stepping in during incidents, or automating repetitive operations. The SRE turns reliability into a manageable property of the product: they define the indicators that genuinely represent the service delivered, set service objectives, measure whether they are met, and use that data to weigh new features against reliability debt, capacity and operational risk.

The boundary with neighbouring roles is deliberately strict. The Cloud Infrastructure Engineer builds the foundations. The DevOps Engineer smooths and automates the path from code to production. The Platform Engineer turns technical capabilities into an internal product. The SRE builds on those foundations, but their object is different: protecting the reliability and performance of services in production, from the point of view of the people using them.

Someone executing can handle an alert or restart a component. Someone holding the role knows whether the alert represents a real degradation, whether the service can continue in a degraded mode, and how much downtime the organisation accepts in order to keep shipping. The central tension of the job pits reliability against speed of change.

Market benchmarks

The SRE title still covers very different realities. Some organisations use it to describe modernised operations or a production-oriented DevOps team; others have genuinely put service objectives, error budgets and shared responsibility with product teams in place.

People who can combine software engineering, operational judgement, incident management and product trade-offs remain rare. Command of an observability stack places someone's experience, but predicts poorly whether they can define a useful objective or keep an on-call rota sustainable.

Context

In an IT services firm, the assessment leans on working inside an inherited estate, with contractual targets sometimes disconnected from the real architecture and limited room to decide. The SRE has to make the risk visible, tell the contractual commitment from the reliability actually achievable, and hand over something operable once the engagement ends.

What separates the levels

Level 1

Junior

Works on a service whose objectives, alerts and operating procedures are already defined. They can read the available signals, qualify a degradation, apply a stabilising measure and ask for help when the impact goes beyond their scope. What separates them from a pure operator is that they check the effect of their action on the service delivered, rather than merely noting that a component has gone green again.

Market benchmark Rarely entrusted at this level

Click to read

Level 2

Mid-level

Owns the reliability of a service end to end. They connect technical indicators to the user experience, improve observability, prepare degraded behaviours, analyse recurring incidents, and reduce the manual operations that eat into on-call. They make failures comprehensible, capacity measurable and recovery verifiable.

Permanent salary · Paris region 50 – 60 k€
Freelance day rate · Paris 470 – 590 €/day

Click to read

Level 3

Senior

Arbitrates between service objectives, pace of change, available capacity, infrastructure cost and human load. They can tell an acceptable degradation from a structural risk, use the error budget to make decisions explicit, and challenge an availability target that is disproportionate to the value of the service. They think in overall reliability, dependencies and blast radius.

Permanent salary · Paris region 62 – 85 k€
Freelance day rate · Paris 570 – 750 €/day

Click to read

Level 4

Expert

Is not distinguished by a bigger catalogue of monitoring tools. They settle the most expensive trade-offs and name what they agree to give up: temporarily slowing releases, cutting a service's functional scope, accepting a less ambitious availability target, or switching off a feature to protect the integrity of the system.

Permanent salary · Paris region 80 – 100 k€
Freelance day rate · Paris 750 – 950 €/day

Click to read

The skills we assess

05
Skill categories
assessed in real situations
Hover a category to see what it measures — and why it carries weight
Service objectives & reliability governance
The role's first priority
What it measures

Without representative indicators, service objectives and an error budget, reliability stays a subjective impression — invoked after every incident but useless for deciding before the next one.

Incidents, observability & operations
The quality of the reasoning
What it measures

The SRE has to be able to qualify a degradation, stabilise the service and run an investigation under pressure. This weighting measures the quality of the reasoning and of the available signals, never how fast someone restarts a component.

Resilience, capacity & performance
Before the incident
What it measures

Reliability is designed before the incident: failure modes, dependencies, saturation, latency, recovery and degraded behaviours all have to be understood at service level.

Reliability automation & reducing toil
Deliberately lower
What it measures

Automation counts when it removes an operation that is repetitive, risky or of no lasting value. Automating a bad practice does not turn it into reliability engineering.

On-call, collaboration & learning
A shared responsibility
What it measures

Reliability is a responsibility shared with the development and product teams. This category checks that the candidate can keep on-call sustainable without creating an SRE team that absorbs all the consequences on its own.

The method in action

The method in action

Y symbol
Annotated interview excerpt

The scorecard at a glance — hover an axis

Cloud, platform & reliability

Other roles in this family

Cloud Infrastructure Engineer DevOps Engineer Platform Engineer