They decide what level of reliability a service actually has to guarantee, which production risks the organisation accepts, and when stability must take precedence over speed of change.
The role is not about watching dashboards, stepping in during incidents, or automating repetitive operations. The SRE turns reliability into a manageable property of the product: they define the indicators that genuinely represent the service delivered, set service objectives, measure whether they are met, and use that data to weigh new features against reliability debt, capacity and operational risk.
The boundary with neighbouring roles is deliberately strict. The Cloud Infrastructure Engineer builds the foundations. The DevOps Engineer smooths and automates the path from code to production. The Platform Engineer turns technical capabilities into an internal product. The SRE builds on those foundations, but their object is different: protecting the reliability and performance of services in production, from the point of view of the people using them.
Someone executing can handle an alert or restart a component. Someone holding the role knows whether the alert represents a real degradation, whether the service can continue in a degraded mode, and how much downtime the organisation accepts in order to keep shipping. The central tension of the job pits reliability against speed of change.
Market benchmarks
The SRE title still covers very different realities. Some organisations use it to describe modernised operations or a production-oriented DevOps team; others have genuinely put service objectives, error budgets and shared responsibility with product teams in place.
People who can combine software engineering, operational judgement, incident management and product trade-offs remain rare. Command of an observability stack places someone's experience, but predicts poorly whether they can define a useful objective or keep an on-call rota sustainable.
Context
In an IT services firm, the assessment leans on working inside an inherited estate, with contractual targets sometimes disconnected from the real architecture and limited room to decide. The SRE has to make the risk visible, tell the contractual commitment from the reliability actually achievable, and hand over something operable once the engagement ends.
Level 1
Junior
Works on a service whose objectives, alerts and operating procedures are already defined. They can read the available signals, qualify a degradation, apply a stabilising measure and ask for help when the impact goes beyond their scope. What separates them from a pure operator is that they check the effect of their action on the service delivered, rather than merely noting that a component has gone green again.
Click to read
Level 2
Mid-level
Owns the reliability of a service end to end. They connect technical indicators to the user experience, improve observability, prepare degraded behaviours, analyse recurring incidents, and reduce the manual operations that eat into on-call. They make failures comprehensible, capacity measurable and recovery verifiable.
Click to read
Level 3
Senior
Arbitrates between service objectives, pace of change, available capacity, infrastructure cost and human load. They can tell an acceptable degradation from a structural risk, use the error budget to make decisions explicit, and challenge an availability target that is disproportionate to the value of the service. They think in overall reliability, dependencies and blast radius.
Click to read
Level 4
Expert
Is not distinguished by a bigger catalogue of monitoring tools. They settle the most expensive trade-offs and name what they agree to give up: temporarily slowing releases, cutting a service's functional scope, accepting a less ambitious availability target, or switching off a feature to protect the integrity of the system.
Click to read
Without representative indicators, service objectives and an error budget, reliability stays a subjective impression — invoked after every incident but useless for deciding before the next one.
The SRE has to be able to qualify a degradation, stabilise the service and run an investigation under pressure. This weighting measures the quality of the reasoning and of the available signals, never how fast someone restarts a component.
Reliability is designed before the incident: failure modes, dependencies, saturation, latency, recovery and degraded behaviours all have to be understood at service level.
Automation counts when it removes an operation that is repetitive, risky or of no lasting value. Automating a bad practice does not turn it into reliability engineering.
Reliability is a responsibility shared with the development and product teams. This category checks that the candidate can keep on-call sustainable without creating an SRE team that absorbs all the consequences on its own.
The method in action
The scorecard at a glance — hover an axis
Cloud, platform & reliability