They decide when data can be considered reliable, how it must be transformed and replayed, and what trade-off to accept between freshness, completeness, cost and continuity of service.
A Data Engineer does far more than move data between sources, a data lake, a warehouse and reporting tools. The job starts once the chain works technically but nobody can yet guarantee that the published data is complete, current, properly modelled and interpretable by the people consuming it.
The difference between someone executing and someone genuinely holding the role shows in how they handle silent failures. A job that finishes without error does not prove the data is right. A successful backfill does not prove nothing was duplicated. An available table says nothing about its freshness. The assessment is therefore less about making a data stack run than about establishing what the chain actually guarantees.
The job runs on permanent tensions: publish fast or wait for complete data, replay or rebuild, keep history or cut costs, tolerate schema drift or stop the feed. The Data Engineer has to make those tensions explicit, measure their consequences, and protect trust in the data without turning the platform into a rigid, over-engineered system.
Market benchmarks
The Data Engineer remains a highly sought-after profile on the French market, in end-user companies as much as in consultancies, IT services firms and SaaS vendors. The title nonetheless covers very different realities: some roles centre on ingestion and operations, others on analytical modelling, the platform, streaming or governance.
Genuinely senior profiles are rarer than people who can build pipelines within a given ecosystem. The difference lies in experience of incidents, imperfect historical data, migrations, definitional conflicts and cost trade-offs.
Context
In an IT services firm, the assessment focuses on grasping an inherited estate quickly, working with several stakeholders, and telling the technically desirable fix from what is contractually feasible. The Data Engineer has to make handovers safe, document the limits of something they did not design, and make responsibilities explicit between the client, the sources, the integrator and the consumers.
Level 1
Junior
A Junior can work on a bounded job, apply an existing model and fix an identified defect. They check the most visible outputs and are starting to understand the dependencies between sources, transformations and published tables. When the data is opaque or several figures contradict each other, they still need a framework to investigate.
Click to read
Level 2
Mid-level
A Mid-level engineer is autonomous within the usual scope. They can trace a chain back, identify a quality defect, reason about the grain of a table, make a job replayable and compare several ways of fixing it. They verify their results instead of treating a successful run as proof. Their usual limit: downstream consumers and the cost of rework are not sufficiently factored in.
Click to read
Level 3
Senior
A Senior engineer links the job to its use. They distinguish restoring the data from fixing the cause, technical freshness from the freshness actually expected, a useful optimisation from a computation that should no longer exist. They anticipate what a change will do to analysts, dashboards and business processes.
Click to read
Level 4
Expert
An Expert is not the person who knows more engines, schedulers or formats. It is the person who settles a structural trade-off while owning what the decision costs: accepting less immediate data to preserve its completeness, halting a publication rather than releasing a doubtful figure, or declaring part of the history unreliable.
Click to read
A data chain has to be diagnosable, recoverable and monitorable over time. A job that is fast but impossible to replay, or unable to spot missing data, is not operable.
Grain, historisation, data contracts and reconciliation determine what a figure means. A mistake at this level can stay invisible for a long time and affect every downstream use.
Volume, batch, streaming, storage and cost all have to be weighed against the actual use. Performance matters, but it never makes up for incorrect data or an architecture with no direction.
Access, retention, lineage, personal data and the ability to delete determine how much control you really have over your data estate. Their weight can rise sharply in a heavily regulated environment.
Reliability also depends on relationships with producers, analysts and business teams. An implicit contract or a badly communicated incident produces failures that technique alone will not fix.
The method in action
The scorecard at a glance — hover an axis
Data & artificial intelligence