They decide when a model deserves to be deployed, what limits should frame its use, and what trade-off to accept between performance, reliability, cost and real-world impact.
A Machine Learning Engineer neither simply trains a model nor simply exposes a prediction behind a service. The job is to turn a statistical capability into a usable system: the data has to represent the real problem, the evaluation has to measure the decision that matters, the preparation chain has to stay consistent between training and production, and the model's behaviour has to be tracked over time.
The difference between someone executing and someone genuinely holding the role appears when the results are convincing but the system is not. Strong offline performance can come from a misleading protocol. A technically correct model can push risk onto users, reproduce past decisions, or sit unused because it was never woven into the actual work.
The job therefore runs on permanent tensions: improve a metric or preserve robustness, serve in real time or precompute, automate or keep a human check. The assessment looks at the ability to make those tensions explicit and to decide on the total cost of the system, not just the theoretical quality of the model.
Market benchmarks
On the French market, the titles Machine Learning Engineer, AI Engineer, Applied Scientist and AI engineer cover very different realities. Some roles centre on modelling, others on industrialisation, operations, or designing systems built around generative models.
Demand remains sustained, but people able to cover the whole chain are fewer than those who can train a model in a controlled environment. Genuinely senior profiles stand out less for the range of architectures they know than for the hard calls they have already had to make.
Context
In an IT services firm, the assessment focuses on getting into an inherited estate of data and models quickly, telling real guarantees from presented results, and making the limits of a contractual scope explicit. The Machine Learning Engineer sometimes has to take over undocumented datasets or models with no traceability, while preparing a handover the client can actually work with.
Level 1
Junior
A Junior can prepare data, train a model within an established framework and check simple results. They apply a protocol they have been given, read the expected metrics and contribute to an existing chain. When the data is ambiguous or offline performance contradicts production, they still need someone to frame it.
Click to read
Level 2
Mid-level
A Mid-level engineer is autonomous on a common use case. They can compare approaches, check data quality, build a coherent evaluation, diagnose a degradation and contribute to going live. Their usual limit: they improve the model properly without always factoring in the running cost or whether users can actually act on the output.
Click to read
Level 3
Senior
A Senior engineer connects the data, the model, the service and the business decision. They question whether the features are genuinely available, what each type of error costs, whether the evaluation is representative, whether training and inference agree, and how drift and feedback loops behave. They know how to restrict a use case or force an abstention.
Click to read
Level 4
Expert
An Expert is not the person who knows more models or always squeezes out the best metric. It is the person who settles a structural trade-off while owning what the decision costs: giving up a performance gain to cut latency, suspending a visible but under-evaluated model, or keeping part of the process in human hands.
Click to read
The quality of the training set, whether the features are genuinely available, the labelling and the protocol together determine whether the performance you observe corresponds to the problem you will meet in production. A mistake here can make everything downstream misleading.
The Machine Learning Engineer has to pick a proportionate level of complexity, make sense of unexpected behaviour, and tell drift from a data defect, a feedback loop or a badly specified problem.
A model has to be reproducible, versioned, consistent between training and serving, monitored and maintainable. Going live is not a final step: it is part of designing the system.
Models can affect people, draw on data whose rights are unclear, or produce decisions that are hard to contest. This weighting means treating those subjects as skills of the role, not as a late external sign-off.
A good practitioner has to be able to reframe a request, explain the limits, and recognise when a rule or a change of process answers the need better than a model would.
The method in action
The scorecard at a glance — hover an axis
Data & artificial intelligence