Before the detail, here's the challenge you'd help us solve.
We build the embodied intelligence that moves real vehicles safely, and the ecosystem a billion machines will run on in the future. Very few people in AI can say this. Every role here, whatever the team, plugs into that.
Here’s what this particular role covers.
The role
Wayve ships a new driving model baseline every week. The Model Integration & Release team runs simulation and on-road testing to decide whether candidates are ready to promote — and there is a significant opportunity to get more from the data we already collect. As our release process matures, the next step is building a deeper analytical layer: strengthening how we measure performance, turning on-road findings into better simulation tests, and aligning what we measure with what operators experience in the vehicle.
This is a senior, high-impact role on the Model Integration & Release team. You'll own that analytical layer — turning evaluation outputs into findings that improve both our models and how we measure them. You'll work across simulation, on-road experiment data, and our internal evaluation tooling, partnering with Validation, Data Science, and Product. You'll have the autonomy and scope to improve how we evaluate — from identifying gaps in our measurement approach through to implementing changes that make our release decisions more confident and our feedback loops faster.
Key Responsibilities
Shape how we learn from evaluation data — define what “going deeper” means in practice, and build the habits, tooling, and workflows that make it part of how we release models
Improve how we measure driving performance — identify blind spots, inconsistencies, and gaps in our simulation and offline metrics; drive improvements to our evaluation methods, suites, and scoring logic
Close the loop between on-road testing and offline evaluation — investigate what happens on-vehicle during release testing, determine whether our offline tests should have caught it, and turn those findings into concrete improvements to coverage and measurement
Perform day-to-day experiment analysis and apply statistical rigor to evaluation data to provide confident, data-driven promotion recommendations for new driving models
Expand what “good” means beyond intervention rates — develop and apply a richer view of model quality using the behavioural and operational signals we already collect
Partner with Data Science and Validation on how we define, implement, and maintain evaluation methods and ground truth
Support partner-facing quality investigations — help quantify and track issues raised through OEM QA workflows so they can be addressed within our release process
Turn investigations into durable improvements — document findings and recommendations in a way the team can act on within a weekly release cadence, and follow through so the same class of issue doesn’t recur
About you
In order to set you up for success as a Senior Data Scientist at Wayve, we’re looking for the following skills and experience.
Essential
Strong analytical and investigative skills — you’re comfortable going from a vague “something looks off” to a clear, evidence-backed conclusion
Hands-on experience working with ML evaluation data, driving performance metrics, or complex test results at scale, including statistical methods for A/B testing and experimentation (e.g., confidence intervals, distributions, and sampling)
Proficiency in Python or SQL; comfortable querying and analysing large datasets
Ability to operate with autonomy in a fast-moving environment — you can set your own priorities within a broader goal, and know when to fix something yourself vs. when to escalate or partner
Clear written communication — you turn messy investigations into concise, actionable recommendations for technical and non-technical audiences
Genuine curiosity about why a model behaves the way it does, and whether we’re measuring the right things
Desirable
Experience in autonomous driving, robotics, or safety-critical ML systems
Familiarity with simulation-based testing, on-road experiment analysis, or statistical methods for comparing model behaviour
Experience working across product, safety, and engineering teams to define what “good” looks like
Exposure to OEM or partner validation workflows
Experience improving evaluation infrastructure or measurement methodology
This is a full-time role based in our office in either in London/UK or Sunnyvale/US. At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home.
A quick, honest note before you apply.
Wayve is not a mature, fully-structured place with the playbook already written. Much of how we work is still being written, and if you join, you’ll help write it. That suits people who want real ownership more than people who need a settled structure from day one.
If that sounds like the kind of problem you want to spend your time on, we’d really like to hear from you.