Lead Evaluation Engineer – Scenario Coverage & Datasets (Autonomous Driving)
Job Description
<h2>Lead Evaluation Engineer Scenario Coverage & Datasets (Autonomous Driving)</h2> <h3>About the team</h3> <p>Avride builds autonomous driving technology for vehicles and delivery robots. Our QA organization measures how well that technology drives, largely in simulation — and an evaluation is only as good as the data behind it. So we build our own datasets, keep track of what they cover, and keep widening that coverage as the technology takes on more of the road.</p> <h3>About the role</h3> <p>You will own <strong>test coverage for the autonomous driving stack, end to end.</strong> Scenario areas are owned by individual QA engineers. You own the whole: what our evaluation dataset must contain for a release decision to be trustworthy, where the gaps are that nobody working inside a single area can see, and how those gaps keep surfacing without you being the one who finds them.</p> <p>Coverage is a data problem before it is a testing problem. A scenario that exists on paper is not covered until there are real scenes behind it, and closing that distance is the craft of this role: metric-driven search across driving data, VLM-assisted retrieval and clustering, criticality-based selection, variation in simulation, and structured capture in the field. You will use that toolkit yourself before you ask anyone else to, and turn what works into something the whole QA team can run without you.</p> <p>You will lead through process first — building it, and delivering through the QA engineers and operational resources already around you — and grow a team of your own as the scope demands it.</p> <h3>The mindset we're hiring for</h3> <p>Most testing work starts from a defined expected output, where the job is to confirm the system produced it. This role is different. The possible driving situations are effectively infinite, and for most of them there is no single right answer to check against.</p> <p>We're looking for someone who thinks in terms of a space rather than a list: which parts of it we've sampled, which we haven't, what's rare but matters, and how you'd even notice that something is missing. If you've had to decide what to test when you couldn't test everything, and defend that decision, this role will feel familiar.</p> <h3>You don't need autonomous vehicle experience</h3> <p>The AV domain is learnable in a couple of quarters. The coverage instinct is what we can't teach. If you've done this kind of work under a different name, we want to hear from you:</p> <ul> <li><strong>Evaluation, validation, or simulation</strong> in AV, robotics, or perception</li> <li><strong>Detection engineering or threat detection</strong>, where the attack space is unbounded, there's no ground truth per event, and you're always trading false positives against false negatives</li> <li><strong>Search relevance or recommendations</strong>, building judged sets with coverage across query types</li> <li><strong>LLM or ML model evaluation</strong>, building eval sets and benchmarks and finding what they fail to catch</li> <li><strong>Speech recognition, medical AI, or fraud and risk modeling</strong>, curating test sets for the long tail, edge cases, and drift</li> </ul> <h3>What you'll do</h3> <ul> <li><strong>Own the evaluation dataset.</strong> What is in it, what is missing, and what it lets us claim about autonomous driving quality.</li> <li><strong>Own the coverage model.</strong> Our scenario taxonomy and ODD parameter space have to stay accurate as the technology matures and the operating environment changes. You own that: spotting where the model no longer describes what our vehicles actually meet, and