Staff Engineer, Autonomous Driving Data Platform & Curation
Job Description
<p><em>We are </em><strong><em>CARIAD</em></strong><em>, an automotive software development team with the Volkswagen Group. Our mission is to make the automotive experience safer, more sustainable, more comfortable, more digital, and more fun. To achieve that we are building the leading tech stack for the automotive industry and creating a unified software platform for over 10 million new vehicles per year. We’re looking for talented, digital minds like you to help us create code that moves the world. Together with you, we’ll build outstanding digital experiences and products for all Volkswagen Group brands that will transform mobility. Join us as we shape the future of the car and everyone around it.</em></p> <p><em></em><strong><u>Role Summary:</u></strong></p> <p><span data-olk-copy-source="MessageBody">The Staff Engineer, Autonomous Driving Data Platform & Curation is a hands-on staff-level individual contributor who owns the path from raw multimodal vehicle data to reliable, versioned, and model-ready datasets. The ideal candidate combines production data or ML systems experience with practical knowledge of autonomous-driving data and can lead ingestion, schema, storage, query, curation, quality, and delivery. This engineer designs and operates datasets containing more than 10 million records or samples, evaluates formats such as Parquet and Lance, and enables efficient filtering, slicing, random access, and sequential retrieval. The role also advances data-quality monitoring, statistical and out-of-distribution detection, rule-based and model-based tagging, model-in-the-loop and human-in-the-loop labeling, and future data preparation for imitation learning and reinforcement learning. </span></p> <p><strong><u>Role Responsibilities:</u></strong></p> <p class="x_MsoNormal"><span data-olk-copy-source="MessageBody">Autonomous Data Architecture, Storage & Query </span></p> <ul type="disc"> <li class="x_MsoNormal">Design canonical representations for drives, scenarios, clips, frames, trajectories, sensor references, vehicle state, map context, labels, predictions, and dataset manifests. </li> </ul> <ul type="disc"> <li class="x_MsoNormal">Build and maintain validated, versioned datasets containing more than 10 million records or samples on local or cloud object storage. </li> </ul> <ul type="disc"> <li class="x_MsoNormal">Evaluate Parquet, Apache Arrow, Lance, and related technologies; own partitioning, indexing, file sizing, compaction, schema evolution, lineage, and reproducibility. </li> </ul> <ul type="disc"> <li class="x_MsoNormal">Optimize filtering, projection, joins, scenario slicing, random sampling, shuffling, sequential retrieval, and model data-loading performance. </li> <li class="x_MsoNormal">Measure and improve ingestion throughput, query latency, training throughput, storage utilization, reliability, and cost per usable sample. </li> </ul> <p> </p> <p class="x_MsoNormal">Autonomous Dataset Lifecycle, Curation & Training Readiness </p> <ul type="disc"> <li class="x_MsoNormal">Work effectively with synchronized camera and other sensor data, ego state, localization, calibration, coordinate frames, map context, control actions, clips, and trajectories. </li> </ul> <ul type="disc"> <li class="x_MsoNormal">Translate perception, planning, VLA, and evaluation needs into schemas, searchable attributes, scenario definitions, sampling strategies, and reproducible dataset splits. </li> </ul> <ul type="disc"> <li class="x_MsoNormal">