Staff Data Engineer
Job Description
<div class="content-intro"><p>Headquartered in Silicon Valley, California, Archer is a leader in the next-gen aerospace sector building an end-to-end advanced air mobility platform that delivers air taxis, unmanned aircraft systems (“UAS”), aviation-related physical artificial intelligence (“AI”) solutions, and other technologies to customers worldwide across the commercial aerospace and defense sectors.</p> <p><span style="font-weight: 400;">Our sights are set high and our problems are hard, and we believe that diversity in the workplace is what makes us smarter, drives better insights, and will ultimately lift us all to success. We are dedicated to cultivating an equitable and inclusive environment that embraces our differences, and supports and celebrates all of our team members.</span></p></div><p><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;">Archer is best known for Midnight, our electric aircraft. This role isn't on that program. The AI Products Org is building software for the general aviation industry. As a Staff Data Engineer on the AI Platform team, you will design, build, and operate the data infrastructure that powers large-scale model training and inference. You will own the pipelines, storage systems, and data quality mechanisms that sit upstream of our ML platform, ensuring models train on clean, high-throughput, well-governed data.</span></p> <p style="line-height: 1.5;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><strong><br>What You’ll Do:</strong></span></p> <ul class="p-rich_text_list p-rich_text_list__bullet p-rich_text_list--nested" data-stringify-type="unordered-list" data-list-tree="true" data-indent="0" data-border="0"> <li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Design and maintain high-throughput, fault-tolerant ingestion and transformation pipelines that feed training workloads at scale, with a focus on latency, throughput, and correctness.</span></li> <li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Build and operate the data lakehouse — defining table formats (Iceberg, Paimon, Parquet), partitioning strategies, and compaction policies optimized for ML consumption patterns.</span></li> <li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Instrument pipelines with data quality checks, lineage tracking, and anomaly detection so that model failures trace back to data problems quickly.</span></li> <li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"> Partner with ML engineers to define feature stores, dataset versioning, and experiment-to-production data contracts; integrate with tools like MLflow for dataset and artifact tracking.</span></li> <li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Work closely with AI researchers, platform engineers, and software engineers to understand data access patterns, optimize query performance, and unblock training runs.</span></li> </ul> <div c