CF CloudFrame Job Scanner Open Dashboard →
Verified Active Opening

Sr Staff Site Reliability Engineer

Archer56 • San Jose, California, United States

Job Description

<div class="content-intro"><p>Headquartered in Silicon Valley, California, Archer is a leader in the next-gen aerospace sector building an end-to-end advanced air mobility platform that delivers air taxis, unmanned aircraft systems (“UAS”), aviation-related physical artificial intelligence (“AI”) solutions, and other technologies to customers worldwide across the commercial aerospace and defense sectors.</p> <p><span style="font-weight: 400;">Our sights are set high and our problems are hard, and we believe that diversity in the workplace is what makes us smarter, drives better insights, and will ultimately lift us all to success. We are dedicated to cultivating an equitable and inclusive environment that embraces our differences, and supports and celebrates all of our team members.</span></p></div><p>We are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role, you will be responsible for the reliability, scalability, performance, and security of our core systems and services. You will leverage your extensive expertise in various technologies to design, implement, and maintain robust infrastructure and automation solutions.</p> <h1><span style="font-size: 14pt;">Responsibilities</span></h1> <ul> <li>Implement and maintain the infrastructure and pipeline required for an internal LLM-powered chat service, potentially leveraging platforms like OpenRouter or similar alternatives.</li> <li>implement and maintain highly available, scalable, and secure cloud-native infrastructure on Amazon Elastic Kubernetes Service (EKS).</li> <li>Develop and implement comprehensive observability strategies, including monitoring, logging, and alerting, to ensure the health and performance of our systems.</li> <li>Architect and optimize data pipelines to ensure efficient and reliable data flow across various platforms.</li> <li>Drive the continuous improvement of our CI/CD pipelines, promoting best practices for automated testing, deployment, and release management.</li> <li>Champion cloud-first strategies, leveraging the full capabilities of cloud platforms for infrastructure, services, and operations.</li> <li>Implement and enforce robust security practices across our infrastructure, applications, and data.</li> <li>Design and maintain Docker-based containerization solutions for our applications.</li> <li>Develop and maintain automation scripts and tools using Python, Bash, and PowerShell.</li> <li>Collaborate with development teams to ensure reliability is built into the software development lifecycle from inception.</li> <li>Troubleshoot complex production issues across various layers of the stack, identifying root causes and implementing preventative measures.</li> <li>Participate in on-call rotations to support production systems.</li> </ul> <h1><span style="font-size: 14pt;">Qualifications</span></h1> <ul> <li>12+ years of experience in Site Reliability Engineering, DevOps, or a similar role with a strong focus on operational excellence.</li> <li>Deep expertise in Amazon EKS, including cluster provisioning, management, and troubleshooting.</li> <li>Extensive experience with observability tools and practices, including Prometheus, Grafana, ELK stack, or similar.</li> <li>Proven track record in designing and implementing robust data pipelines (e.g., Kafka, Airflow, Spark).</li> <li>Strong background in CI/CD methodologies and tools (e.g., Jenkins, GitLab CI, ArgoCD).</li> <li>Expert-level knowledge of cloud platforms (AWS preferred), including infrastructure-as-code principles.</li> <li>Comprehensive understanding of securit

Job Reference ID: CF-154657 • Posted on CloudFrame Job Scanner