Senior Site Reliability Engineer
Job Description
Senior Site Reliability Engineer Location: Global Remote / San Francisco · Full-Time About Andromeda Andromeda is a market and infrastructure platform to buy, sell, and operate compute. We believe demand for compute will grow exponentially. So fast that a handful of vertically integrated providers won't be able to scale across operations, capital, supply chains, and politics to serve it. The result is a massive wave of fragmentation, with AI factories of every shape and size coming to market to fill this demand. Our job is to enable all of that fragmented compute to flow through one platform, delivering reliable capacity to model builders, research labs, and inference providers when they need it. We believe every spare electron should be made productive for AI and we're building the platform that makes that possible. We sit at the center of three forces: Companies that need reliable, high-performance compute fast A fragmented global supply of GPUs across hyperscalers, neoclouds, and independent data centers Capital, risk, and operational complexity that most teams are not equipped to manage When we succeed, trillions of dollars of compute will flow through Andromeda. Builders get capacity when they need it. Providers get a reliable way to monetize, operate, and finance infrastructure at scale. Capital gets an easy way to deploy, hedge, and underwrite. In five years, Andromeda won't just participate in the AI infrastructure market. We will shape it. The Role This is not a generalist SRE role. You will design, operate, and debug large-scale GPU infrastructure used for distributed training and inference, working directly with customers pushing the limits of modern AI systems. We’re looking for engineers who have personally run GPU clusters in production, understand the failure modes of distributed training, and can reason about performance from network fabric → kernel → framework. What You’ll Own GPU Cluster Architecture: Design and evolve multi-provider, multi-region GPU compute clusters optimized for large-scale training. Make topology-aware scheduling, networking, and storage decisions that directly impact training throughput and cost efficiency. Customer Technical Partnership: Serve as the primary technical point of contact for customers running large-scale training workloads. Onboard, troubleshoot, and optimize, often in real time. Reliability & Performance Engineering: Define SLOs and error budgets that account for the unique failure modes of GPU infrastructure (ECC errors, NVLink degradation, NCCL timeouts). Own capacity planning across heterogeneous GPU fleets optimized for training throughput. Networking & Fabric Health: Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) that underpin distributed training. Diagnose and resolve fabric-level issues that degrade collective operations. Observability: Build deep visibility into GPU utilization, memory pressure, interconnect throughput, training job performance, and hardware health. Go well beyond standard infrastructure metrics. Automation & Tooling: Build production-grade automation for cluster provisioning, GPU health checks, job scheduling, self-healing, and firmware/driver lifecycle management. Incident Leadership: Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks. Drive blameless postmortems and systemic fixes. What We’re Looking For GPU Systems Expertise: Deep, hands-on experience operating large-scale GPU clusters (NVIDIA A100/H100/B200 or equivalent). You understand GPU memory hierarchies, ECC behavior, thermal throttling, and hardware failure modes from direct experience not documentation. High-Performance Networking: Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training. You can diagnose why an all-reduce is slow, identify a degraded link in a fat-tree topology, and reason about congestion control at scale. Dist