CF CloudFrame Job Scanner Open Dashboard →
Verified Active Opening

Lead Site Reliability Engineer

Alloy • New York City

Job Description

<h2>Alloy is where you belong!</h2> <p>Alloy is the AI-powered identity and fraud prevention platform that accelerates onboarding, stops fraud, and scales compliance across the customer lifecycle so financial organizations can grow without limits. More than 900 of the world's leading financial institutions and fintechs trust Alloy for smarter risk management that drives growth.<br><br>Through our values: Be Bold, Go Fast, Collaborate, and Celebrate Our Differences, we are creating a workplace where you can grow, thrive, and belong. See how we’ve been continuously recognized and named one of <a href="https://www.inc.com/profile/alloy_ny" target="_blank">Inc. Magazine’s Best Workplaces</a>, <a href="https://www.forbes.com/companies/alloy/?list=americas-best-startup-employers&sh=55c076453c61" target="_blank">Forbes America’s Best Startup Employers</a>, <a href="https://www.americanbanker.com/list/best-places-to-work-in-fintech-2022" target="_blank">Best Fintech to Work for by American Banker</a>, year after year.</p> <p><span style="font-weight: 400;">Check out our investors and read more about us <a class="c-link" href="https://www.alloy.com/about" target="_blank" data-stringify-link="https://www.alloy.com/about" data-sk="tooltip_parent">here</a>.</span></p> <h2>About the team</h2> <p>Alloy’s Infrastructure Team is a small team (6 engineers) responsible for a large and growing infrastructure footprint: 15+ Kubernetes clusters, 100+ databases, dozens of services, and complex data organization.</p> <p>Our challenge isn’t just scale—it’s making that scale reliable, secure, and operable with less manual work.</p> <p>We’re looking for engineers who enjoy turning complex, fragile systems into automated, self-service platforms with strong safety guarantees.</p> <h2><strong>What you'll be doing</strong></h2> <p>Reporting to the Engineering Manager of Infrastructure, you'll:</p> <ul> <li>Design and build systems to automate infrastructure management at scale (provisioning, upgrades, migrations)</li> <li>Reduce operational toil by turning manual processes into reliable, repeatable workflows</li> <li>Build internal tooling and platforms that enable safe self-service changes for other engineers</li> <li>Improve the reliability and resilience of our infrastructure (Kubernetes, databases, services)</li> <li>Implement and evolve systems for deploying and running applications in Kubernetes</li> <li>Contribute to architecture decisions across infrastructure, reliability, and security</li> <li>Write and review production-quality code</li> <li>Participate in on-call rotations—but focus on building systems that prevent incidents, not just respond to them</li> </ul> <h2><strong>Who we’re looking for</strong></h2> <ul> <li>10+ years of experience in infrastructure, SRE, or software engineering roles</li> <li>Strong software engineering skills—you build systems, not just scripts</li> <li>Experience managing production infrastructure at scale (cloud + containerized systems)</li> <li>Experience with Infrastructure as Code (e.g., Terraform)</li> <li>Experience running and troubleshooting distributed systems (Docker/Kubernetes)</li> <li>Experience with observability and debugging tools (Datadog, CloudWatch, ELK/EFK, etc.)</li> <li>Proficiency in at least one programming language (Python, Go, JavaScript, etc.)</li> <li>Experience participating in on-call rotations and improving systems based on incidents</li> <li>Strong communication and collaboration skills</li> </ul> <p><strong>You might be a great

Job Reference ID: CF-148233 • Posted on CloudFrame Job Scanner