CF CloudFrame Job Scanner Open Dashboard →
Verified Active Opening

Senior Site Reliability Engineer

Alpaca • Remote - Global

Job Description

<div class="content-intro"><p><strong>Who We Are:</strong></p> <p><strong>Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure </strong>for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more.<br><br>Amongst our subsidiaries, Alpaca is a licensed financial services company, serving hundreds of financial institutions across 40 countries with our institutional-grade APIs. This includes broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, totalling over 10 million brokerage accounts.<br><br>Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to achieve our mission of opening financial services to everyone on the planet. We're deeply committed to open-source contributions and fostering a vibrant community, continuously enhancing our award-winning, developer-friendly API and the robust infrastructure behind it.<br><br><strong>Alpaca is proudly backed by $400 million in funding from top-tier global investors including Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, Horizons Ventures, Opera Tech Ventures, SBI Group, Derayah Financial, Unbound, Peak XV, Elefund, and Y Combinator.</strong><br><br><strong>Our Team Members:</strong></p> <p>We're a dynamic team of 400+ globally distributed members who thrive working from our favorite places around the world, with teammates spanning the USA, Canada, Japan, Hungary, Nigeria, Brazil, the UK, and beyond!<br><br>We're searching for passionate individuals eager to contribute to Alpaca's rapid growth. If you align with our core values—Stay Curious, Have Empathy, and Be Accountable—and are ready to make a significant impact, we encourage you to apply.</p></div><h2><span style="font-size: 10pt;">Your Role:</span></h2> <p>As a Site Reliability Engineer at Alpaca, you'll help keep our brokerage platform reliable, observable, and operable as we grow - working across our cloud infrastructure, Kubernetes platform, observability stack, messaging layer, and data layer. We're especially interested in candidates with strong PostgreSQL fundamentals who'd like to grow into deeper ownership of our database reliability posture: PostgreSQL sits on the trading-critical path, and we want this person to spend a meaningful share of their time leveling it up while still being a well-rounded SRE the rest of the week.</p> <h2><strong>Things You Get To Do</strong></h2> <ul> <li><strong>Operate production day-to-day</strong> - oncall, incident response, postmortems, and the follow-ups that actually close the loop.</li> <li><strong>Own reliability practice</strong> - define and refine SLIs/SLOs and error budgets, and help product teams live within them.</li> <li><strong>Strengthen our observability</strong> across metrics, logs, traces, and alerting.</li> <li><strong>Ship infrastructure through code</strong> in a GitOps workflow - cloud resources and Kubernetes workloads alike.</li> <li><strong>Look after PostgreSQL</strong>: performance tuning, schema and migration review, online migrations on large tables, HA/DR, and CDC pipelines.</li> <li><strong>Mentor engineers</strong> on reliability and database fundamentals through code review, design review, and pairing.</li> </ul> <h2><strong>Who You Are (must-haves)</strong></h2> <ul> <li><strong>4+ years</strong> in SRE, DevOps, Platform/Infrastructure, or backend engineering with significant production operations ownership.</li> <li>Hands-on experience operating production services on &

Job Reference ID: CF-150309 • Posted on CloudFrame Job Scanner