Incident Operations Lead (EMEA/AMER)
Job Description
<div class="content-intro"><p><strong>Who We Are:</strong></p> <p><strong>Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure </strong>for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more.<br><br>Amongst our subsidiaries, Alpaca is a licensed financial services company, serving hundreds of financial institutions across 40 countries with our institutional-grade APIs. This includes broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, totalling over 10 million brokerage accounts.<br><br>Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to achieve our mission of opening financial services to everyone on the planet. We're deeply committed to open-source contributions and fostering a vibrant community, continuously enhancing our award-winning, developer-friendly API and the robust infrastructure behind it.<br><br><strong>Alpaca is proudly backed by $400 million in funding from top-tier global investors including Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, Horizons Ventures, Opera Tech Ventures, SBI Group, Derayah Financial, Unbound, Peak XV, Elefund, and Y Combinator.</strong><br><br><strong>Our Team Members:</strong></p> <p>We're a dynamic team of 400+ globally distributed members who thrive working from our favorite places around the world, with teammates spanning the USA, Canada, Japan, Hungary, Nigeria, Brazil, the UK, and beyond!<br><br>We're searching for passionate individuals eager to contribute to Alpaca's rapid growth. If you align with our core values—Stay Curious, Have Empathy, and Be Accountable—and are ready to make a significant impact, we encourage you to apply.</p></div><h3>Role</h3> <p>Lead the team that commands Alpaca's most critical incidents. You will build the function and then keep raising its bar: the severity model, the escalation and communication paths, 24x7 follow-the-sun coverage, and the KPIs that prove it is improving. You will do that across boundaries - with the engineering teams who own the services, with SRE on reliability standards and on-call readiness, with Risk on financial and regulatory materiality, with our partner communications teams on what reaches a customer, and with reliability programme management on what happens after.</p> <p>You will own how well we respond. Not the fix, not the partner communication, and not the reliability standard. Holding that line is a deliberate part of the design and a core part of the job.</p> <p><strong>Things You Get To Do</strong></p> <ul> <li><strong>Build the team and stand up 24x7 command.</strong> Recruit and certify Incident Commanders, build a follow-the-sun rotation across APAC, EMEA and AMER with warm handoffs at every regional boundary, and carry a rostered slot yourself. Keep the team sharp between real incidents with game days, tabletop exercises and simulations, and coach them through the live ones. Build a blameless review culture that treats an outlier as a process gap rather than a person's failure.</li> <li><strong>Own the process, and keep raising it.</strong> Drive severity maturity with Risk on financial and regulatory materiality - in a regulated brokerage a severity call can also start a reporting clock, so the model has to map cleanly onto those thresholds. Own the escalation path and what happens when a page goes unanswered, agree the thresholds for taking an incident to engineering leadership, and keep the service catalogue and its ownership current - time spent working out who owns a failing service is customer impact.</li> <li><strong>Own both bridges.</strong> Your team c