Senior Staff Engineer, SRE
Job Description
<p>We’re a global team of over 400 people, working together to push the boundaries of open-source technology and multi-cloud solutions. Our vision is to help developers, builders, and creators bring their ideas to life with speed and simplicity, by providing a cloud data platform that makes open-source databases, search, streaming, and application infrastructure easily accessible to everyone. </p> <h3><strong>The Role:</strong></h3> <p>We are seeking a <strong>Senior Staff Engineer, Site Reliability Engineering</strong> to help drive the reliability and operational excellence of the Aiven platform globally. You will work closely with our SRE teams, helping set the technical direction and strategy to ensure resilient, scalable, and highly automated systems across our 24/7/365 operations.</p> <p>You will work hands-on with SRE and engineering teams on platform health, incident response and cross-functional coordination, and drive continuous improvement in reliability and performance. As a senior technical leader, you will partner closely with engineering, product, and support teams worldwide, influence system architecture, and invest in tooling and automation to reduce toil and enhance production reliability.</p> <p>This role combines technical leadership, customer focus, and deep operational expertise, with a focus on delivering reliable services at global scale and helping teams solve some of our most challenging reliability problems.</p> <h3><strong>What You'll Do:</strong></h3> <ul> <li>Help define and drive global SRE operating strategy in partnership with regional SRE leaders across EMEA, AMER and APAC, ensuring alignment on reliability goals, operating models, and execution across a 24/7/365 follow-the-sun organization.</li> <li>Provide hands-on technical leadership across our multi-regional SRE teams, working closely with SRE teams and engineers across geographies.</li> <li>Help set the vision and roadmap for reliability engineering, enabling teams to deliver high-impact tools, automation, and process initiatives that improve platform resilience, scalability, and efficiency.</li> <li>Drive improvements to global incident management strategy and operating model, including on-call design, coverage, and escalation frameworks, ensuring seamless coordination and high availability across regions.</li> <li>Help establish a metrics-driven operating cadence, defining KPIs/SLIs/SLOs/Error Budgets, driving data-informed prioritization, and embedding operational rigor and continuous improvement across the SRE organization.</li> </ul> <h3><strong>What We're Looking For:</strong></h3> <ul> <li>Proven experience working in senior SRE or infrastructure engineering roles, ideally across multiple regions and time zones.</li> <li>Strong track record of defining and executing reliability strategy at scale, including ownership of SLIs/SLOs, incident management frameworks, and operational excellence programs.</li> <li>Demonstrated ability to provide technical leadership, mentor engineers, and influence teams without relying on organizational authority.</li> <li>Experience operating in a 24/7/365 production environment, with deep understanding of follow-the-sun models, on-call design, and large-scale incident response.</li> <li>Ability to partner cross-functionally with senior Engineering, Product, and Support stakeholders to influence architecture, prioritization, and long-term platform investments.</li> <li>Strong data-driven approach, with experience defining SLIs/SLOs and using metrics to drive prioritization, accountability, and continuous improvement.</li> <li>Solid technical foundation in distributed systems, cloud infrastructure, and automation, wi