Senior Site Reliability Engineer, IaaS
Job Description
<div class="content-intro"><div><span id="m_8220454926977230902gmail-docs-internal-guid-9a9f34f4-7fff-014a-c97c-98ca90610a47">Algolia is the retrieval intelligence layer that turns intent into trusted, decision-grade outcomes. Powering more than 1.7 trillion queries a year for over 18,000 customers with millisecond latency and 99.999% reliability, we are the recognized leader for Search and Product Discovery by top industry analyst firms. The Algolia platform turns a company's products, content, and business rules into data that humans, applications and AI agents can act upon. The result is trusted customer experiences with stronger conversions for measurable business impact.</span></div> <div> </div></div><h3><strong>The team</strong></h3> <p>The Infrastructure as a Service team is at the center of one of Algolia’s most consequential engineering transformations.</p> <p>For years, Algolia has operated a production fleet of approximately 4,000 bare-metal servers to deliver the reliability, low latency, and scalability that our customers expect. We are now building the foundations of a unified cloud and Kubernetes platform designed to support Algolia’s growth for years to come.</p> <p>This is not a lift-and-shift project. It is an opportunity to rethink how Algolia provisions, secures, operates, observes, upgrades, and scales production infrastructure and to build it as a platform that engineers can safely consume, rather than a queue of manual requests.</p> <h3><strong>The opportunity</strong></h3> <p>As a Senior Site Reliability Engineer in IaaS, you will help shape the next generation of Algolia’s production infrastructure.</p> <p>You will lead major parts of the Cloud Baseline and the reliable lifecycle capabilities that enable teams to operate and migrate workloads safely on a cloud-native platform. You will work across cloud foundations, Kubernetes, platform engineering, automation, reliability, and large-scale production operations.</p> <p>This role is for an engineer who enjoys solving infrastructure problems where the answer must work not once, but hundreds or thousands of times: creating repeatable cloud environments, enabling a growing fleet of production clusters, reducing manual operations, and maintaining the reliability and cost efficiency our customers expect throughout the transition.</p> <h3><strong>YOU WILL:</strong></h3> <ul> <li>Lead the design and evolution of Cloud Baseline capabilities across cloud providers, including identity and access, networking, account structure, security, auditability, tagging, inventory, and cost visibility.</li> <li>Design and automate cloud infrastructure foundations that enable a growing fleet of production Kubernetes clusters.</li> <li>Lead complex infrastructure initiatives, such as cloud-environment standardisation, cluster lifecycle automation, upgrade strategies, or infrastructure-drift reduction.</li> <li>Ensure cloud and Kubernetes foundations, lifecycle operations, and operational guardrails are reliable and scalable enough to support large-scale workload migration without compromising customer experience.</li> <li>Treat the platform as a product: define clear interfaces, reusable modules, self-service workflows, documentation, and reliable operational standards for the engineers who consume it.</li> <li>Build automated guardrails for security, compliance, reliability, and safe change management, allowing teams to move faster without weakening production protections.</li> <li>Improve platform efficiency through capacity planning, rightsizing, autoscaling, resource governance, and clear cost visibility.</li> <li>Use automation and AI-assisted engineering tools where appropriate to i