CF CloudFrame Job Scanner Open Dashboard →
Verified Active Opening

Site Reliability Engineer, IaaS

Algolia • Paris, France

Job Description

<div class="content-intro"><div><span id="m_8220454926977230902gmail-docs-internal-guid-9a9f34f4-7fff-014a-c97c-98ca90610a47">Algolia is the retrieval intelligence layer that turns intent into trusted, decision-grade outcomes. Powering more than 1.7 trillion queries a year for over 18,000 customers with millisecond latency and 99.999% reliability, we are the recognized leader for Search and Product Discovery by top industry analyst firms. The Algolia platform turns a company's products, content, and business rules into data that humans, applications and AI agents can act upon. The result is trusted customer experiences with stronger conversions for measurable business impact.</span></div> <div> </div></div><h3><strong>The team</strong></h3> <p>The Infrastructure as a Service team is at the center of one of Algolia’s most consequential engineering transformations.</p> <p>For years, Algolia has operated a production fleet of approximately 4,000 bare-metal servers to deliver the reliability, low latency, and scalability that our customers expect. We are now building the foundations of a unified cloud and Kubernetes platform designed to support Algolia’s growth for years to come.</p> <p>. We are now building the foundations of a unified cloud and Kubernetes platform designed to support Algolia’s growth for years to come.</p> <p>This is not a lift-and-shift project. It is an opportunity to rethink how Algolia provisions, secures, operates, observes, upgrades, and scales production infrastructure and to build it as a platform that engineers can safely consume, rather than a queue of manual requests.</p> <h3><strong>The opportunity</strong></h3> <p>As a Site Reliability Engineer in IaaS, you will help build the next generation of Algolia’s production infrastructure.</p> <p>You will contribute to the cloud baseline and reliable lifecycle capabilities that enable teams to operate and migrate workloads safely on a cloud-native platform. You will work across cloud foundations, Kubernetes, automation, reliability, and large-scale production operations.</p> <p>As a P3 engineer, you will be a hands-on contributor. You will build, operate, and improve production infrastructure while developing deep expertise in cloud, Kubernetes, reliability, and automation.</p> <h3><strong>YOU WILL:</strong></h3> <ul> <li>Build and improve Cloud Baseline capabilities, including identity and access, networking, security, resource inventory, tagging, and auditability.</li> <li>Develop and maintain infrastructure as code and automation for cloud environments and Kubernetes infrastructure.</li> <li>Contribute to reliable, repeatable cloud and cluster lifecycle operations.</li> <li>Help build self-service capabilities, reusable modules, and clear documentation that make the safe path the easy path for platform consumers.</li> <li>Reduce manual work and configuration drift through automation, testing, GitOps practices, and standardisation.</li> <li>Use automation and AI-assisted engineering tools where appropriate to improve infrastructure analysis, documentation, and safe, repeatable changes.</li> <li>Improve observability, monitoring, alerting, capacity management, and operational documentation.</li> <li>Investigate production issues, participate in the on-call rotation, and turn lessons learned into lasting improvements.</li> <li>Work with Infrastructure, Security, FinOps, and engineering teams to deliver reliable, secure, and cost-aware platform capabilities.</li> </ul> <h3><strong>YOU MIGHT BE A FIT IF YOU HAVE:</strong></h3> <ul> <li>Hands-on production knowledge of AWS or GCP.</li> <li>

Job Reference ID: CF-147728 • Posted on CloudFrame Job Scanner