Senior Reliability Engineer
Job Description
<p><strong><span data-contrast="auto">About BEES</span></strong><span data-ccp-props="{"201341983":0,"335559739":0,"335559740":240}"> </span></p> <p><span data-contrast="auto">Join us to build the future of B2B commerce!</span><span data-ccp-props="{"201341983":0,"335559739":0,"335559740":240}"> </span></p> <p><span data-contrast="auto">BEES is AB InBev’s B2B platform. Through our ecosystem, merchants and retailers across 29 countries can stock their businesses quickly, easily, and securely.</span> <br><span data-contrast="auto">At BEES, we dream big, lead with purpose, and develop technology that transforms the way retailers and sellers grow.</span><span data-ccp-props="{"201341983":0,"335559739":0,"335559740":240}"> </span></p> <p><span data-contrast="auto">Every line of code and every partnership is built in service of a single mission: to make commerce better for retailers and sellers around the world. Here, your work is not just important. It makes a difference!</span><span data-ccp-props="{"201341983":0,"335559739":0,"335559740":240}"> </span></p> <p><strong>What you will do:</strong></p> <ul> <li>Architect and design highly available, fault-tolerant, and scalable infrastructure and systems.</li> <li>Collaborate with software development teams to influence design decisions and ensure reliability and scalability from the ground up.</li> <li>Develop and implement best practices, guidelines, and standards for SRE processes, infrastructure, and automation.</li> <li>Define and enforce service-level objectives (SLOs) and error budgets to drive system reliability and availability.</li> <li>Evaluate and recommend suitable technologies, tools, and platforms for infrastructure, monitoring, and observability.</li> <li>Drive automation efforts through the development and maintenance of infrastructure-as-code (IaC) solutions.</li> <li>Lead incident response and post-incident analysis efforts, identifying root causes and implementing preventive measures.</li> <li>Implement robust monitoring, logging, and alerting solutions to proactively detect and resolve issues.</li> <li>Collaborate with cross-functional teams, including development, operations, and quality assurance, to ensure alignment and successful project delivery.</li> <li>Stay updated with emerging technologies, industry trends, and best practices in SRE and cloud infrastructure.</li> </ul> <p><strong><span class="TextRun SCXW225092077 BCX0" lang="EN-US" data-contrast="auto"><span class="NormalTextRun SCXW225092077 BCX0">We are looking for people with:</span></span></strong></p> <ul> <li>Bachelor's or Master's degree in computer science, engineering, or a related field.</li> <li>3 years of experience in SRE or systems engineering/architecture roles, with a focus on designing and architecting highly reliable and scalable systems.</li> <li>Solid understanding of cloud platforms (e.g., AWS, Azure, GCP) and experience with infrastructure-as-code (IaC) tools like Terraform.</li> <li>Proficiency in at least one programming language (e.g., Python, Go, Java) and experience with scripting for automation.</li> <li>Strong knowledge of containerization and orchestration technologies (e.g., Docker, Kubernetes).</li> <li>Experience with monitoring and observability tools (e.g., New Relic, Dynatrace, Prometheus, Grafana, ELK stack).</li> <li>Proven track record of architecting and implementing highly available and scalable systems in a production environment.</li> <li>Strong problem-sol