CF CloudFrame Job Scanner Open Dashboard →
Verified Active Opening

Senior Engineer - Platform

Buildkite • ANZ Region

Job Description

<h2><strong>About Buildkite</strong></h2> <p>At Buildkite, our mission is to unblock every developer on the planet. We’ve rethought how software delivery should work and built a platform that is fast, reliable, secure, and able to scale to the needs of demanding high-growth technology companies globally, including Airbnb, Shopify, Canva, PagerDuty, Lyft, and Pinterest.</p> <h2><strong>Job Overview</strong></h2> <p>We’re hiring a Senior Platform Engineer to join our Platform engineering organisation. The Platform team builds and operates the shared foundations that help Buildkite teams ship and scale products securely and reliably.</p> <p>This is a hands-on role for a senior platform engineer with strong cloud and infrastructure foundations. You’ll design and operate reliable Kubernetes platforms across multiple clusters, solve distributed-systems challenges, automate infrastructure at scale, and take ownership of the systems you build. </p> <p>We’re also hiring a Platform Engineer at intermediate level. If you’re earlier in your career than this role describes but the work below appeals, apply anyway and we’ll consider you.</p> <h2><strong>🔧 About the Team</strong></h2> <p>Our mission is to build the foundations that let Buildkite teams ship, operate, and scale products securely and reliably. We make the reliable path the easiest path through paved roads for deployment, observability and security.</p> <p>Buildkite’s customers run some of the world’s most demanding CI/CD and emerging AI workloads. As software development accelerates, CI/CD is becoming a major scaling bottleneck. A shared-platform change can affect every product team, while one unusual workload can expose the platform’s next limit. The infrastructure underneath must behave accordingly.</p> <h3><strong>Our current challenges include:</strong></h3> <ul> <li>Self-service, repeatable foundations. Build secure defaults, reusable architecture, automation, tooling, and documentation so teams can provision and operate consistent environments across regions and isolated use cases without tickets or operational hand-offs.</li> <li>Kubernetes at production scale. Improve networking, upgrades, patching, capacity management, and recovery without disrupting teams or customers.</li> <li>Global and regional resilience. Develop repeatable regional environments, customer routing, disaster recovery, and tested failover. Design for traffic growth, partial failures, hot shards, noisy neighbours, queue pressure, and datastore limits.</li> <li>Safe, observable operations. Make infrastructure and application changes easier to deploy, understand, and reverse. Improve observability, capacity planning, incident response, documentation, and runbooks, and demonstrate readiness through tests, dashboards, and recovery exercises.</li> </ul> <p>We measure our impact through cluster availability, customer-routing latency, deployment speed and success, self-service adoption, and recovery time.</p> <h2><strong>🚀 What You’ll Do</strong></h2> <ul> <li>Own substantial parts of Buildkite’s AWS and multi-cluster Kubernetes platform from design through production operation.</li> <li>Design reusable, self-service automation for provisioning and operating secure, consistent, and isolated platform environments.</li> <li>Deliver regional architecture, request routing, disaster recovery, and tested failover capabilities.</li> <li>Improve Kubernetes networking, upgrades, patching, scaling, observability, capacity management, and recovery.</li> <li>Make infrastructure and deployments safer through secure defaults, progressive delivery, automated guardrails, clear health signals, and fast

Job Reference ID: CF-167272 • Posted on CloudFrame Job Scanner