Platform Lead - ML Ops
Job Description
<div class="content-intro"><p>At Anaplan, we are a team of innovators focused on optimizing business decision-making through our leading AI-infused scenario planning and analysis platform so our customers can outpace their competition and the market.</p> <p>What unites Anaplanners across teams and geographies is our collective commitment to our customers’ success and to our Winning Culture.</p> <p style="padding-left: 40px;">Our customers rank among the who’s who in the Fortune 50. Coca-Cola, LinkedIn, Adobe, LVMH and Bayer are just a few of the 2,400+ global companies who rely on our best-in-class platform.</p> <p style="padding-left: 40px;">Our Winning Culture is the engine that drives our teams of innovators. We champion diversity of thought and ideas, we behave like leaders regardless of title, we are committed to achieving ambitious goals, and we love celebrating<em> </em>our wins – big and small.</p> <p>Supported by operating principles of being strategy-led, <a href="https://www.anaplan.com/careers/">values</a>-based and disciplined in execution, you’ll be inspired, connected, developed and rewarded here. Everything that makes you unique is welcome; join us and let’s build what’s next - together!</p></div><p><strong>Role Overview</strong></p> <p>We are seeking a Platform Lead to spearhead our ML Ops infrastructure, cost-optimisation, and deployment strategies. You will manage a talented team of DevOps engineers while remaining deeply technical and hands-on. Your primary mission is to build, scale, and secure the foundational platforms for our machine learning (ML) and generative AI (GenAI) models while maintaining financial accountability.<br><br><strong>Your Impact</strong></p> <ul> <li><strong><span data-markdown-start-index="157">Team Leadership & Collaboration:</span></strong><span data-markdown-start-index="191"> Lead and manage a dedicated DevOps team, mentoring both junior and senior engineers while collaborating closely with Data Science and Engineering leaders.</span></li> <li><strong><span data-markdown-start-index="352">Infrastructure Strategy & Automation:</span></strong><span data-markdown-start-index="391"> Define the infrastructure roadmap for AI/ML workloads and automate provisioning across cloud environments using Infrastructure as Code (IaC).</span></li> <li><strong><span data-markdown-start-index="539">MLOps & LLMOps Engineering:</span></strong><span data-markdown-start-index="568"> Architect, maintain, and optimise robust MLOps/LLMOps pipelines and CI/CD frameworks for continuous model deployment.</span></li> <li><strong><span data-markdown-start-index="692">GenAI Production Deployment:</span></strong><span data-markdown-start-index="722"> Deploy Large Language Models (LLMs) into production environments, ensuring high availability, low latency, and optimal performance for GenAI applications.</span></li> <li><strong><span data-markdown-start-index="883">FinOps & Budget Management:</span></strong><span data-markdown-start-index="912"> Establish FinOps frameworks to track, allocate, and forecast AI infrastructure spend, managing high-cost GPU/CPU cloud budgets.</span></li> <li><strong><span data-markdown-start-index="1046">Resource Efficiency & Unit Economics:</span></strong><span data-markdown-start-index="1085"> Implement auto-scaling, spot instances, and down-scaling policies to eliminate waste, while providing full visibility into the unit economics of training and serving LLM models.</span></li> <li><str