QA Automation Engineer - AI
Job Description
<div class="content-intro"><p>At Anaplan, we are a team of innovators focused on optimizing business decision-making through our leading AI-infused scenario planning and analysis platform so our customers can outpace their competition and the market.</p> <p>What unites Anaplanners across teams and geographies is our collective commitment to our customers’ success and to our Winning Culture.</p> <p style="padding-left: 40px;">Our customers rank among the who’s who in the Fortune 50. Coca-Cola, LinkedIn, Adobe, LVMH and Bayer are just a few of the 2,400+ global companies who rely on our best-in-class platform.</p> <p style="padding-left: 40px;">Our Winning Culture is the engine that drives our teams of innovators. We champion diversity of thought and ideas, we behave like leaders regardless of title, we are committed to achieving ambitious goals, and we love celebrating<em> </em>our wins – big and small.</p> <p>Supported by operating principles of being strategy-led, <a href="https://www.anaplan.com/careers/">values</a>-based and disciplined in execution, you’ll be inspired, connected, developed and rewarded here. Everything that makes you unique is welcome; join us and let’s build what’s next - together!</p></div><p><strong>About The Role</strong></p> <p>We're pioneering a new role focused exclusively on quality assurance for GenAI/Agentic systems. As our QA Automation Engineer within AI, you'll develop testing strategies, evaluation frameworks, and quality metrics specifically designed for LLM-powered applications. This role requires a unique blend of QA expertise, understanding of GenAI behaviour, and automation skills to ensure our AI features are reliable, accurate, and trustworthy.</p> <p><strong>Your Impact</strong></p> <ul> <li><span data-markdown-start-index="153">Develop and execute end-to-end automated test coverage</span><span data-markdown-start-index="209"> across evaluations, APIs, UIs, and performance tests using Python.</span></li> <li><span data-markdown-start-index="280">Implement testing strategies</span><span data-markdown-start-index="310"> for GenAI features, including conversational AI, agentic systems, and LLM-powered workflows within your workstream.</span></li> <li><span data-markdown-start-index="430">Run and maintain automated test suites</span><span data-markdown-start-index="470"> for prompt testing, including regression tests designed to detect unintended changes in model behaviour.</span></li> <li><span data-markdown-start-index="579">Support and maintain evaluation frameworks</span><span data-markdown-start-index="623"> to measure GenAI performance across accuracy, relevance, safety, consistency, and latency.</span></li> <li><span data-markdown-start-index="718">Contribute to curating high-quality test datasets</span><span data-markdown-start-index="769"> and "golden examples" that represent diverse user scenarios and edge cases.</span></li> <li><span data-markdown-start-index="849">Assist in adversarial testing</span><span data-markdown-start-index="880"> to identify potential model hallucinations or biases, and support the integration of production alerting systems.</span></li> <li><span data-markdown-start-index="998">Utilise internal tools and testing frameworks</span><span data-markdown-start-index="1045"> to help developers easily validate their GenAI code, and assist with user acceptance testing.</span></li> <li><span data-markdown-start-index="1143">Take accountability for your workstream's automated coverage</span><span data-markdown-start-index="1205">, do