ML Data Infrastructure Engineer
Job Description
<div class="content-intro"><h3><span style="font-weight: 400;"><strong>About AppLovin</strong></span></h3> <p><a href="https://cts.businesswire.com/ct/CT?id=smartlink&url=http%3A%2F%2Fwww.applovin.com&esheet=54549558&newsitemid=20260608414458&lan=en-US&anchor=AppLovin&index=3&md5=a4ac8b1dcb8b5502d0154df68e322df6" target="_blank">AppLovin</a> makes technologies that help businesses of every size connect to their ideal customers. The company provides end-to-end advertising solutions for businesses to reach, monetize and grow their global audiences. For more information about AppLovin, visit: <a href="http://www.applovin.com" target="_blank">www.applovin.com</a>.</p> <p><span style="font-weight: 400;">To deliver on this mission, our global team is composed of team members with life experiences, backgrounds, and perspectives that mirror our developers and customers around the world. At AppLovin, we are intentional about the team and culture we are building, seeking candidates who are outstanding in their own right and also demonstrate their support of others.</span></p></div><p>As a member of our ML Data Platform team, you'll solve technical challenges, including upgrading and implementing state-of-the-art software infrastructure. The team builds a high-performance, high availability, globally distributed ecosystem platform of services that in turn provide the foundation for rapid development of novel new systems that integrate into that ecosystem and improve it.</p> <h3><strong>The Impact You'll Make</strong></h3> <ul> <li style="color: #1d1c1d !important;"><span style="color: #1d1c1d;">Design and build data processing infrastructure for model training and feature serving, optimizing for performance, reproducibility, and traceability</span></li> <li style="color: #1d1c1d !important;"><span style="color: #1d1c1d;">Collaborate closely with research teams to design and implement novel data processing architectures for emerging model and training paradigms</span></li> <li style="color: #1d1c1d !important;"><span style="color: #1d1c1d;">Identify and resolve performance bottlenecks across the training data pipeline, from raw data ingestion to feature delivery</span></li> <li style="color: #1d1c1d !important;"><span style="color: #1d1c1d;">Establish best practices, tooling for data infrastructure used across ML teams</span></li> </ul> <h3><span style="color: #1d1c1d;"><strong>Required Qualifications</strong></span></h3> <ul> <li>Have 1 - 3 years of experience and a minimum of a BS and/or MS in Computer Science</li> <li style="color: #1d1c1d !important;"><span style="color: #1d1c1d;">Strong software engineering fundamentals, with experience building high-throughput, fault-tolerant distributed systems</span></li> <li style="color: #1d1c1d !important;"><span style="color: #1d1c1d;">Hands-on experience with distributed computing frameworks such as Apache Spark or Flink</span></li> <li style="color: #1d1c1d !important;"><span style="color: #1d1c1d;">Solid grounding in data structures, systems design, and performance optimization</span></li> <li style="color: #1d1c1d !important;"><span style="color: #1d1c1d;">Strong problem-solving skills and attention to detail</span></li> </ul> <h3><span style="color: #1d1c1d;"><strong>Preferred Qualifications</strong></span></h3> <ul> <li style="color: #1d1c1d !important;"><span style="color: #1d1c1d;">Background in MLOps, Data Infrastructure, or ML Infrastructure</span></li> <li style="color: #1d1c1d !importan