CF CloudFrame Job Scanner Open Dashboard →
Verified Active Opening

Cloud Systems Engineer

Alarmcom • Tysons, Virginia

Job Description

<p>We are seeking a Cloud Systems Engineer to support and operate large-scale AI and high-performance computing (HPC) environments. This role will be responsible for the deployment, maintenance, performance, and lifecycle management of GPU-accelerated compute infrastructure that powers critical AI, machine learning, and data-intensive workloads.</p> <p>The ideal candidate is a hands-on infrastructure professional with strong Linux administration skills, deep hardware troubleshooting experience, and expertise supporting enterprise-class compute platforms. This individual will work closely with infrastructure, networking, storage, and AI engineering teams to ensure the reliability, scalability, and operational excellence of our AI infrastructure.</p> <p><strong>Responsibilities:</strong></p> <p><strong>AI Infrastructure Operations</strong></p> <ul> <li>Deploy, configure, and maintain GPU-accelerated compute infrastructure.</li> <li>Manage operating system, firmware, BIOS, BMC, driver, and software lifecycle updates.</li> <li>Monitor system health, performance, utilization, and capacity across AI infrastructure environments.</li> <li>Support infrastructure utilized for AI model training, inference, and data processing workloads.</li> <li>Develop and maintain operational standards, runbooks, and maintenance procedures.</li> <li>Participate in on-call support and incident response activities.</li> </ul> <p><strong>Linux Systems Administration</strong></p> <ul> <li>Administer enterprise Linux environments, including Ubuntu and Red Hat-based distributions.</li> <li>Perform system patching, hardening, and operating system lifecycle management.</li> <li>Troubleshoot operating system, kernel, storage, networking, and application-level issues.</li> <li>Develop automation to streamline deployment, monitoring, and operational processes.</li> <li>Support security and compliance initiatives across AI infrastructure platforms.</li> </ul> <p><strong>Hardware and Datacenter Operations</strong></p> <ul> <li>Install, configure, maintain, and troubleshoot enterprise compute hardware.</li> <li>Diagnose and resolve issues involving GPUs, CPUs, memory, storage, power, and networking components.</li> <li>Perform firmware upgrades and hardware lifecycle management activities.</li> <li>Coordinate hardware replacements, vendor support engagements, and warranty services.</li> <li>Participate in rack-and-stack deployments, datacenter expansions, and technology refresh projects.</li> <li>Maintain accurate asset inventories and operational documentation.</li> <li>Support high-performance networking technologies, including Ethernet and InfiniBand environments.</li> <li>Collaborate with networking, storage, cloud, and AI engineering teams on infrastructure design and operations.</li> <li>Assist with scalability, resiliency, and performance optimization initiatives.</li> <li>Perform root-cause analysis of infrastructure failures and develop preventative measures.</li> <li>Other duties as assigned. </li> </ul> <p><strong>Required Qualifications</strong></p> <p><strong>Experience</strong></p> <ul> <li>Bachelor’s degree required</li> <li>3-5 years of Linux systems administration experience in production environments.</li> <li>3-5 years of experience supporting enterprise server infrastructure.</li> <li>Experience supporting large-scale compute environments, HPC platforms, AI infrastructure, or GPU-enabled systems.</li> <li>Experience perform

Job Reference ID: CF-147575 • Posted on CloudFrame Job Scanner