CF CloudFrame Job Scanner Open Dashboard →
Verified Active Opening

Reinforcement Learning Engineer

Bugcrowd • Remote - US

Job Description

<div class="content-intro"><p>Founded in 2012, Bugcrowd is the preemptive security platform that unifies exposure discovery and assessment, offensive testing, and intelligence shaped by AI and human insight to help organizations avoid, discover, and validate real-world risk. Bugcrowd helps security teams move faster by identifying the exposures that matter most so they can act first and stay ahead of attackers. By combining the power of humans and AI, teams can preempt attack paths and prevent breaches. Based in San Francisco and New Hampshire, Bugcrowd is supported by General Catalyst, Rally Ventures, Costanoa Ventures, and others. Visit <span><a href="http://www.bugcrowd.com" target="_blank">www.bugcrowd.com</a></span>.</p></div><p><strong>Job Summary</strong></p> <p>The Bugcrowd RL and Reasoning Team focuses on pushing the boundaries of autonomous cybersecurity by building authentic, verifiable reinforcement learning environments for world-leading foundational AI companies. As a Reinforcement Learning Engineer specializing in Reinforcement Learning from Verifiable Rewards (RLVR), you will design and scale automated verification pipelines that transform real-world software vulnerabilities into deterministic reward functions. In this role, you will bridge the gap between low-level security analysis and modern LLM reasoning models, engineering environments where AI agents learn to discover, exploit, and remediate software vulnerabilities with mathematical certainty. Instead of relying on subjective human feedback, your work directly powers the rigorous, verifiable reward signals that teach next-generation frontier AI models how to master complex cybersecurity domain logic. You will work at the intersection of fuzzing, dynamic program analysis, system exploitation, and scalable ML infrastructure to shape the safety and offensive/defensive capabilities of future artificial intelligence. </p> <p><strong>Essential Duties and Responsibilities </strong></p> <ul> <li>Design, build, and deploy high-throughput RLVR (Reinforcement Learning from Verifiable Rewards) environments that evaluate LLM action sequences against deterministic execution outcomes.</li> <li>Develop automated test harnesses, sandboxes, and verification engines that convert complex vulnerability research (e.g., memory corruption, web security, logic bugs) into binary pass/fail reward signals.</li> <li>Integrate Bugcrowd’s Mayhem automated analysis platform and real-world vulnerability feeds into continuous, scalable RL environment generation pipelines.</li> <li>Architect safe, isolated, and highly reproducible execution environments (using Docker, BuildKit, or Nix) capable of running thousands of simultaneous agent-driven exploitation and patching trajectories.</li> <li>Collaborate directly with researchers at frontier AI labs including Anthropic, OpenAI, and Cohere to define standard benchmark formats, observation spaces, and verifiable evaluation metrics for cybersecurity tasks.</li> <li>Implement precise telemetry, ground-truth verification algorithms, and trajectory logging to analyze agent reasoning paths and prevent reward hacking or false positives.</li> <li>Build low-level instrumentation and debugging tools to monitor memory states, process executions, and network behaviors during agent interaction cycles.</li> <li>Optimize infrastructure performance and environment reset latency to support massive-scale parallel sampling and distributed RL training workflows.</li> <li>Benchmark and evaluate frontier AI model performance across diverse offensive and defensive security challenges, such as automated fuzzing, exploit payload generation, and patch validation.</li> </ul> <p><strong>Education, Experience, Knowledge, Ski

Job Reference ID: CF-167263 • Posted on CloudFrame Job Scanner