Research Engineer - RL Infrastructure
Prime Intellect provides an open superintelligence stack for training, evaluating, deploying, and continuously improving AI agents and models. Its platform combines RL environments, hosted training, inference, GPU compute, secure sandboxes, and open-source research tooling for researchers, startups, and enterprises.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Prime Intellect, Inc.
Prime Intellect operates an integrated AI infrastructure platform spanning Lab, hosted reinforcement-learning training, evaluations, environments, inference, secure sandboxes, and on-demand or reserved GPU compute. It also develops open-source tools including Verifiers, prime-rl, and Prime Agent, supporting workflows from environment creation and model evaluation through post-training and production deployment. The company serves researchers, startups, enterprises, and teams building agentic AI systems, with customer examples including Ramp and Zapier.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and optimize the systems infrastructure behind large-scale reinforcement learning and distributed training workloads. You will improve efficiency across compute, memory, networking, and scheduling, implement low-level optimizations, develop RL training systems, and collaborate with researchers on frontier-scale model training.
Requirements
- AI/ML infrastructure engineering experience
- Large-scale model training or inference experience
- PyTorch
- PyTorch Distributed
- DeepSpeed
- FSDP
- Megatron
- vLLM
- Ray
- Training performance optimization
- Data parallelism
- Tensor parallelism
- Pipeline parallelism
- GPU architecture
- Profiling
- Performance debugging
- CUDA
- Triton
- Compiler optimization
- Runtime optimization
- RL training infrastructure
- Rollout systems
- Asynchronous training pipelines
- Multi-node GPU clusters
- High-performance networking
- Open-source contributions
Responsibilities
- Build and optimize systems infrastructure for large-scale RL and distributed training workloads
- Improve training efficiency across compute, memory, networking, and scheduling layers
- Design and implement kernel, communication path, and runtime optimizations
- Develop distributed training systems for data, tensor, and pipeline parallel workloads
- Shape the architecture of the RL training stack
- Contribute to open-source libraries and internal infrastructure
- Translate system bottlenecks into concrete improvements
- Track advances in training systems, inference systems, compiler tooling, runtime tooling, and hardware-aware optimization techniques
Benefits
- Equity
- Flexible work arrangements
- Remote or in-person work options
- Visa sponsorship
- Relocation support
- Quarterly team offsites
- Hackathons
- Conferences
- Learning opportunities
