Research Engineer - Distributed Training
Prime IntellectVisit Prime Intellect website
AI infrastructure company providing an integrated stack for training, evaluating, deploying, and continuously improving agentic models.
Prime Intellect on X (Twitter)Prime Intellect on DiscordPrime Intellect on GitHubPrime Intellect on Documentation
Series ARecently funded34 current maintainers27 active leads7 new active leads9 lead step-downsTeam intelligence
Maintainer signals as of 9/25/2026
San Francisco, United States
Funding history
Projects
About Prime Intellect
Prime Intellect, Inc. operates AI infrastructure spanning RL environments, hosted training and evaluations, inference, secure sandboxes, and globally sourced GPU compute.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Build and optimize distributed training infrastructure for pre-training and large-scale reinforcement learning workloads, improving efficiency across compute, memory, networking, scheduling, and runtime systems.
Requirements
- AI/ML infrastructure engineering experience
- Large-scale model training or inference experience
- PyTorch
- PyTorch Distributed
- DeepSpeed
- FSDP
- Megatron
- vLLM
- Ray
- Training performance optimization
- Data parallelism
- Tensor parallelism
- Pipeline parallelism
- GPU architecture
- Profiling
- Performance debugging
- CUDA
- Triton
- Compiler optimization
- Runtime optimization
- RL training infrastructure
- Multi-node GPU clusters
- High-performance networking
- Open-source contributions
Responsibilities
- Build and optimize distributed training infrastructure
- Improve training efficiency across compute, memory, networking, and scheduling layers
- Design and implement kernel, communication path, and runtime optimizations
- Develop distributed training systems for data, tensor, and pipeline parallel workloads
- Shape the architecture of the RL training stack
- Contribute to open-source libraries and internal infrastructure
- Translate system bottlenecks into concrete improvements
- Track advances in training systems, inference systems, compiler/runtime tooling, and hardware-aware optimization
Benefits
- Equity incentives
- Flexible work arrangements
- Remote or in-person work options
- Visa sponsorship
- Relocation assistance
- Quarterly team off-sites
- Hackathons
- Conferences
- Learning opportunities
