Research Engineer - Distributed Training

Prime Intellect provides an open superintelligence stack for training, evaluating, deploying, and continuously improving AI agents and models. Its platform combines RL environments, hosted training, inference, GPU compute, secure sandboxes, and open-source research tooling for researchers, startups, and enterprises.

Maintainer signals as of 8/23/2026

San Francisco, USA
About Prime Intellect, Inc.

Prime Intellect operates an integrated AI infrastructure platform spanning Lab, hosted reinforcement-learning training, evaluations, environments, inference, secure sandboxes, and on-demand or reserved GPU compute. It also develops open-source tools including Verifiers, prime-rl, and Prime Agent, supporting workflows from environment creation and model evaluation through post-training and production deployment. The company serves researchers, startups, enterprises, and teams building agentic AI systems, with customer examples including Ramp and Zapier.

View jobs by Prime Intellect, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Build and optimize distributed training infrastructure for pre-training and large-scale reinforcement learning workloads, improving efficiency across compute, memory, networking, scheduling, and runtime systems.

Requirements

  • AI/ML infrastructure engineering experience
  • Large-scale model training or inference experience
  • PyTorch
  • PyTorch Distributed
  • DeepSpeed
  • FSDP
  • Megatron
  • vLLM
  • Ray
  • Training performance optimization
  • Data parallelism
  • Tensor parallelism
  • Pipeline parallelism
  • GPU architecture
  • Profiling
  • Performance debugging
  • CUDA
  • Triton
  • Compiler optimization
  • Runtime optimization
  • RL training infrastructure
  • Multi-node GPU clusters
  • High-performance networking
  • Open-source contributions

Responsibilities

  • Build and optimize distributed training infrastructure
  • Improve training efficiency across compute, memory, networking, and scheduling layers
  • Design and implement kernel, communication path, and runtime optimizations
  • Develop distributed training systems for data, tensor, and pipeline parallel workloads
  • Shape the architecture of the RL training stack
  • Contribute to open-source libraries and internal infrastructure
  • Translate system bottlenecks into concrete improvements
  • Track advances in training systems, inference systems, compiler/runtime tooling, and hardware-aware optimization

Benefits

  • Equity incentives
  • Flexible work arrangements
  • Remote or in-person work options
  • Visa sponsorship
  • Relocation assistance
  • Quarterly team off-sites
  • Hackathons
  • Conferences
  • Learning opportunities