Research Engineer Post Training
Pluralis Research is a research lab developing decentralized AI through Protocol Learning, a communication-efficient approach to collaborative model training. It operates open training systems that allow distributed contributors to provide compute for collectively owned foundation models.
Funding history
About Pluralis Research
Pluralis Research develops Protocol Learning, which enables foundation models to be trained and served across globally distributed participants without requiring a single participant to hold the complete model. Its work covers low-bandwidth model parallelism, asynchronous distributed optimization, fault-tolerant training, privacy-preserving unextractable models, and collective ownership. The organization operates Agora and Node0 training systems, publishes research, provides participation documentation, and releases open-source software for distributed training and reinforcement-learning workflows.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build an end-to-end RL post-training stack, including rollout ingestion, reward computation, policy updates, and weight distribution. You will adapt RL algorithms for asynchronous, high-latency, partially trusted generation, create evaluations, and deliver decentralized post-trained model releases.
Requirements
- Hands-on RL post-training experience for large language models
- Experience with rollout generation, asynchronous training loops, and weight synchronization
- Production-quality Python and PyTorch skills
- Experience with concurrency, failure handling, and profiling
- Research ability in RL post-training, asynchronous RL, distributed RL, or related fields
- Professional-level written and spoken English
Responsibilities
- Build the end-to-end RL post-training stack
- Adapt RL algorithms for asynchronous and high-latency generation
- Build evaluations that measure model improvement
- Deliver decentralized post-trained model releases
Benefits
- Significant ownership for key technical contributors
- Flexible remote work environment
- Optional visa sponsorship and relocation support to Australia or the US
