Research Engineer or Research Scientist RL Frontiers
2 days agoSalary: 500K - 850KSan Francisco, CA | New York City, NY | Seattle, WAHybridResearchJobs by Anthropic
AnthropicVisit Anthropic website
AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.
Maintainer signals as of 9/25/2026
San Francisco, United States
Funding history
About Anthropic
Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will study reinforcement learning at increasing model, context, and compute scale. You will develop architectures and algorithms, build reproducible experimental infrastructure, diagnose training issues, and improve the performance and cost of large-scale training runs.
Requirements
- Modern transformer language model architecture, training dynamics, and large-scale optimization knowledge
- Experience training large distributed models using data, tensor, and pipeline parallelism
- Track record of original technical work in machine learning training or systems
- Ability to design rigorous large-scale experiments with baselines and ablations
- Ability to reason quantitatively about compute, memory, and communication costs
- Python and JAX or PyTorch programming skills
Responsibilities
- Study how reinforcement learning training and sampling scale with model size, context length, and compute
- Develop model architectures and reinforcement learning algorithms for frontier-scale execution
- Scale promising small-scale results and diagnose numerical, algorithmic, and systemic differences
- Build reproducible experimental infrastructure for architecture and algorithm comparisons
- Own end-to-end performance of large reinforcement learning runs
- Build performance and cost models for architecture and algorithm changes
- Investigate and trace training instabilities, divergence, and throughput regressions to root causes
Benefits
- Optional equity donation matching
- Generous vacation
- Parental leave
- Flexible working hours
