Research Engineer RL Scaling Science
AnthropicVisit Anthropic website
AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.
Maintainer signals as of 9/23/2026
San Francisco, United States
Funding history
About Anthropic
Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design, run, and interpret large-scale reinforcement learning experiments. You will build and maintain long-horizon benchmarks, investigate scaling behavior, debug research and infrastructure issues, and turn robust findings into production training recipes.
Requirements
- Empirical research skills in reinforcement learning, large-scale machine learning training, or a closely related area
- Ability to own large experiments end to end
- Python proficiency
- Experience with large-scale or distributed machine learning systems
- Ability to work at the research and systems boundary
- Knowledge of responsible AI scaling
Responsibilities
- Design, run, and interpret large-scale reinforcement learning experiments
- Investigate reinforcement learning scaling across horizon, compute, and model size
- Build and maintain benchmarks for long-horizon reinforcement learning
- Translate validated findings into production training recipes
- Debug complex research and infrastructure issues at scale
- Partner with adjacent reinforcement learning teams
Benefits
- Visa sponsorship
- Equity donation matching
- Generous vacation
- Parental leave
- Flexible working hours
