Research Engineer RL Scaling Science

3 weeks agoSalary: 375K - 640KLondon, UKHybridResearchJobs by Anthropic

AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.

Series F+Recently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/23/2026

San Francisco, United States
About Anthropic

Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.

View jobs by Anthropic

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design, run, and interpret large-scale reinforcement learning experiments. You will build and maintain long-horizon benchmarks, investigate scaling behavior, debug research and infrastructure issues, and turn robust findings into production training recipes.

Requirements

  • Empirical research skills in reinforcement learning, large-scale machine learning training, or a closely related area
  • Ability to own large experiments end to end
  • Python proficiency
  • Experience with large-scale or distributed machine learning systems
  • Ability to work at the research and systems boundary
  • Knowledge of responsible AI scaling

Responsibilities

  • Design, run, and interpret large-scale reinforcement learning experiments
  • Investigate reinforcement learning scaling across horizon, compute, and model size
  • Build and maintain benchmarks for long-horizon reinforcement learning
  • Translate validated findings into production training recipes
  • Debug complex research and infrastructure issues at scale
  • Partner with adjacent reinforcement learning teams

Benefits

  • Visa sponsorship
  • Equity donation matching
  • Generous vacation
  • Parental leave
  • Flexible working hours
Research Engineer RL Scaling Science at Anthropic | JobStash