Research Engineer RL Engineering

AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.

Series F+Recently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/24/2026

San Francisco, United States
About Anthropic

Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.

View jobs by Anthropic

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build, maintain, and improve the algorithms and systems used to train models with reinforcement learning. You will improve their speed, reliability, and usability, profile training pipelines, automate test training jobs, support new model architectures, and diagnose performance problems.

Requirements

  • Software engineering
  • System
  • Tooling
  • Large-scale distributed system
  • Large-scale large language model training
  • Python
  • Large language model fine-tuning algorithm
  • RLHF

Responsibilities

  • Build, maintain, and improve model-training algorithms and systems
  • Improve the speed, reliability, and usability of training systems
  • Profile reinforcement-learning pipelines
  • Build systems that launch test training jobs
  • Adapt fine-tuning systems to new model architectures
  • Build instrumentation to detect and eliminate Python GIL contention
  • Diagnose and fix training-performance issues
  • Implement stable and fast training algorithms

Benefits

  • Visa sponsorship
  • Equity donation matching
  • Vacation leave
  • Parental leave
  • Flexible working hours