Research Engineer Code Reinforcement Learning

AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.

Series F+Recently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/23/2026

San Francisco, United States
About Anthropic

Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.

View jobs by Anthropic

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design coding tasks, reinforcement-learning environments, rewards, and verifiers. You will run frontier-model training experiments, diagnose model behavior, and improve the speed and reliability of training pipelines for safe, capable software-engineering work.

Requirements

  • Software engineering skills
  • Deep Python expertise
  • Async and concurrent programming
  • End-to-end systems ownership
  • Debugging across the stack
  • Experimental design and results interpretation
  • Code quality, testing, and performance knowledge

Responsibilities

  • Design reinforcement-learning environments and coding tasks
  • Build reward signals and verifiers for code quality
  • Run training experiments on frontier models
  • Diagnose model performance on software-engineering tasks
  • Improve training-pipeline speed and reliability

Benefits

  • Optional equity donation matching
  • Generous vacation
  • Parental leave
  • Flexible working hours
Research Engineer Code Reinforcement Learning at Anthropic | JobStash