Research Engineer Domain Scaling

AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.

Series F+Recently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/24/2026

San Francisco, United States
About Anthropic

Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.

View jobs by Anthropic

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own data strategy from task sourcing through reinforcement-learning training. You will manage technical data-vendor relationships, design reward signals, data pipelines, evaluations, and QA frameworks, run generalization experiments, and translate capability goals into training environments.

Requirements

  • Experience fine-tuning large language models for domains or real-world use cases
  • Experience with reinforcement learning, reward design, or LLM training-data curation
  • Experience managing technical vendor relationships
  • Ability to assess datasets and identify issues
  • Cross-functional collaboration skills

Responsibilities

  • Own data strategy from task sourcing through reinforcement learning training
  • Manage technical relationships with external data vendors
  • Evaluate data quality and reward design
  • Collaborate with domain experts to design data pipelines and evaluations
  • Create reinforcement learning environments for high-value tasks
  • Develop QA frameworks to detect reward hacking and ensure environment quality
  • Run generalization experiments
  • Translate capability goals into training environments and evaluations

Benefits

  • Optional equity donation matching
  • Generous vacation
  • Parental leave
  • Flexible working hours