Research Engineer Pretraining Scaling

3 weeks agoSalary: 260K - 630KLondon, UKOnsiteAiJobs by Anthropic

AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.

Series F+Recently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/23/2026

San Francisco, United States
About Anthropic

Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.

View jobs by Anthropic

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own and improve production pretraining systems, including performance, reliability, observability, and model operations. You will resolve hardware, networking, training, and evaluation issues; run efficiency experiments; respond to launch incidents; maintain monitoring systems; extend training code; and document operational knowledge.

Requirements

  • Large language model training
  • JAX
  • TPU
  • PyTorch
  • Distributed systems
  • Performance optimization
  • Hardware debugging
  • Experimental design
  • Observability
  • Production systems
  • On-call support
  • AI safety

Responsibilities

  • Own production pretraining pipeline operations, performance optimization, observability, and reliability
  • Debug and resolve hardware, networking, training, and evaluation issues
  • Design and run experiments to improve training efficiency and model performance
  • Respond to on-call incidents during model launches
  • Build and maintain production logging, monitoring dashboards, and evaluation infrastructure
  • Add capabilities to the training codebase
  • Document systems, debugging approaches, and lessons learned

Benefits

  • Optional equity donation matching
  • Generous vacation
  • Parental leave
  • Flexible working hours