Staff Software Engineer, AI Reliability Engineering

3 weeks agoHeadSalary: 235K - 295KDublin, IEHybridDevopsJobs by Anthropic

AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.

Series F+Recently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/23/2026

San Francisco, United States
About Anthropic

Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.

View jobs by Anthropic

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will define service-level objectives for language-model serving, build monitoring and observability across the token path, and help implement highly available multi-region infrastructure. You will lead responses to critical incidents, conduct reviews, and improve the reliability of safeguard-model serving.

Requirements

  • Distributed systems, infrastructure, or reliability background
  • Communication skills
  • Collaboration skills

Responsibilities

  • Develop service-level objectives for language-model serving systems
  • Design and implement monitoring and observability systems across the token path
  • Assist with high-availability serving infrastructure across regions and cloud providers
  • Lead incident response, recovery, incident reviews, and systematic improvements
  • Support the reliability of safeguard-model serving

Benefits

  • Optional equity donation matching
  • Generous vacation
  • Parental leave
  • Flexible working hours
Staff Software Engineer, AI Reliability Engineering at Anthropic | JobStash