Site Reliability Engineer Production

Artificial-intelligence research and product company building customizable AI systems, including the Tinker training API and Inkling open-weight models.

Distributed
About Thinking Machines Lab

Thinking Machines Lab develops AI products that let researchers and developers fine-tune and use models, while also releasing open-weight multimodal models.

View jobs by Thinking Machines Lab

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will define end-to-end reliability across CI/CD, observability, and incident response. You will develop service-level objectives, implement monitoring, lead incident recovery and reviews, improve multi-tenant isolation and scheduling, and address production vulnerabilities with security teams.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or a similar field
  • Experience in distributed systems, cloud infrastructure, or site reliability engineering
  • Experience writing software, tooling, and automation to solve reliability problems
  • Experience with production incident response, postmortems, and systematic reliability improvement
  • Strong communication and cross-functional coordination skills

Responsibilities

  • Define and own end-to-end reliability across CI/CD, observability, and incident response
  • Develop service-level objectives for distributed training systems
  • Design and implement monitoring and observability across the training path
  • Drive incident response, recovery, incident reviews, and preventive improvements
  • Harden multi-tenant isolation and resource scheduling
  • Collaborate with security teams to address production vulnerabilities

Benefits

  • Visa sponsorship
  • Health benefits
  • Dental benefits
  • Vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
Site Reliability Engineer Production at Thinking Machines Lab | JobStash