Reliability Engineer Supercomputing
39 minutes agoSalary: 350K - 475KSan Francisco, USAHybridFull TimeDevopsJobs by Thinking Machines Lab
Thinking Machines LabVisit Thinking Machines Lab website
Artificial-intelligence research and product company building customizable AI systems, including the Tinker training API and Inkling open-weight models.
Distributed
Funding history
About Thinking Machines Lab
Thinking Machines Lab develops AI products that let researchers and developers fine-tune and use models, while also releasing open-weight multimodal models.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will investigate and remediate failures across GPU clusters, owning diagnostics across hardware, firmware, drivers, and operating systems. You will automate reliability monitoring, manage firmware lifecycles, work directly with vendors, oversee hardware RMAs, and create postmortems that lead to lasting fixes.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, or a similar field
- Proficiency in at least one backend language
- Experience operating large-scale clusters and container orchestration systems
- Ability to operate across the stack and own projects end-to-end
Responsibilities
- Investigate, reproduce, and remediate issues across large GPU clusters
- Own drivers, kernel surfaces, and diagnostics spanning hardware, firmware, and operating systems
- Automate fleet-reliability monitoring and analyze error rates
- Drive firmware tracking, qualification, staged rollout, and regression analysis
- Engage GPU, server, NIC, and storage vendors to resolve issues
- Manage RMA flows for hardware replacements
- Monitor and improve GPU hardware-health signals
- Write postmortems and vendor cases
Benefits
- Visa sponsorship
- Health benefits
- Dental benefits
- Vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
