Staff Reliability Engineer
Tenstorrent is an AI-computing company that sells AI hardware and licenses AI and RISC-V intellectual property.
Funding history
Investors
About Tenstorrent
Tenstorrent builds computers for AI, including AI processors, scalable server systems, and open-source software and compiler tooling. It also licenses AI and RISC-V IP for customers building customized silicon.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will define and implement reliability strategy for next-generation AI computing systems. You will develop MTBF models, lead root-cause investigations, oversee accelerated life testing, assess design risks, and translate test findings into improvements. You will partner with engineering, validation, manufacturing, and supply-chain stakeholders and travel to third-party test facilities.
Requirements
- 8+ years of reliability engineering experience
- Knowledge of HALT, HASS, ALT, MTBF, Weibull analysis, and FMEA
- Ability to analyze thermal-lab technical problems and communicate risks and trade-offs
- Ability to coordinate mechanical, electrical, thermal, software, and validation teams
- Bachelor's or Master's degree in Mechanical Engineering, Electrical Engineering, Reliability Engineering, or a related field
- Eligibility to access U.S. export-controlled technology
Responsibilities
- Set reliability strategy for next-generation AI computing systems
- Develop MTBF models and predictive reliability frameworks
- Lead root-cause investigations and drive corrective actions
- Summarize findings and feed them back to Systems Engineering
- Partner with System Development and Compliance Validation teams
- Plan and oversee HALT and HASS testing at third-party facilities
- Assess design risk and resolve device-under-test issues on site
