Head of Platform Product Reliability

Etched is an AI-hardware company building rack-scale frontier inference clusters.

San Jose, United States
About Etched

Etched co-designs chips, racks, software, and manufacturing systems for efficient inference of frontier AI models, targeting throughput, latency, cost, and power efficiency across prefill and decode workloads.

View jobs by Etched

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own end-to-end product reliability for server, rack, and datacenter hardware from architecture through field deployment. You will define requirements and qualification methods, lead stress testing and failure analysis, develop reliability models and fleet monitoring, drive release readiness, and build a reliability engineering organization.

Requirements

  • Degree in electrical engineering, mechanical engineering, reliability engineering, or a related technical field
  • 10+ years of reliability engineering experience in hardware-centric organizations and complex systems
  • Experience leading reliability programs for AI accelerator or GPU compute systems, server infrastructure, networking platforms, storage systems, or rack-scale infrastructure
  • Knowledge of system-level thermal, power-delivery, mechanical, and connector or interconnect failure mechanisms
  • Experience with FMEA, Weibull analysis, HALT, HASS, qualification planning, failure analysis, and reliability statistics and modeling
  • Experience driving cross-functional root-cause investigations
  • Ability to make tradeoffs among reliability targets, cost, schedule, and performance
  • Communication skills
  • Experience with liquid-cooled systems, high-density power delivery, or thermal management
  • Experience supporting hyperscale or cloud datacenter deployments, customer-facing reliability commitments, or SLA management
  • Experience building a reliability organization
  • Familiarity with fleet telemetry, field reliability analytics, and proactive reliability management
  • Experience working with ODM or JDM partners, including NPI support and on-site qualification
  • Background in high-speed digital systems, GPU compute platforms, or accelerator-based architectures

Responsibilities

  • Define and own reliability strategy for AI servers, accelerator platforms, rack systems, and datacenter infrastructure
  • Establish reliability requirements, qualification standards, and validation methodologies
  • Build reliability engineering processes across the product lifecycle
  • Define EVT, DVT, and PVT qualification gates and exit criteria
  • Lead accelerated life and stress testing programs
  • Lead environmental, HALT, HASS, vibration, shock, transportation, power-cycling, thermal-cycling, and soak testing
  • Lead root-cause investigations and corrective actions for reliability failures
  • Develop MTBF projections, FIT-rate analysis, Weibull lifetime models, component derating methods, and reliability growth tracking
  • Ensure reliability requirements inform early design decisions
  • Work with ODMs, JDMs, contract manufacturers, and suppliers on reliability commitments
  • Build fleet reliability infrastructure, telemetry analysis pipelines, field feedback loops, and monitoring frameworks
  • Drive reliability signoff and product release-readiness reviews
  • Build and lead a product reliability engineering organization

Benefits

  • Medical, dental, and vision coverage
  • USD 500 monthly credit for waiving medical benefits
  • USD 2,500 monthly housing subsidy for employees living within walking distance of the office
  • Relocation support to San Jose
  • Wellness benefits
  • Daily office lunch and dinner
  • Unlimited compute budget subject to ROI justification
Head of Platform Product Reliability at Etched | JobStash