Hardware Diagnostics Engineer

TensorWave is an AMD-exclusive AI cloud provider for large-model training, fine-tuning, and inference.

Las Vegas, United States
About TensorWave

TensorWave provides AMD Instinct GPU-based cloud infrastructure, managed services, storage, networking, and operational support for AI workloads.

View jobs by TensorWave

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will validate server and GPU hardware through burn-in and stress testing, diagnose failures, and determine production readiness. You will manage servers through out-of-band tools, apply firmware updates, coordinate vendor RMAs, maintain asset records, identify fleet failure patterns, automate repeatable work, support datacenter expansions, and participate in hardware on-call escalation.

Requirements

  • 3–6 years of experience in datacenter operations, systems administration, hardware support, or infrastructure engineering
  • Hands-on experience with enterprise server hardware, component replacement, POST and boot failures, and rack-level hardware behavior
  • Experience with BMCs and out-of-band management, including IPMI, Redfish, iDRAC, iLO, or equivalent
  • Strong Linux troubleshooting skills, including boot processes, drivers, devices, dmesg, lspci, ipmitool, and SMART
  • Ability to interpret sensor data, event logs, thermal telemetry, and power telemetry
  • Working Bash or Python scripting ability
  • Experience managing hardware RMAs with vendors or driving external issues to closure
  • Methodical troubleshooting and evidence-based diagnosis
  • Clear written communication for tickets, runbooks, and vendor cases

Responsibilities

  • Run server and GPU burn-in and stress testing and determine production readiness
  • Triage failures across GPUs, memory, drives, NICs, PSUs, and cabling
  • Manage servers out-of-band through IPMI and Redfish
  • Apply qualified firmware updates across the fleet
  • Drive vendor RMAs through replacement, installation, and failed-part return
  • Maintain asset, serial, and replacement records in NetBox
  • Track and escalate fleet failure patterns
  • Improve runbooks and script repeatable steps
  • Support datacenter turn-ups and expansions
  • Participate in an on-call and escalation rotation for hardware issues

Benefits

  • Stock options
  • 100% paid medical, dental, and vision insurance for employees
  • Company Health Savings Account contributions
  • 100% paid short-term and long-term disability insurance for employees
  • Life and voluntary supplemental insurance options
  • Pet and legal insurance
  • Supplementary health benefits, including discounted virtual healthcare appointments and serious illness support
  • Flexible Spending Account
  • 401(k)
  • Employee Assistance Program
  • Flexible PTO
  • Paid holidays
  • Parental leave
  • In-office perks