Technical Support Engineer

TensorWave is an AMD-exclusive AI cloud provider for large-model training, fine-tuning, and inference.

Las Vegas, United States
About TensorWave

TensorWave provides AMD Instinct GPU-based cloud infrastructure, managed services, storage, networking, and operational support for AI workloads.

View jobs by TensorWave

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own escalated level 2 support tickets and diagnose Linux, GPU, Kubernetes, Slurm, storage, networking, and workload issues. You will communicate findings to customers, coordinate hardware remediation, maintain runbooks and documentation, build diagnostic tools, support account health, join on-call coverage, and share recurring patterns with engineering and product.

Requirements

  • 3+ years of Linux systems administration, site reliability engineering, infrastructure operations, or technical support engineering experience
  • Strong Linux troubleshooting skills across systems, networking, file systems, storage, processes, resources, and logs
  • Working knowledge of Kubernetes cluster and pod state, scheduling, placement, and workload diagnosis
  • Experience with a batch scheduler in a shared compute environment, ideally Slurm
  • Python or Bash scripting ability
  • Experience using incident and tracking systems such as JIRA or PagerDuty
  • Excellent written communication skills
  • Willingness to participate in an on-call rotation
  • Background in GPU cloud, HPC, or AI/ML infrastructure operations
  • Familiarity with AMD Instinct platforms, ROCm, or NVIDIA/CUDA
  • Experience diagnosing RDMA and high-speed networking issues
  • Exposure to Weka, Vast, or high-performance storage environments
  • Experience with Docker, Enroot, Pixis, or Aptainer
  • Experience with Grafana and Prometheus
  • Understanding of PyTorch, JAX, and inference-serving patterns
  • Experience with ITIL or another incident and change management framework

Responsibilities

  • Own escalated level 2 tickets through resolution or a well-evidenced engineering handoff
  • Diagnose Linux hosts, GPU health, Kubernetes workloads, Slurm scheduling, storage, and high-speed networking issues
  • Investigate degraded and failed training and inference workloads using logs, metrics, and cluster telemetry
  • Triage GPU and node hardware faults and coordinate remediation with data center operations
  • Communicate directly with customer engineering teams during investigations
  • Write and maintain runbooks for recurring issues
  • Build diagnostic scripts and tools that reduce investigation time and repeat work
  • Contribute to customer-facing knowledge base documentation
  • Partner with technical account managers on account health
  • Participate in an on-call rotation
  • Provide evidence of repeated patterns to engineering and product

Benefits

  • Stock options
  • 100% paid medical, dental, and vision insurance for employees
  • Company Health Savings Account contributions
  • 100% paid short-term and long-term disability insurance for employees
  • Life and voluntary supplemental insurance options
  • Pet and legal insurance options
  • Discounted virtual healthcare appointments and serious illness support
  • Flexible Spending Account
  • 401(k)
  • Employee Assistance Program
  • Flexible PTO
  • Paid holidays
  • Parental leave
  • In-office perks
Technical Support Engineer at TensorWave | JobStash