Senior Machine Learning Engineer

TensorWave is an AMD-exclusive AI cloud provider for large-model training, fine-tuning, and inference.

Las Vegas, United States
About TensorWave

TensorWave provides AMD Instinct GPU-based cloud infrastructure, managed services, storage, networking, and operational support for AI workloads.

View jobs by TensorWave

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design, operate, and improve ML infrastructure for distributed training and inference. You will build workload-execution and orchestration patterns across shared GPU environments, troubleshoot performance and reliability issues, and improve developer experience and operational efficiency.

Requirements

  • Experience supporting production ML systems using SLURM and Kubernetes
  • Understanding of GPU-accelerated workloads and distributed systems
  • Linux debugging experience
  • Ability to build automation and tooling using Python or Go

Responsibilities

  • Design, operate, and improve ML infrastructure for distributed training and inference
  • Build workload-execution and orchestration patterns across shared GPU environments
  • Troubleshoot ML-stack performance, reliability, and scalability issues
  • Improve developer experience and operational efficiency with ML, systems, and platform partners

Benefits

  • Stock options
  • 100% paid medical, dental, and vision insurance for employees
  • Company health savings account contributions
  • 100% paid short-term and long-term disability insurance for employees
  • Life and voluntary supplemental insurance options
  • Pet and legal insurance options
  • Supplementary health benefits
  • Flexible spending account
  • 401(k)
  • Employee assistance program
  • Flexible PTO
  • Paid holidays
  • Parental leave
  • In-office perks