Customer Success Engineer GPU Cluster

Together AI operates an AI-native cloud platform for open and custom AI models.

San Francisco, United States
About Together AI

Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.

View jobs by Together AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own the technical relationship with a strategic customer across compute, networking, storage, and facilities. You will manage incidents, RMAs, observability, capacity expansions, operational reviews, and data center events while providing infrastructure guidance and coordinating delivery through production acceptance.

Requirements

  • 5+ years in a customer-facing technical role
  • 2+ years in dedicated technical account management or solutions architecture for large-scale AI or HPC infrastructure
  • GPU infrastructure expertise, including diagnostics, RMA workflows, and acceptance testing
  • Experience with large-scale Ethernet and InfiniBand fabric architecture
  • Knowledge of enterprise storage systems, NVMe, parallel file systems, and metadata infrastructure
  • Experience with data center operations, facilities coordination, and hosting-provider SLA management
  • Incident management and root-cause analysis experience
  • Experience with Prometheus, Grafana, or equivalent observability tooling
  • Ability to manage concurrent workstreams
  • Python, Bash, or infrastructure automation proficiency preferred

Responsibilities

  • Serve as the technical point of contact for a strategic customer
  • Drive status reporting, technical steering meetings, QBRs, and EBRs
  • Translate customer feedback into roadmap input
  • Lead incident lifecycle management, escalations, and root-cause analysis
  • Coordinate RMAs, hardware lifecycle management, acceptance testing, and spare inventory
  • Advise on GPU compute, fabric, storage, and incident resolution
  • Define alert policies, develop dashboards, and manage infrastructure health
  • Coordinate data center operations and facilities events
  • Manage capacity expansions and node deployment through production acceptance

Benefits

  • Startup equity
  • Health insurance
  • Remote-work flexibility
Customer Success Engineer GPU Cluster at Together AI | JobStash