Technical Support Engineer GPU Cluster
Together AIVisit Together AI website
Together AI operates an AI-native cloud platform for open and custom AI models.
San Francisco, United States
About Together AI
Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will provide technical support for customers using GPU clusters. You will diagnose complex issues, act as a product expert, coordinate with engineering and product partners, identify support trends, and maintain troubleshooting documentation. You will also provide support coverage outside standard business hours when required.
Requirements
- 3+ years of customer-facing technical experience
- At least 1 year of support experience in AI or mission-critical SaaS APIs
- Knowledge of AI, machine learning, GPU technologies, and HPC environments
- Familiarity with Kubernetes, SLURM, Ansible, network fabrics, NFS storage, containers, and scripting
- Knowledge of compute-cluster installation, configuration, administration, troubleshooting, and security
- Technical problem-solving and troubleshooting skills
- Cross-functional collaboration skills
- Communication skills for technical and non-technical stakeholders
Responsibilities
- Resolve customer technical challenges involving Kubernetes GPU clusters
- Serve as a product expert and escalate issues to engineering and product teams
- Collaborate with engineering, research, product, and sales partners on customer concerns
- Identify support-case patterns and inform roadmap decisions
- Maintain system configuration, procedure, troubleshooting, and FAQ documentation
- Provide support coverage during holidays, nights, and weekends as needed
Benefits
- Startup equity
- Health insurance
- Remote-work flexibility
