Technical Support Engineer Inference
Together AIVisit Together AI website
Together AI operates an AI-native cloud platform for open and custom AI models.
San Francisco, United States
About Together AI
Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will support customers using GPU clusters, inference, and fine-tuning services. You will maintain endpoint health, investigate incidents, validate migrations, communicate customer impacts, contribute infrastructure changes through pull requests, diagnose bugs, and document troubleshooting procedures. You will work a four-day India daytime weekend shift after ramping up.
Requirements
- 6+ years of customer-facing technical, SRE, DevOps, or infrastructure engineering experience
- At least 1 year of AI-service support experience
- Knowledge of AI, machine learning, GPU technologies, and HPC environments
- Production experience with Kubernetes, SLURM, Ansible, networking, NFS storage, and containers
- Experience operating Vast or Weka storage systems
- Ability to diagnose network-layer issues and read traces
- Knowledge of Python, TypeScript, or JavaScript
- Experience with curl, API testing, observability, Prometheus, and Grafana
- Knowledge of REST APIs and HTTP semantics
- Experience with LLM inference frameworks and LoRA fine-tuning
- Experience with infrastructure as code and Git workflows
- GPU cluster management experience
- AWS, GCP, or Azure experience
Responsibilities
- Resolve customer technical challenges involving GPU clusters, inference, and fine-tuning services
- Maintain the health, stability, and performance of customer inference endpoints
- Validate system health and traffic routing during hardware and platform migrations
- Monitor dashboards and escalate anomalies with data-backed analysis
- Communicate customer impacts during incidents and degradations
- Contribute infrastructure changes through pull requests
- Flag engine-level bugs with logs and reproduction steps
- Identify support patterns and inform roadmap decisions
- Maintain system and troubleshooting documentation
- Provide support coverage during holidays, nights, and weekends as needed
Benefits
- Startup equity
- Health insurance
- Remote-work flexibility
