Forward Deployed Engineer Inference and Post Training

Together AI operates an AI-native cloud platform for open and custom AI models.

San Francisco, United States
About Together AI

Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.

View jobs by Together AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will optimize inference engines and configurations for customer workloads, tune performance for throughput and latency targets, and guide post-training and fine-tuning pipelines into production. You will support strategic accounts, lead technical onboarding, and share field insights to improve product and model roadmaps.

Requirements

  • 5+ years in a technical role focused on inference systems, open-source LLM deployment, or post-training workflows
  • Expert hands-on experience with inference engines including vLLM, TensorRT-LLM, or SGLang
  • Knowledge of KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization
  • Hands-on experience with LoRA, SFT, DPO, RLHF, and GRPO pipelines
  • Knowledge of open-source models and model selection
  • Strong Python skills and comfort in production environments

Responsibilities

  • Select, configure, and optimize inference engines for hardware, model architectures, and workload profiles
  • Develop configuration updates and tune KV cache, speculative decoding, tensor parallelism, and quantization
  • Drive RL training runs and guide LoRA, SFT, DPO, RLHF, and GRPO pipelines into production
  • Serve as the primary technical contact for strategic accounts and optimize endpoint configurations
  • Align customers on inference and post-training configurations during onboarding
  • Surface field insights and contribute to product, model, feature, and research adoption

Benefits

  • Startup equity
  • Health insurance
  • Remote-work flexibility
Forward Deployed Engineer Inference and Post Training at Together AI | JobStash