Member of Technical Staff, Inference & Serving

AI research and product company building diffusion-based language models for production applications.

Palo Alto, United States

Funding history

About Inception

Inception develops and deploys the Mercury family of diffusion LLMs, which generate and refine output in parallel to target lower-latency, lower-cost production AI workloads.

View jobs by Inception

Skills

About the Role

You will build low-latency model serving systems and distributed inference infrastructure. You will manage traffic, autoscaling, versioning, deployments, and observability, while converting model advances into production serving improvements.

Requirements

  • Machine learning serving
  • SGLang
  • vLLM
  • Triton Inference Server
  • TensorRT-LLM
  • PyTorch
  • TensorFlow
  • High-performance computing
  • GPU programming
  • CUDA
  • Docker
  • Kubernetes
  • CI/CD
  • Performance optimization
  • Profiling

Responsibilities

  • Build and optimize low-latency model serving systems
  • Extend distributed inference and serving orchestration frameworks
  • Manage load balancing autoscaling and traffic routing
  • Build model versioning canary deployment and zero-downtime rollout systems
  • Develop monitoring alerting and observability tooling
  • Translate model advances into production serving improvements

Benefits

  • Equity
  • Flexible vacation and paid time off
  • Health insurance
  • Dental insurance
  • Vision insurance
  • 401k match
  • Catered meals
  • Commuter subsidies
Member of Technical Staff, Inference & Serving at Inception | JobStash