Member of Technical Staff Backend LLM Applications

AI research and product company building diffusion-based language models for production applications.

Palo Alto, United States

Funding history

About Inception

Inception develops and deploys the Mercury family of diffusion LLMs, which generate and refine output in parallel to target lower-latency, lower-cost production AI workloads.

View jobs by Inception

Skills

About the Role

You will design, build, and operate scalable backend services and model-serving infrastructure for diffusion LLMs. You will optimize latency, throughput, cost, and reliability; manage traffic; support deployments; and develop observability tooling.

Requirements

  • 5+ years of experience building production backend systems
  • Proficiency in Python, including asynchronous programming and concurrent systems
  • Understanding of distributed systems, networking, and load balancing at scale
  • Familiarity with Kubernetes, CI/CD pipelines, and AWS and/or Azure

Responsibilities

  • Design, build, and operate scalable backend services and model-serving infrastructure
  • Implement and manage load balancing, autoscaling, and traffic routing
  • Build systems for model versioning, canary deployments, and zero-downtime rollouts
  • Develop monitoring, alerting, and observability tooling
  • Benchmark serving frameworks and hardware configurations

Benefits

  • Equity
  • Flexible vacation and paid time off
  • Health, dental, and vision insurance
  • 401k match
  • Catered meals
  • Commuter subsidies
Member of Technical Staff Backend LLM Applications at Inception | JobStash