Member of Technical Staff Site Reliability Engineer

Vapi is a developer platform for building, deploying, and improving voice AI agents.

San Francisco, United States
About Vapi

Vapi provides voice-agent infrastructure and orchestration across transcription, language models, and text-to-speech, including telephony, real-time monitoring, scaling, and conversation-flow tooling.

View jobs by Vapi

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will join the on-call rotation, define SLOs and error budgets, lead incident command and postmortems, and improve platform reliability. You will build services for capacity forecasting, auto-remediation, and on-call tooling, while tuning autoscaling and running production load tests.

Requirements

  • Experience running incident command and postmortem discipline on an on-call rotation
  • Experience operating SLOs and error budgets with Chronosphere, Prometheus, Grafana, or Datadog
  • Experience with capacity planning and production load testing
  • Fluency in Kubernetes production operations, including HPA/VPA tuning, PodDisruptionBudgets, and graceful shutdown
  • Knowledge of backpressure and autoscaling patterns, including KEDA and custom metrics scaling

Responsibilities

  • Join the on-call rotation
  • Define SLOs for the call-completion path
  • Establish error budgets and SLO-based alerting
  • Run load tests and audit provider rate limits and concurrency
  • Tune autoscaling
  • Build platform services for capacity forecasting, auto-remediation, or on-call tooling
  • Own the postmortem process
  • Improve p99 call completion and MTTR

Benefits

  • Equity ownership
  • Medical, dental, and vision coverage
  • Quarterly off-sites
  • Flexible time off
  • Catered meals
  • Transportation
  • Gym access
Member of Technical Staff Site Reliability Engineer at Vapi | JobStash