Member of Technical Staff Site Reliability Engineer
Vapi is a developer platform for building, deploying, and improving voice AI agents.
San Francisco, United States
Funding history
About Vapi
Vapi provides voice-agent infrastructure and orchestration across transcription, language models, and text-to-speech, including telephony, real-time monitoring, scaling, and conversation-flow tooling.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will join the on-call rotation, define SLOs and error budgets, lead incident command and postmortems, and improve platform reliability. You will build services for capacity forecasting, auto-remediation, and on-call tooling, while tuning autoscaling and running production load tests.
Requirements
- Experience running incident command and postmortem discipline on an on-call rotation
- Experience operating SLOs and error budgets with Chronosphere, Prometheus, Grafana, or Datadog
- Experience with capacity planning and production load testing
- Fluency in Kubernetes production operations, including HPA/VPA tuning, PodDisruptionBudgets, and graceful shutdown
- Knowledge of backpressure and autoscaling patterns, including KEDA and custom metrics scaling
Responsibilities
- Join the on-call rotation
- Define SLOs for the call-completion path
- Establish error budgets and SLO-based alerting
- Run load tests and audit provider rate limits and concurrency
- Tune autoscaling
- Build platform services for capacity forecasting, auto-remediation, or on-call tooling
- Own the postmortem process
- Improve p99 call completion and MTTR
Benefits
- Equity ownership
- Medical, dental, and vision coverage
- Quarterly off-sites
- Flexible time off
- Catered meals
- Transportation
- Gym access
