Site Reliability Engineer

Baseten is an AI inference platform for deploying, optimizing, and scaling custom, open-source, and fine-tuned models in production.

San Francisco, United States
About Baseten

Baseten provides model runtimes, inference infrastructure, developer workflows, and deployment options including managed cloud, self-hosted, and hybrid environments.

View jobs by Baseten

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own the reliability of multi-cloud Kubernetes infrastructure, respond to incidents, and track remediation. You will build observability systems, develop runbooks and self-healing automations, diagnose runtime issues, and define service-level objectives and indicators.

Requirements

  • Kubernetes
  • Cloud infrastructure
  • Observability
  • VictoriaMetrics
  • Prometheus
  • Loki
  • ELK
  • Grafana
  • Alerting
  • Infrastructure as code
  • Terraform
  • Helm
  • GitOps
  • Flux CD
  • ArgoCD
  • Runbook
  • Incident response
  • Post-mortem analysis
  • incident.io

Responsibilities

  • Own the reliability of multi-cloud Kubernetes infrastructure
  • Lead incident response, post-mortems, and remediation tracking
  • Build and maintain metrics, logging, dashboards, and alerting infrastructure as code
  • Author and improve runbooks for recurring failures
  • Convert recurring failures into automated mitigations and self-healing automations
  • Diagnose runtime issues involving latency, memory, GPU utilization, concurrency, and model lifecycle management
  • Define and instrument SLOs and SLIs

Benefits

  • Meaningful equity
  • Medical, dental, and vision insurance for U.S. employees and dependents
  • Flexible PTO
  • Company-wide winter break
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Company-facilitated 401(k) for U.S. employees
Site Reliability Engineer at Baseten | JobStash