Software Engineer, Site Reliability (SRE)

Enterprise conversational-AI platform that helps businesses build and operate customer-experience agents across channels.

San Francisco, United States
About Sierra

Sierra provides an AI platform for businesses to build, deploy, analyze, and optimize customer-facing agents across chat, SMS, WhatsApp, email, voice, and ChatGPT.

View jobs by Sierra

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own observability, including monitoring, alerting, logging, and tracing. You will design reliable AWS infrastructure with Terraform, improve LLM deployment reliability, advance CI/CD and incident management, and establish SRE practices.

Requirements

  • 5+ years of Site Reliability or Infrastructure engineering experience
  • Availability, scalability, and reliability design
  • Terraform
  • AWS
  • Container orchestration
  • Cloud networking
  • IAM
  • VPC architecture
  • Observability systems
  • Enterprise customer compliance and networking knowledge

Responsibilities

  • Own the observability stack
  • Design reliable and scalable systems with engineering partners
  • Implement AWS infrastructure using Terraform and DevOps tooling
  • Improve LLM deployment reliability and scalability
  • Improve deployment pipelines, CI/CD tooling, and incident management
  • Define SRE practices, tooling, and best practices

Benefits

  • Flexible unlimited paid time off
  • Medical, dental, and vision benefits for employees and families
  • Life insurance and disability benefits
  • Retirement plan dependent on country of employment
  • Parental leave
  • Fertility and family-building benefits through Carrot
  • Lunch, snacks, and coffee
  • Discretionary benefit stipend
  • Free alphorn lessons
  • Equity plan participation