Software Engineer, Site Reliability (SRE)
SierraVisit Sierra website
Enterprise conversational-AI platform that helps businesses build and operate customer-experience agents across channels.
San Francisco, United States
About Sierra
Sierra provides an AI platform for businesses to build, deploy, analyze, and optimize customer-facing agents across chat, SMS, WhatsApp, email, voice, and ChatGPT.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own observability, including monitoring, alerting, logging, and tracing. You will design reliable AWS infrastructure with Terraform, improve LLM deployment reliability, advance CI/CD and incident management, and establish SRE practices.
Requirements
- 5+ years of Site Reliability or Infrastructure engineering experience
- Availability, scalability, and reliability design
- Terraform
- AWS
- Container orchestration
- Cloud networking
- IAM
- VPC architecture
- Observability systems
- Enterprise customer compliance and networking knowledge
Responsibilities
- Own the observability stack
- Design reliable and scalable systems with engineering partners
- Implement AWS infrastructure using Terraform and DevOps tooling
- Improve LLM deployment reliability and scalability
- Improve deployment pipelines, CI/CD tooling, and incident management
- Define SRE practices, tooling, and best practices
Benefits
- Flexible unlimited paid time off
- Medical, dental, and vision benefits for employees and families
- Life insurance and disability benefits
- Retirement plan dependent on country of employment
- Parental leave
- Fertility and family-building benefits through Carrot
- Lunch, snacks, and coffee
- Discretionary benefit stipend
- Free alphorn lessons
- Equity plan participation
