Site Reliability Engineer

Gamma is an AI design platform for creating and sharing presentations, websites, documents, social posts, and graphics.

San Francisco, United States
About Gamma

Gamma Tech, Inc. provides an AI-native content-creation product that turns prompts, outlines, and imported material into designed, shareable presentations, documents, websites, graphics, and social content. It also offers a developer API for programmatic generation.

View jobs by Gamma

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own the reliability, availability, and performance of production systems. You will build observability and deployment automation, lead incident response and post-mortems, drive systemic fixes, partner on architecture and SLO design, and optimize compute, networking, databases, and managed services.

Requirements

  • Site reliability engineering
  • DevOps
  • Systems engineering
  • AWS
  • Python
  • Go
  • TypeScript
  • Node.js
  • Infrastructure as code
  • Terraform
  • CloudFormation
  • Observability
  • Networking
  • Distributed system
  • Containerization
  • Docker
  • Kubernetes
  • Database performance
  • Incident management
  • Debugging

Responsibilities

  • Own production-system reliability, availability, and performance across AWS infrastructure
  • Build metrics, logging, tracing, and alerting infrastructure
  • Design and ship automation that reduces toil and improves deployment safety
  • Lead incident response and blameless post-mortems
  • Implement systemic fixes following incidents
  • Partner on architecture reviews, SLO design, SLI design, and reliability practices
  • Manage and optimize compute, networking, databases, and managed services

Benefits

  • Equity
Site Reliability Engineer at Gamma | JobStash