Site Reliability Engineer

Cognition is an applied AI company that operates Devin, an autonomous software-engineering agent.

San Francisco, United States
About Cognition

Cognition builds AI agents and models for software engineering. Its flagship product, Devin, can plan, write, test, and ship code in a customer’s existing codebase and tools.

View jobs by Cognition

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own production reliability and platform engineering. You will define service-level objectives and indicators, build monitoring and alerting, lead incident response and postmortems, manage CI/CD and cloud infrastructure through code, plan capacity, improve performance, reduce toil, and strengthen reliability practices.

Requirements

  • Production systems
  • SLOs
  • Error budgets
  • On-call rotations
  • Incident command
  • Software engineering
  • AWS, GCP, or Azure
  • Kubernetes
  • Infrastructure as code
  • Terraform
  • CI/CD
  • Deployment infrastructure
  • Observability
  • Automation
  • Incident detection, triage, mitigation, resolution, and postmortem
  • Developer-facing product or platform experience

Responsibilities

  • Define and own SLOs, SLIs, and error budgets
  • Build monitoring, alerting, and observability systems
  • Lead incident response and run blameless postmortems
  • Build runbooks and on-call tooling
  • Own CI/CD pipelines, deployment infrastructure, and developer tooling
  • Manage cloud infrastructure through code
  • Plan capacity and improve system performance
  • Reduce toil through automation
  • Identify and remediate security and reliability issues

Benefits

  • Fully paid medical, dental, and vision coverage for you and your dependents
  • 401(k) company match
  • Private chef
  • Cozy slippers
  • Endless snacks
  • Early-stage equity
Site Reliability Engineer at Cognition | JobStash