Site Reliability Engineer

Runpod is an AI developer cloud providing GPU compute, serverless inference, Pods, and clusters for building, training, fine-tuning, deploying, and scaling AI workloads.

Recently fundedCompany intelligence
San Francisco, United States
About Runpod

Runpod Inc. operates a globally distributed GPU cloud platform for AI developers, offering on-demand GPU infrastructure and serverless services across the AI development lifecycle.

View jobs by Runpod

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will define reliability indicators and objectives, lead incident response, improve monitoring and alerting, automate operational work, harden deployments, and partner with engineering teams to improve resilience, scalability, and production readiness.

Requirements

  • 5+ years of SRE, reliability engineering, or production engineering experience
  • Linux systems expertise
  • Networking expertise
  • Containerized production systems experience
  • Distributed systems and failure-mode knowledge
  • SLI and SLO experience
  • Incident response and postmortem leadership
  • Scripting or programming skills
  • Monitoring and alerting systems experience
  • Written communication
  • Successful completion of a background check

Responsibilities

  • Define and implement SLIs and SLOs for critical services
  • Lead incident response and cross-team mitigation
  • Conduct blameless postmortems and ensure corrective actions
  • Perform production-readiness reviews
  • Identify systemic risks and drive preventative improvements
  • Design monitoring, alerting, and dashboards
  • Improve alert signal-to-noise ratio
  • Build reliability tracking and reporting tools
  • Improve GPU and distributed-systems visibility
  • Automate recurring operational workflows
  • Build tools and scripts to eliminate manual processes
  • Improve deployment safety, CI/CD reliability, and releases
  • Guide teams on fault tolerance, scalability, and failure handling

Benefits

  • Equity through stock options
  • Medical, dental, and vision plans
  • Flexible PTO
  • Remote-first work