Senior Staff Lead Site Reliability Engineer

Shield AI is a U.S. defense-technology company developing mission-autonomy software and autonomous aircraft for military and allied operations.

San Diego, United States
About Shield AI

Founded in 2015, Shield AI builds Hivemind autonomy software and V-BAT and X-BAT aircraft for operations in contested, GPS- and communications-denied environments. Its current site also presents Aechelon synthetic-reality simulation and Vision Systems detection and tracking products.

View jobs by Shield AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will establish and mature reliability practices for cloud infrastructure and platform services. You will define reliability targets, improve observability, investigate complex failures, lead incident response, automate operational work, and turn incident lessons into engineering improvements. You will mentor engineers and guide the SRE roadmap.

Requirements

  • 7+ years of experience in SRE, software engineering, infrastructure engineering, or related fields
  • Experience operating production services with availability and reliability requirements
  • Experience implementing SLIs, SLOs, monitoring, alerting, and incident-response practices
  • Experience designing and operating infrastructure in AWS or another major cloud environment
  • Experience with infrastructure as code and automated infrastructure provisioning
  • Experience supporting containerized applications and distributed systems
  • Experience developing operational tooling or automation using Python, Go, or a similar language
  • Ability to diagnose complex failures across applications, infrastructure, networking, and dependent services
  • Experience leading incident response and root-cause analysis across engineering teams
  • Experience leading and executing a technical vision over multi-quarter timelines

Responsibilities

  • Define and implement SLIs, SLOs, and service-reliability measures
  • Build and improve monitoring, alerting, logging, and tracing
  • Lead responses to complex incidents and drive root-cause analysis
  • Identify recurring failure modes and help eliminate them
  • Improve resilience through automation, testing, capacity planning, and recovery
  • Develop tooling that reduces manual operational work
  • Incorporate reliability requirements into system design
  • Establish incident-response practices
  • Mentor engineers in reliability and operational practices
  • Define and manage the SRE roadmap

Benefits

  • Bonus
  • Equity
  • Benefits
Senior Staff Lead Site Reliability Engineer at Shield AI | JobStash