Senior Staff Lead Site Reliability Engineer
Shield AI is a U.S. defense-technology company developing mission-autonomy software and autonomous aircraft for military and allied operations.
Funding history
Investors
About Shield AI
Founded in 2015, Shield AI builds Hivemind autonomy software and V-BAT and X-BAT aircraft for operations in contested, GPS- and communications-denied environments. Its current site also presents Aechelon synthetic-reality simulation and Vision Systems detection and tracking products.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will establish and mature reliability practices for cloud infrastructure and platform services. You will define reliability targets, improve observability, investigate complex failures, lead incident response, automate operational work, and turn incident lessons into engineering improvements. You will mentor engineers and guide the SRE roadmap.
Requirements
- 7+ years of experience in SRE, software engineering, infrastructure engineering, or related fields
- Experience operating production services with availability and reliability requirements
- Experience implementing SLIs, SLOs, monitoring, alerting, and incident-response practices
- Experience designing and operating infrastructure in AWS or another major cloud environment
- Experience with infrastructure as code and automated infrastructure provisioning
- Experience supporting containerized applications and distributed systems
- Experience developing operational tooling or automation using Python, Go, or a similar language
- Ability to diagnose complex failures across applications, infrastructure, networking, and dependent services
- Experience leading incident response and root-cause analysis across engineering teams
- Experience leading and executing a technical vision over multi-quarter timelines
Responsibilities
- Define and implement SLIs, SLOs, and service-reliability measures
- Build and improve monitoring, alerting, logging, and tracing
- Lead responses to complex incidents and drive root-cause analysis
- Identify recurring failure modes and help eliminate them
- Improve resilience through automation, testing, capacity planning, and recovery
- Develop tooling that reduces manual operational work
- Incorporate reliability requirements into system design
- Establish incident-response practices
- Mentor engineers in reliability and operational practices
- Define and manage the SRE roadmap
Benefits
- Bonus
- Equity
- Benefits
