Senior Staff Lead Site Reliability Engineer
Shield AI is a U.S. defense-technology company developing mission-autonomy software and autonomous aircraft for military and allied operations.
Funding history
Investors
About Shield AI
Founded in 2015, Shield AI builds Hivemind autonomy software and V-BAT and X-BAT aircraft for operations in contested, GPS- and communications-denied environments. Its current site also presents Aechelon synthetic-reality simulation and Vision Systems detection and tracking products.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will establish and mature reliability practices for cloud infrastructure and platform services. You will define service objectives, improve observability, investigate failures, lead incident response, automate operational work and recovery, improve resilience, partner on reliable system design, mentor engineers, and manage the reliability roadmap.
Requirements
- 7+ years of SRE, software engineering, infrastructure engineering, or related experience
- Experience operating production services with availability and reliability requirements
- Experience with SLIs, SLOs, monitoring, alerting, and incident response
- Experience designing and operating infrastructure in AWS or another major cloud environment
- Infrastructure-as-code and automated infrastructure-provisioning experience
- Experience supporting containerized applications and distributed systems
- Experience using Python, Go, or a similar language for operational tooling or automation
- Ability to diagnose failures across applications, infrastructure, networking, and dependent services
- Experience leading incident response and root-cause analysis
- Experience leading a technical vision over multi-quarter timelines
Responsibilities
- Define and implement SLIs, SLOs, and service-reliability measures
- Build and improve monitoring, alerting, logging, and tracing
- Lead technical response to complex incidents and root-cause analysis
- Identify recurring failure modes and eliminate them with engineering teams
- Improve resilience through automation, testing, capacity planning, and failure recovery
- Develop tooling and automation that reduces manual operational work
- Incorporate reliability requirements into system design
- Establish incident-response practices
- Mentor product engineers, SREs, and cloud engineers
- Define and manage the SRE roadmap and distribute work across teammates
Benefits
- Bonus
- Benefits
- Equity
