Staff Engineer Platform
Shield AI is a U.S. defense-technology company developing mission-autonomy software and autonomous aircraft for military and allied operations.
Funding history
Investors
About Shield AI
Founded in 2015, Shield AI builds Hivemind autonomy software and V-BAT and X-BAT aircraft for operations in contested, GPS- and communications-denied environments. Its current site also presents Aechelon synthetic-reality simulation and Vision Systems detection and tracking products.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and operate Kubernetes-native services, controllers, operators, deployment patterns, and runtime integrations for distributed workloads. You will develop orchestration and data-processing capabilities, establish platform standards and reference architectures, and improve observability, reliability, and operational diagnostics. You will partner with downstream teams to turn recurring infrastructure needs into reusable platform components.
Requirements
- Experience designing and operating production distributed systems, cloud-native platforms, backend infrastructure, or data-intensive services
- Software engineering skills delivering production systems in Go and Python
- Understanding of distributed-systems fundamentals including failure handling, idempotency, consistency tradeoffs, retries, ordering, delivery semantics, backpressure, partitioning, state management, and fault tolerance
- Experience with workflow orchestration, distributed job execution, asynchronous processing, event-driven systems, or long-running service workflows
- Ability to define architecture and technical standards while remaining hands-on in implementation, troubleshooting, performance analysis, and reliability improvement
- Experience creating reusable, documented platform capabilities across multiple teams
- Technical communication skills
Responsibilities
- Build and operate Kubernetes-native platform services, controllers, operators, deployment patterns, and runtime integrations
- Design reusable primitives for authoring, scheduling, and scaling pipeline work
- Develop data storage, ingestion, validation, transformation, and governance capabilities
- Develop extensible platform components for authentication, authorization, observability, networking, routing, and secret management
- Create deployment patterns, operating profiles, capacity guidance, benchmarks, reliability practices, and distribution approaches
- Establish metrics, logs, traces, dashboards, alerting, service-level objectives, diagnostics, and runbooks
- Partner with downstream teams to turn recurring distributed-systems problems into reusable platform capabilities
Benefits
- Bonus
- Equity
