Staff Engineer Platform

Shield AI is a U.S. defense-technology company developing mission-autonomy software and autonomous aircraft for military and allied operations.

San Diego, United States
About Shield AI

Founded in 2015, Shield AI builds Hivemind autonomy software and V-BAT and X-BAT aircraft for operations in contested, GPS- and communications-denied environments. Its current site also presents Aechelon synthetic-reality simulation and Vision Systems detection and tracking products.

View jobs by Shield AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build and operate Kubernetes-native services, controllers, operators, deployment patterns, and runtime integrations for distributed workloads. You will develop orchestration and data-processing capabilities, establish platform standards and reference architectures, and improve observability, reliability, and operational diagnostics. You will partner with downstream teams to turn recurring infrastructure needs into reusable platform components.

Requirements

  • Experience designing and operating production distributed systems, cloud-native platforms, backend infrastructure, or data-intensive services
  • Software engineering skills delivering production systems in Go and Python
  • Understanding of distributed-systems fundamentals including failure handling, idempotency, consistency tradeoffs, retries, ordering, delivery semantics, backpressure, partitioning, state management, and fault tolerance
  • Experience with workflow orchestration, distributed job execution, asynchronous processing, event-driven systems, or long-running service workflows
  • Ability to define architecture and technical standards while remaining hands-on in implementation, troubleshooting, performance analysis, and reliability improvement
  • Experience creating reusable, documented platform capabilities across multiple teams
  • Technical communication skills

Responsibilities

  • Build and operate Kubernetes-native platform services, controllers, operators, deployment patterns, and runtime integrations
  • Design reusable primitives for authoring, scheduling, and scaling pipeline work
  • Develop data storage, ingestion, validation, transformation, and governance capabilities
  • Develop extensible platform components for authentication, authorization, observability, networking, routing, and secret management
  • Create deployment patterns, operating profiles, capacity guidance, benchmarks, reliability practices, and distribution approaches
  • Establish metrics, logs, traces, dashboards, alerting, service-level objectives, diagnostics, and runbooks
  • Partner with downstream teams to turn recurring distributed-systems problems into reusable platform capabilities

Benefits

  • Bonus
  • Equity