Senior Staff Engineer ML Ops

Shield AI is a U.S. defense-technology company developing mission-autonomy software and autonomous aircraft for military and allied operations.

San Diego, United States
About Shield AI

Founded in 2015, Shield AI builds Hivemind autonomy software and V-BAT and X-BAT aircraft for operations in contested, GPS- and communications-denied environments. Its current site also presents Aechelon synthetic-reality simulation and Vision Systems detection and tracking products.

View jobs by Shield AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead the architecture and implementation of an AI platform for distributed training, simulation, evaluation, and deployment. You will build GPU infrastructure and model-lifecycle capabilities, create self-service workflows, package deployments with infrastructure as code, and work with researchers to scale AI workflows.

Requirements

  • Experience building Kubernetes-native AI or MLOps platforms for distributed machine-learning workloads
  • Understanding of PyTorch, Hugging Face Transformers, and distributed training
  • Experience operating GPU-accelerated infrastructure and distributed training systems
  • Understanding of Kubernetes, Linux, networking, security, storage, and distributed systems
  • Experience with GPU scheduling and large-scale AI workloads
  • Experience using Terraform and Helm for cloud-native infrastructure
  • Software engineering skills in Python and Golang
  • Experience translating ML research workflows into scalable platform capabilities

Responsibilities

  • Lead the design and implementation of a Kubernetes-native AI platform
  • Partner with ML researchers to support training workflows and AI frameworks
  • Design self-service AI development workflows
  • Build infrastructure for distributed training, simulation, inference, and reinforcement learning
  • Design and optimize GPU infrastructure across cloud and on-premises environments
  • Build dataset, experiment, artifact, model-versioning, evaluation, deployment, and monitoring capabilities
  • Develop repeatable deployment and lifecycle management solutions using infrastructure as code
  • Evaluate AI infrastructure technologies and establish architectural patterns
  • Collaborate with researchers, autonomy teams, infrastructure engineers, and product teams

Benefits

  • Bonus
  • Benefits
  • Equity
  • Temporary benefits package after 60 days of employment
Senior Staff Engineer ML Ops at Shield AI | JobStash