Senior Staff Engineer ML Ops

Shield AI is a U.S. defense-technology company developing mission-autonomy software and autonomous aircraft for military and allied operations.

San Diego, United States
About Shield AI

Founded in 2015, Shield AI builds Hivemind autonomy software and V-BAT and X-BAT aircraft for operations in contested, GPS- and communications-denied environments. Its current site also presents Aechelon synthetic-reality simulation and Vision Systems detection and tracking products.

View jobs by Shield AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design and implement an AI platform for distributed training, simulation, evaluation, and deployment. You will enable researcher workflows, build GPU infrastructure and data-model lifecycle capabilities, develop repeatable deployments, evaluate platform technologies, and collaborate across AI and infrastructure disciplines.

Requirements

  • Experience building Kubernetes-native AI or MLOps platforms for distributed machine-learning workloads.
  • Understanding of AI training frameworks including PyTorch, Hugging Face Transformers, and distributed training techniques.
  • Experience operating GPU-accelerated infrastructure and distributed training systems.
  • Knowledge of Kubernetes, Linux, networking, security, storage, and distributed systems.
  • Experience with GPU scheduling and large-scale AI workloads.
  • Experience packaging and deploying cloud-native infrastructure with Terraform and Helm.
  • Software engineering skills in Python and Golang.
  • Experience translating ML research workflows into scalable platform capabilities.

Responsibilities

  • Lead the design and implementation of a Kubernetes-native platform for AI development, distributed training, simulation, evaluation, and deployment.
  • Partner with ML researchers to support training workflows, AI frameworks, foundation-model development, reinforcement learning, and distributed training.
  • Design self-service AI development workflows from local experimentation to distributed execution.
  • Build infrastructure for distributed training, simulation, inference, and reinforcement-learning workloads.
  • Design and optimize shared GPU infrastructure across cloud and on-premises environments.
  • Build capabilities for dataset management, experiment tracking, artifact management, model versioning, evaluation, deployment, monitoring, and continuous model improvement.
  • Develop repeatable deployment and lifecycle-management solutions using infrastructure as code.
  • Evaluate AI infrastructure technologies and establish scalable platform architecture patterns.
  • Collaborate with AI researchers, autonomy teams, infrastructure engineers, and product teams.

Benefits

  • Bonus
  • Equity
Senior Staff Engineer ML Ops at Shield AI | JobStash