Senior Staff Engineer ML Ops
Shield AI is a U.S. defense-technology company developing mission-autonomy software and autonomous aircraft for military and allied operations.
Funding history
Investors
About Shield AI
Founded in 2015, Shield AI builds Hivemind autonomy software and V-BAT and X-BAT aircraft for operations in contested, GPS- and communications-denied environments. Its current site also presents Aechelon synthetic-reality simulation and Vision Systems detection and tracking products.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead the architecture and implementation of an AI platform for distributed training, simulation, evaluation, and deployment. You will build GPU infrastructure and model-lifecycle capabilities, create self-service workflows, package deployments with infrastructure as code, and work with researchers to scale AI workflows.
Requirements
- Experience building Kubernetes-native AI or MLOps platforms for distributed machine-learning workloads
- Understanding of PyTorch, Hugging Face Transformers, and distributed training
- Experience operating GPU-accelerated infrastructure and distributed training systems
- Understanding of Kubernetes, Linux, networking, security, storage, and distributed systems
- Experience with GPU scheduling and large-scale AI workloads
- Experience using Terraform and Helm for cloud-native infrastructure
- Software engineering skills in Python and Golang
- Experience translating ML research workflows into scalable platform capabilities
Responsibilities
- Lead the design and implementation of a Kubernetes-native AI platform
- Partner with ML researchers to support training workflows and AI frameworks
- Design self-service AI development workflows
- Build infrastructure for distributed training, simulation, inference, and reinforcement learning
- Design and optimize GPU infrastructure across cloud and on-premises environments
- Build dataset, experiment, artifact, model-versioning, evaluation, deployment, and monitoring capabilities
- Develop repeatable deployment and lifecycle management solutions using infrastructure as code
- Evaluate AI infrastructure technologies and establish architectural patterns
- Collaborate with researchers, autonomy teams, infrastructure engineers, and product teams
Benefits
- Bonus
- Benefits
- Equity
- Temporary benefits package after 60 days of employment
