Staff Engineer AI Operations and Governance Workplace AI
Shield AI is a U.S. defense-technology company developing mission-autonomy software and autonomous aircraft for military and allied operations.
Funding history
Investors
About Shield AI
Founded in 2015, Shield AI builds Hivemind autonomy software and V-BAT and X-BAT aircraft for operations in contested, GPS- and communications-denied environments. Its current site also presents Aechelon synthetic-reality simulation and Vision Systems detection and tracking products.
Skills
About the Role
You will configure and operate AI platforms, models, routes, guardrails, policies, and integrations. You will maintain observability, secrets, access controls, production changes, technical evaluations, incident response, and operational documentation while translating technical signals into governance and risk views.
Requirements
- 4–7+ years of experience in platform engineering, DevOps, SRE, ML or AI operations, or technical SaaS operations.
- Fluency in APIs, integrations, infrastructure-as-code concepts, code, JSON, YAML, and automation.
- Experience with monitoring and observability tools, including logs, metrics, and alerts.
- Experience with AI or automation platforms such as LLM providers, AI productivity tools, or workflow engines.
- Experience with secure secrets management and access control practices.
- Ability to document technical work and decisions for technical and non-technical audiences.
Responsibilities
- Implement and maintain AI platform configurations, models, routes, guardrails, tenants, policies, role mappings, and prompt libraries.
- Design, configure, and maintain connectors and extensions to SaaS systems, data sources, and workflow tools.
- Set up and maintain logging, metrics, and alerts for AI workflows and tools.
- Implement secure storage and rotation for API keys, tokens, and credentials.
- Maintain access-control configurations and support security and IT reviews.
- Execute production model, policy, prompt, version, and rollout changes with change records and rollback paths.
- Run experiments and benchmarks for models, tools, and configurations.
- Translate logs, metrics, and incidents into risk, reliability, and compliance views.
- Co-author operational playbooks, runbooks, and training materials.
- Triage, investigate, mitigate, resolve, and document incidents and changes.
Benefits
- Bonus
- Equity
