Engineering Manager Production Platform and Orchestration
Positron AI builds hardware and software for energy-efficient generative-AI and Transformer-model inference.
Funding history
Investors
About Positron AI
Reno-based AI-infrastructure company whose currently shipping Atlas inference server is complemented by planned Asimov custom accelerator silicon and Titan inference systems.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead and grow engineers responsible for the production platform, fleet orchestration, reliability, and operational automation. You will set the roadmap for provisioning, deployment, observability, upgrades, capacity, and rollback; improve incident response and release automation; and partner on data-center and production-service operations.
Requirements
- Experience managing and growing engineering teams responsible for distributed systems, cloud infrastructure, production platforms, SRE, or a related domain
- Technical judgment across Linux systems, networking, orchestration, deployment systems, observability, and production reliability
- Experience translating ambiguous operational demands into roadmaps, ownership, and measurable outcomes
- Ability to recruit, coach, and retain engineers
- Ability to operate across software, hardware, data center, customer, and business boundaries
- Written and verbal communication skills, prioritization, and tradeoff management
- Hands-on leadership experience close to architecture and operations
- Authorization to work in the United States
Responsibilities
- Lead, coach, and grow engineers across production platform, fleet orchestration, reliability, and operational automation
- Hire engineers, develop emerging leaders, and establish scalable ownership boundaries
- Establish a roadmap for provisioning, deployment, environment lifecycle, fleet health, capacity, upgrades, and rollback
- Build orchestration and control-plane capabilities for inventory, placement, configuration, health management, and multi-system operations
- Define service-level objectives, operational metrics, alerting standards, and an on-call model
- Own incident, escalation, postmortem, corrective-action, launch-readiness, and reliability-review processes
- Increase automation for deployment, upgrades, remediation, diagnostics, capacity planning, and support workflows
- Partner on rack bring-up, networking, firmware, sparing, failure handling, and platform transitions
- Collaborate to make new capabilities supportable production services
Benefits
- Fully company-paid medical, dental, and vision insurance for employees and dependents
- Company-paid life and disability coverage
- Voluntary supplemental hospital, critical illness, and accident coverage
- Unlimited paid time off
- 13 paid company holidays
- Remote-first work
- Company-provided computer and home office setup
- Equity
- 401(k) with company matching from day one
