Engineering Manager Serving API
Positron AIVisit Positron AI website
Positron AI builds hardware and software for energy-efficient generative-AI and Transformer-model inference.
Reno, Nevada, United States
Funding history
Investors
About Positron AI
Reno-based AI-infrastructure company whose currently shipping Atlas inference server is complemented by planned Asimov custom accelerator silicon and Titan inference systems.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead and grow engineers responsible for API serving, tokenization, structured generation, and model integration. You will set the serving-layer roadmap, ensure OpenAI-compatible API behavior, lead support for reasoning models and speculative decoding, and establish release gates for models and features.
Requirements
- Experience managing and growing engineering teams responsible for LLM inference, model serving, ML systems, or a related domain
- Systems expertise in C++
- Working proficiency in Python
- Ability to review performance- and correctness-critical code
- Experience translating ambiguous product demands into roadmaps, ownership, and measurable outcomes
- Ability to recruit, coach, and retain engineers
- Written and verbal communication, prioritization, and tradeoff management
- Hands-on leadership experience close to architecture and code
- Authorization to work in the United States
Responsibilities
- Lead, coach, and grow engineers across API serving, tokenization, structured generation, and model integration
- Establish a roadmap for the API surface, chat templates, tool calling, reasoning, budgets, and new-model support
- Ensure OpenAI-compatible behavior, streaming, usage accounting, and error handling
- Lead reasoning-model support, reasoning-effort controls, token limits, and per-request accounting
- Partner on speculative decoding, draft-model integration, acceptance metrics, and policies
- Lead vision-language model support and SGLang interoperability
- Keep serving-layer overhead off the critical path for time-to-first-token and streaming latency
- Keep the API layer independent of hardware topology
- Define ownership for the shared host-level load-balancing layer
- Build API conformance, tool-call, reasoning, and regression release gates
- Collaborate to make serving capabilities supportable production endpoints
- Hire engineers, develop emerging leaders, and establish scalable ownership boundaries
Benefits
- Fully company-paid medical, dental, and vision insurance for employees and dependents
- Company-paid life and disability coverage
- Voluntary supplemental hospital, critical illness, and accident coverage
- Unlimited paid time off
- 13 paid company holidays
- Remote-first work
- Company-provided computer and home office setup
- Equity
- 401(k) with company matching from day one
