Engineering Manager Fleet Engineering
Lambda is an AI infrastructure company providing GPU supercomputers and cloud capacity for AI training and inference.
About Lambda
Lambda, Inc. builds and operates AI-focused compute infrastructure, including single-tenant Superclusters, deployable 1-Click Clusters, and on-demand GPU Instances for researchers, enterprises, and hyperscalers.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead and grow engineers responsible for deploying and operating production systems infrastructure. You will align cross-functional stakeholders, improve tools and automation, manage priorities and staffing, communicate progress and risks, support incident management, and develop your team through feedback and regular one-to-ones.
Requirements
- 3+ years leading or managing engineers in AI/ML infrastructure or a large-scale compute environment
- Experience owning production systems with SLAs
- Linux troubleshooting across operating system, hardware, and networking layers
- Technical design leadership for medium-to-large efforts
- Project planning and timeline management
- Experience building high-performing teams
Responsibilities
- Lead and grow a distributed engineering team
- Deliver projects and deployments on time with cross-functional stakeholders
- Identify efficiency improvements in tools, processes, and automation
- Communicate project progress, risks, and outcomes to stakeholders
- Participate in qualification efforts for new production technologies
- Manage staff allocation, priorities, deadlines, and deliverables
- Conduct regular one-to-ones and provide constructive feedback
- Participate in incident management and review programs
Benefits
- Cash and equity compensation
- Health, dental, and vision coverage for employees and dependents
- Wellness and commuter stipends for select roles
- 401k plan with 2% company match for USA employees
- Flexible paid time off
