Data Center Operations Systems Engineer III
47 minutes agoLeadSalary: 115K - 153KVernon, CA - Data CenterOnsiteFull TimeEngineeringJobs by Lambda
LambdaVisit Lambda website
Lambda is an AI infrastructure company providing GPU supercomputers and cloud capacity for AI training and inference.
San Francisco, United States
About Lambda
Lambda, Inc. builds and operates AI-focused compute infrastructure, including single-tenant Superclusters, deployable 1-Click Clusters, and on-demand GPU Instances for researchers, enterprises, and hyperscalers.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will deploy and troubleshoot GPU, networking, server, and data center infrastructure. You will maintain DCIM documentation, manage parts and RMA processes, resolve complex hardware incidents, develop maintenance procedures, align operational capabilities with requirements, and train junior staff.
Requirements
- Experience with data center critical infrastructure systems, including power distribution, airflow management, environmental monitoring, capacity planning, DCIM software, structured cabling, and cable management
- Carrier DIA circuit testing and turn-ups
- Fiber testing and troubleshooting
- Cable optics
- Single-phase and three-phase power theory
- PDU balancing
- Cable media
- Cold aisle and hot aisle containment
- Server hardware and boot processes
- Maintenance MOPs
- Willingness to train junior staff
- Willingness to travel for new data center bring-up
- 5+ years of experience with data center critical infrastructure systems
- Network topology and configuration
- 400Gb InfiniBand architectures
- DDP or SCM cluster storage systems
- 5+ years of ticketing-system experience
- JIRA
- Zendesk
- Advanced Linux administration
- High-performance computing GPU systems
- Nvidia NVL72
Responsibilities
- Rack, label, cable, and configure server, storage, and network infrastructure
- Troubleshoot GPU, networking, hardware, and software issues
- Document and update data center layouts and network topology in DCIM software
- Coordinate timely system deployments and project plans with supply chain and manufacturing teams
- Manage parts depot inventory and track equipment through deployment
- Resolve, report on, and share solutions for complex data center hardware incidents
- Return faulty parts and order replacements with the RMA team
- Follow installation standards for placement, labeling, and cabling
- Align operational capabilities with company goals
- Translate business priorities into technical and operational requirements
- Support cross-functional infrastructure projects
- Train junior staff on best practices
Benefits
- Equity compensation
- Health, dental, and vision coverage for employees and dependents
- Wellness and commuter stipends for select roles
- 401k plan with 2% company match for USA employees
- Flexible paid time off
