Senior Data Center Operations Systems Engineer

Lambda is an AI infrastructure company providing GPU supercomputers and cloud capacity for AI training and inference.

San Francisco, United States
About Lambda

Lambda, Inc. builds and operates AI-focused compute infrastructure, including single-tenant Superclusters, deployable 1-Click Clusters, and on-demand GPU Instances for researchers, enterprises, and hyperscalers.

View jobs by Lambda

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will rack, cable, label, configure, and troubleshoot server, storage, and network infrastructure. You will maintain data center layouts and network topology in DCIM software, coordinate deployments and inventory, and support hardware incident resolution and RMA processes. You will follow installation standards, align operational requirements with cross-functional partners, train junior staff, and travel for new data center launches as needed.

Requirements

  • Experience with critical data center infrastructure systems
  • Knowledge of power distribution, airflow management, environmental monitoring, capacity planning, DCIM, structured cabling, and cable management
  • Familiarity with carrier DIA circuit testing, fiber testing, and troubleshooting
  • Knowledge of cable optics and cable media types
  • Understanding of single- and three-phase power, PDU balancing, and containment
  • Understanding of server hardware and boot processes
  • Ability to develop and improve maintenance procedures
  • Willingness to train junior staff
  • Willingness to travel

Responsibilities

  • Rack, label, cable, and configure server, storage, and network infrastructure
  • Troubleshoot hardware and software issues
  • Document data center layouts and network topology in DCIM software
  • Coordinate large-scale system deployments with supply chain and manufacturing teams
  • Manage parts inventory and equipment lifecycle tracking
  • Resolve and report complex data center hardware incidents
  • Coordinate faulty-part returns and replacement orders
  • Follow installation, labeling, and cabling standards
  • Train junior staff on best practices
  • Travel for new data center launches as needed

Benefits

  • Cash and equity compensation
  • Health, dental, and vision coverage for employees and dependents
  • Wellness and commuter stipends for select roles
  • 401k plan with 2% company match for USA employees
  • Flexible paid time off
Senior Data Center Operations Systems Engineer at Lambda | JobStash