AI Operations Engineer IT Internal Infrastructure

Lambda is an AI infrastructure company providing GPU supercomputers and cloud capacity for AI training and inference.

San Francisco, United States
About Lambda

Lambda, Inc. builds and operates AI-focused compute infrastructure, including single-tenant Superclusters, deployable 1-Click Clusters, and on-demand GPU Instances for researchers, enterprises, and hyperscalers.

View jobs by Lambda

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design and deliver software and services that improve the availability, scalability, reliability, and efficiency of internal IT platforms. You will automate operational responses, shape distributed-system designs, perform capacity and performance analysis, and document the systems you support.

Requirements

  • Experience with system design, performance, scalability, and cloud infrastructure platforms
  • Knowledge of AWS, GCP, or Azure
  • Knowledge of configuration management tools such as Chef, Ansible, Terraform, and GitHub Actions
  • Programming skills in Python or Go
  • Ability to document issues and solutions and collaborate asynchronously

Responsibilities

  • Design and deliver software and services for internal IT systems
  • Automate responses to non-exceptional operational events
  • Solve problems in mission-critical services and prevent recurrence
  • Influence distributed-system designs, architectures, standards, and methods
  • Perform capacity planning, demand forecasting, performance analysis, and system tuning
  • Produce documentation for supported systems

Benefits

  • Equity compensation
  • Health, dental, and vision coverage for employees and dependents
  • Wellness stipend for select roles
  • Commuter stipend for select roles
  • 401k plan with 2% company match for US employees
  • Flexible paid time off
AI Operations Engineer IT Internal Infrastructure at Lambda | JobStash