Member of Technical Staff GPU Infrastructure Engineer

Liquid AI is an efficiency-first foundation-model company building device-native Liquid Foundation Models (LFMs) and tools to customize and deploy them.

Cambridge, Massachusetts, United States
About Liquid AI

An MIT CSAIL spinout, Liquid AI develops general-purpose AI models focused on efficient deployment across CPUs, GPUs, NPUs, edge devices, and cloud or on-premises environments.

View jobs by Liquid AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will operate and improve GPU clusters supporting training and research. You will debug compute, storage, networking, scheduler, and workload issues; improve resource utilization; migrate workloads; and build monitoring, validation, automation, and platform abstractions.

Requirements

  • Software engineering experience building production-quality infrastructure tooling and automation
  • Knowledge of distributed systems, Linux, networking, and storage
  • Experience operating a shared compute cluster or distributed training platform
  • Experience supporting production users and converting recurring failures into durable solutions
  • Ability to partner with senior research and infrastructure engineers

Responsibilities

  • Own the reliability and operation of GPU clusters
  • Debug compute, storage, networking, scheduler, and distributed-workload issues
  • Improve CPU, GPU, and storage utilization through tooling and automation
  • Onboard and migrate workloads across GPU providers and hardware platforms
  • Build monitoring, validation, and platform abstractions
  • Contribute to training infrastructure and GPU platform architecture

Benefits

  • Equity
  • Medical, dental, and vision premiums fully paid for employees and dependents
  • 401(k) matching up to 4% of base pay
  • Unlimited PTO
  • Company-wide Refill Days
Member of Technical Staff GPU Infrastructure Engineer at Liquid AI | JobStash