Member of Technical Staff GPU Infrastructure Engineer
Liquid AIVisit Liquid AI website
Liquid AI is an efficiency-first foundation-model company building device-native Liquid Foundation Models (LFMs) and tools to customize and deploy them.
Cambridge, Massachusetts, United States
Funding history
About Liquid AI
An MIT CSAIL spinout, Liquid AI develops general-purpose AI models focused on efficient deployment across CPUs, GPUs, NPUs, edge devices, and cloud or on-premises environments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will operate and improve GPU clusters supporting training and research. You will debug compute, storage, networking, scheduler, and workload issues; improve resource utilization; migrate workloads; and build monitoring, validation, automation, and platform abstractions.
Requirements
- Software engineering experience building production-quality infrastructure tooling and automation
- Knowledge of distributed systems, Linux, networking, and storage
- Experience operating a shared compute cluster or distributed training platform
- Experience supporting production users and converting recurring failures into durable solutions
- Ability to partner with senior research and infrastructure engineers
Responsibilities
- Own the reliability and operation of GPU clusters
- Debug compute, storage, networking, scheduler, and distributed-workload issues
- Improve CPU, GPU, and storage utilization through tooling and automation
- Onboard and migrate workloads across GPU providers and hardware platforms
- Build monitoring, validation, and platform abstractions
- Contribute to training infrastructure and GPU platform architecture
Benefits
- Equity
- Medical, dental, and vision premiums fully paid for employees and dependents
- 401(k) matching up to 4% of base pay
- Unlimited PTO
- Company-wide Refill Days
