Member of Technical Staff Distributed Training Engineer
Liquid AIVisit Liquid AI website
Liquid AI is an efficiency-first foundation-model company building device-native Liquid Foundation Models (LFMs) and tools to customize and deploy them.
Cambridge, Massachusetts, United States
Funding history
About Liquid AI
An MIT CSAIL spinout, Liquid AI develops general-purpose AI models focused on efficient deployment across CPUs, GPUs, NPUs, edge devices, and cloud or on-premises environments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design and optimize distributed systems for large-scale training. You will implement parallelism and sharding strategies, improve distributed efficiency, remove data-loading bottlenecks, develop checkpointing, and create monitoring, profiling, and debugging tools.
Requirements
- Experience building distributed training infrastructure with PyTorch Distributed DDP or FSDP, DeepSpeed ZeRO, or Megatron-LM TP or PP
- Experience diagnosing performance bottlenecks and failure modes
- Understanding of hardware accelerators and networking topologies
- Experience optimizing ML data pipelines
Responsibilities
- Design and build reliable systems for large training runs
- Build scalable distributed training infrastructure for GPU clusters
- Implement and tune parallelism and sharding strategies
- Optimize distributed efficiency
- Build data-loading systems for multimodal datasets
- Develop checkpointing mechanisms
- Create monitoring, profiling, and debugging tools
Benefits
- Equity
- Medical, dental, and vision premiums fully paid for employees and dependents
- 401(k) matching up to 4% of base pay
- Unlimited PTO
- Company-wide Refill Days
