Member of Technical Staff Distributed Training Engineer

Liquid AI is an efficiency-first foundation-model company building device-native Liquid Foundation Models (LFMs) and tools to customize and deploy them.

Cambridge, Massachusetts, United States
About Liquid AI

An MIT CSAIL spinout, Liquid AI develops general-purpose AI models focused on efficient deployment across CPUs, GPUs, NPUs, edge devices, and cloud or on-premises environments.

View jobs by Liquid AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design and optimize distributed systems for large-scale training. You will implement parallelism and sharding strategies, improve distributed efficiency, remove data-loading bottlenecks, develop checkpointing, and create monitoring, profiling, and debugging tools.

Requirements

  • Experience building distributed training infrastructure with PyTorch Distributed DDP or FSDP, DeepSpeed ZeRO, or Megatron-LM TP or PP
  • Experience diagnosing performance bottlenecks and failure modes
  • Understanding of hardware accelerators and networking topologies
  • Experience optimizing ML data pipelines

Responsibilities

  • Design and build reliable systems for large training runs
  • Build scalable distributed training infrastructure for GPU clusters
  • Implement and tune parallelism and sharding strategies
  • Optimize distributed efficiency
  • Build data-loading systems for multimodal datasets
  • Develop checkpointing mechanisms
  • Create monitoring, profiling, and debugging tools

Benefits

  • Equity
  • Medical, dental, and vision premiums fully paid for employees and dependents
  • 401(k) matching up to 4% of base pay
  • Unlimited PTO
  • Company-wide Refill Days
Member of Technical Staff Distributed Training Engineer at Liquid AI | JobStash