Member of Technical Staff ML Research Engineer Data
Liquid AI is an efficiency-first foundation-model company building device-native Liquid Foundation Models (LFMs) and tools to customize and deploy them.
Funding history
About Liquid AI
An MIT CSAIL spinout, Liquid AI develops general-purpose AI models focused on efficient deployment across CPUs, GPUs, NPUs, edge devices, and cloud or on-premises environments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build scalable data processing, filtering, and selection pipelines for foundation-model training. You will create datasets for pretraining, post-training, supervised fine-tuning, and preference optimization; develop synthetic-data systems; run evaluations and ablations; and support text, vision, and audio data needs.
Requirements
- Strong Python skills
- ML fundamentals, including training, evaluating, and iterating on models
- Experience with PyTorch
- Ability to learn new technical domains quickly
- 3+ years of relevant experience with an M.S., 1+ year with a Ph.D., or 5+ years with a B.S.
Responsibilities
- Build and maintain scalable data processing, filtering, and selection pipelines
- Create datasets for pretraining, midtraining, supervised fine-tuning, and preference optimization
- Design synthetic-data generation systems using LLMs, structured prompting, and domain-specific generators
- Design and run evaluations and ablations to measure dataset impact on model performance
- Monitor public datasets across text, vision, and audio
- Collaborate on modality-specific data needs
Benefits
- Equity
- Medical, dental, and vision premiums fully paid for employees and dependents
- 401(k) matching up to 4% of base pay
- Unlimited PTO
- Company-wide Refill Days
