Member of Technical Staff, Distributed Training Systems

Bagel Labs is a physical AI research lab developing compact world-action models and distributed training systems for autonomous robot control.

San Francisco, California, United States; Toronto, Ontario, Canada
About Bagel Labs

Bagel Labs currently describes itself as a physical AI research lab focused on a General World Action Model for autonomous robot control across tasks, environments, and robot types. Its current work centers on PARIS, a distributed training architecture, and WorldDiT, a compact diffusion-transformer architecture that unifies world-state prediction with robot-action generation.

View jobs by Bagel Labs

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Build and operate systems that support distributed training across heterogeneous compute, create reliable experiment infrastructure and evaluation harnesses, maintain traceable data and model pipelines, add observability, and turn fragile research prototypes into repeatable runs and trustworthy artifacts.

Requirements

  • Hands-on experience with distributed training, GPU workloads, experiment infrastructure, or large-scale machine learning systems.
  • Ability to debug performance, reliability, and reproducibility problems in complex training and evaluation workflows.
  • Taste for simple tools that researchers will adopt.
  • Clear communication.
  • Strong ownership.

Responsibilities

  • Build and operate distributed training for diffusion-heavy workloads across heterogeneous compute.
  • Make experiments reliable with launchers, configurations, checkpointing, logging, metrics, run comparison, and reproducibility.
  • Build benchmark and evaluation harnesses for physical AI research, including robotics and world-model experiments.
  • Own data and model pipelines so results trace back to a dataset, version, and configuration.
  • Add observability for GPU utilization, failure modes, data quality, routing behavior, model quality, and training stability.
  • Turn fragile research prototypes into repeatable runs and trustworthy artifacts.

Benefits

  • Competitive compensation.
  • Meaningful equity.
  • Deeply technical culture connecting research and systems work.
  • Ownership of foundational infrastructure for frontier AI research.
  • Paid travel to top machine learning and systems conferences around the world.