Research Scientist or Engineer Foundation Model Agent

Luma AI develops multimodal creative AI agents, video and image generation models, and an API for generative media workflows.

Redwood City, United States
About Luma AI

Luma AI is an active AI company whose Luma Agents product plans, generates, iterates, and refines creative work across video, image, audio, and text. Its first-party materials describe proprietary Ray video and Uni image models, as well as developer API access.

View jobs by Luma AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will architect, train, and evaluate large-scale multimodal agentic models that reason, plan, code, and call tools. You will build data pipelines for pixel datasets, train models on GPU clusters, and develop evaluation frameworks for multimodal agents.

Requirements

  • Foundation in machine learning, foundation models, and agentic systems
  • Understanding of agentic systems, LLM and VLM reasoning, coding models, and tool calling
  • Hands-on experience with PyTorch and large-scale distributed training, mixed precision, and large datasets

Responsibilities

  • Architect large-scale multimodal agentic models for complex multi-step work
  • Design, build, and run data pipelines for massive pixel datasets
  • Train large-scale multimodal models on GPU clusters
  • Define and build evaluation frameworks for multimodal agents