Forward Deployed Engineer ML

Modal Labs, Inc. operates Modal, a serverless cloud and AI infrastructure platform for developers running inference, training, batch processing, and isolated sandboxes.

New York City, United States
About Modal

Modal provides code-first, elastic CPU/GPU compute infrastructure for AI workloads, including model inference, fine-tuning and training, large-scale batch jobs, and secure ephemeral execution environments.

View jobs by Modal

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will work directly with AI companies to architect and optimize production workloads, including model serving and training. You will conduct demos, experiments, and proofs of concept, publish technical content, contribute to open source, and collaborate with product and sales stakeholders.

Requirements

  • 2+ years of professional ML engineering experience
  • Experience with inference optimization, model training, GPU programming, or ML infrastructure
  • Familiarity with serving toolchains including vLLM or SGLang
  • Familiarity with training toolchains including slime, verl, or TRL
  • Ability to communicate technical architecture and tradeoffs to engineering teams and technical leadership
  • Interest in working directly with customers
  • Willingness to work in person in New York City, San Francisco, or Stockholm

Responsibilities

  • Architect and optimize production AI workloads with customers
  • Contribute to open-source projects and publish technical content
  • Collaborate with product and sales teams as an engineer and product stakeholder
  • Build relationships with technical leaders
  • Conduct technical demos, experiments, and proofs of concept

Benefits

  • Equity