Forward Deployed Engineer - ML

Modal Labs, Inc. operates Modal, a serverless cloud and AI infrastructure platform for developers running inference, training, batch processing, and isolated sandboxes.

New York City, United States
About Modal

Modal provides code-first, elastic CPU/GPU compute infrastructure for AI workloads, including model inference, fine-tuning and training, large-scale batch jobs, and secure ephemeral execution environments.

View jobs by Modal

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will architect and optimize production AI workloads with customers, including LLM serving, model training, audio pipelines, and scientific computing. You will contribute to open-source projects and technical content, run demos and proof-of-concepts, collaborate on product needs, and build trusted technical customer relationships.

Requirements

  • 2+ years of professional ML engineering experience
  • Experience with inference optimization, model training, GPU programming, or ML infrastructure
  • Familiarity with serving or training toolchains such as vLLM, SGLang, slime, verl, or TRL
  • Technical architecture communication skills
  • Interest in working directly with customers
  • Willingness to work in person in Stockholm

Responsibilities

  • Architect and optimize production AI workloads with customers
  • Contribute to open-source projects and publish technical content
  • Collaborate with product and sales teams on the platform
  • Build trusted relationships with technical leaders
  • Conduct technical demos, experiments, and proof-of-concepts
Forward Deployed Engineer - ML at Modal | JobStash