Member of Technical Staff Research Inference

Modal Labs, Inc. operates Modal, a serverless cloud and AI infrastructure platform for developers running inference, training, batch processing, and isolated sandboxes.

New York City, United States
About Modal

Modal provides code-first, elastic CPU/GPU compute infrastructure for AI workloads, including model inference, fine-tuning and training, large-scale batch jobs, and secure ephemeral execution environments.

View jobs by Modal

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead end-to-end inference research, from selecting high-impact bets to delivering results. You will optimize serving cost and latency, train custom speculators, tune customer deployments, collaborate with research labs, and translate frontier serving techniques into production capabilities.

Requirements

  • Research or systems background in LLM inference
  • Fluency in the LLM serving stack, including kernels, quantization, schedulers, and autoscaling
  • Record of shipping research or systems that others use
  • Ability to independently take research from idea to result
  • Ability to work in person in the New York City or San Francisco office

Responsibilities

  • Own end-to-end inference research bets
  • Train custom speculators using production traffic
  • Deploy and tune models with customers
  • Maintain collaborations with external research labs
  • Translate serving techniques into production products
  • Help shape the research agenda

Benefits

  • Equity
Member of Technical Staff Research Inference at Modal | JobStash