Member of Technical Staff Research Inference
Modal Labs, Inc. operates Modal, a serverless cloud and AI infrastructure platform for developers running inference, training, batch processing, and isolated sandboxes.
Funding history
About Modal
Modal provides code-first, elastic CPU/GPU compute infrastructure for AI workloads, including model inference, fine-tuning and training, large-scale batch jobs, and secure ephemeral execution environments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead end-to-end inference research, from selecting high-impact bets to delivering results. You will optimize serving cost and latency, train custom speculators, tune customer deployments, collaborate with research labs, and translate frontier serving techniques into production capabilities.
Requirements
- Research or systems background in LLM inference
- Fluency in the LLM serving stack, including kernels, quantization, schedulers, and autoscaling
- Record of shipping research or systems that others use
- Ability to independently take research from idea to result
- Ability to work in person in the New York City or San Francisco office
Responsibilities
- Own end-to-end inference research bets
- Train custom speculators using production traffic
- Deploy and tune models with customers
- Maintain collaborations with external research labs
- Translate serving techniques into production products
- Help shape the research agenda
Benefits
- Equity
