Inference Engineering and Product Lead

Modal Labs, Inc. operates Modal, a serverless cloud and AI infrastructure platform for developers running inference, training, batch processing, and isolated sandboxes.

New York City, United States
About Modal

Modal provides code-first, elastic CPU/GPU compute infrastructure for AI workloads, including model inference, fine-tuning and training, large-scale batch jobs, and secure ephemeral execution environments.

View jobs by Modal

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead the technical and product direction of an LLM inference platform while contributing hands-on. You will recruit, coach, and manage engineers; set technical standards; guide design and code reviews; turn customer learnings into roadmaps; and partner on compute strategy, product launches, and platform roadmaps.

Requirements

  • 10+ years of industry experience
  • 3+ years in a leadership role
  • Experience building high-performance systems at scale
  • Cloud infrastructure background
  • Low-level operating system knowledge including Linux kernel, file systems, and containers
  • Production LLM inference experience and knowledge of inference engines, kernels, routing, KV cache management, or speculative decoding

Responsibilities

  • Recruit, hire, coach, and grow engineers
  • Set performance expectations and foster accountability
  • Drive technical and product decisions through reviews and architectural discussions
  • Lead work with customers on novel workloads
  • Translate customer learnings into product roadmaps
  • Establish reliability and product excellence standards
  • Partner on compute purchase strategy
  • Collaborate on product launches, positioning, and inference opportunities
  • Guide infrastructure and adjacent product roadmaps

Benefits

  • Equity