Member of Technical Staff Inference Runtime

Modal Labs, Inc. operates Modal, a serverless cloud and AI infrastructure platform for developers running inference, training, batch processing, and isolated sandboxes.

Series CRecently funded33 current maintainers8 active leads16 lead step-downs5 early lead departuresTeam intelligence

Maintainer signals as of 9/25/2026

New York City, United States
About Modal

Modal provides code-first, elastic CPU/GPU compute infrastructure for AI workloads, including model inference, fine-tuning and training, large-scale batch jobs, and secure ephemeral execution environments.

View jobs by Modal

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build production container runtime infrastructure in Rust and Go. You will improve startup, checkpointing, storage, GPU, filesystem, and networking performance; debug low-level Linux and GPU issues; and safely deploy runtime and kernel changes across the fleet.

Requirements

  • Strong Linux systems knowledge, including processes, virtual memory, filesystems, system calls, scheduling, namespaces, cgroups, and signals
  • Experience building or debugging container runtimes, sandboxes, Linux kernel, or similar low-level infrastructure
  • Strong Rust, Go, or systems-language programming ability
  • Experience profiling and improving memory, I/O, synchronization, or kernel-interaction performance
  • Ability to debug across application, runtime, driver, and kernel layers
  • Ability to turn loosely defined production problems into reliable systems

Responsibilities

  • Accelerate container startup, checkpoint, and restore for inference and training workloads
  • Build multi-GPU and accelerator-aware snapshotting
  • Design efficient data paths across container memory, filesystems, storage, and the runtime
  • Optimize snapshot pipelines and filesystem performance
  • Extend the sandboxed runtime for GPUs, drivers, profiling tools, and device capabilities
  • Debug Linux, kernel, GPU driver, and container-isolation failures
  • Roll out runtime and kernel changes using compatibility controls, scheduling constraints, feature flags, and observability
  • Investigate and deploy solutions for production runtime, scheduler, storage, and GPU problems

Benefits

  • Equity