Member of Technical Staff Inference Runtime
Modal Labs, Inc. operates Modal, a serverless cloud and AI infrastructure platform for developers running inference, training, batch processing, and isolated sandboxes.
Maintainer signals as of 9/25/2026
Funding history
About Modal
Modal provides code-first, elastic CPU/GPU compute infrastructure for AI workloads, including model inference, fine-tuning and training, large-scale batch jobs, and secure ephemeral execution environments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build production container runtime infrastructure in Rust and Go. You will improve startup, checkpointing, storage, GPU, filesystem, and networking performance; debug low-level Linux and GPU issues; and safely deploy runtime and kernel changes across the fleet.
Requirements
- Strong Linux systems knowledge, including processes, virtual memory, filesystems, system calls, scheduling, namespaces, cgroups, and signals
- Experience building or debugging container runtimes, sandboxes, Linux kernel, or similar low-level infrastructure
- Strong Rust, Go, or systems-language programming ability
- Experience profiling and improving memory, I/O, synchronization, or kernel-interaction performance
- Ability to debug across application, runtime, driver, and kernel layers
- Ability to turn loosely defined production problems into reliable systems
Responsibilities
- Accelerate container startup, checkpoint, and restore for inference and training workloads
- Build multi-GPU and accelerator-aware snapshotting
- Design efficient data paths across container memory, filesystems, storage, and the runtime
- Optimize snapshot pipelines and filesystem performance
- Extend the sandboxed runtime for GPUs, drivers, profiling tools, and device capabilities
- Debug Linux, kernel, GPU driver, and container-isolation failures
- Roll out runtime and kernel changes using compatibility controls, scheduling constraints, feature flags, and observability
- Investigate and deploy solutions for production runtime, scheduler, storage, and GPU problems
Benefits
- Equity
