Member of Technical Staff Platform Engineering
Modal Labs, Inc. operates Modal, a serverless cloud and AI infrastructure platform for developers running inference, training, batch processing, and isolated sandboxes.
Funding history
About Modal
Modal provides code-first, elastic CPU/GPU compute infrastructure for AI workloads, including model inference, fine-tuning and training, large-scale batch jobs, and secure ephemeral execution environments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will identify architectural improvements for reliability and performance, implement operational processes, and operate production infrastructure. You will work with Kubernetes, Postgres, Redis, monitoring, CI/CD, fleet management, and capacity planning while participating in on-call incident response.
Requirements
- 5+ years of experience writing high-quality production code
- 2+ years of on-call experience for critical production services
- Strong cloud skills and familiarity with a hyperscaler cloud
- Familiarity with auto scaling, fleet management, and capacity planning
- Experience operating databases, monitoring, CI/CD, and infrastructure at scale
- Ability to work in-person in New York City or Stockholm
Responsibilities
- Identify architectural changes that improve reliability and performance
- Foster reliability practices across engineering
- Define and implement deployment and upgrade processes
- Operate Kubernetes, Postgres, Redis, and other infrastructure
- Participate in on-call rotations and respond to production incidents
Benefits
- Equity
