Member of Technical Staff Platform Engineering

Modal Labs, Inc. operates Modal, a serverless cloud and AI infrastructure platform for developers running inference, training, batch processing, and isolated sandboxes.

New York City, United States
About Modal

Modal provides code-first, elastic CPU/GPU compute infrastructure for AI workloads, including model inference, fine-tuning and training, large-scale batch jobs, and secure ephemeral execution environments.

View jobs by Modal

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will identify architectural improvements for reliability and performance, implement operational processes, and operate production infrastructure. You will work with Kubernetes, Postgres, Redis, monitoring, CI/CD, fleet management, and capacity planning while participating in on-call incident response.

Requirements

  • 5+ years of experience writing high-quality production code
  • 2+ years of on-call experience for critical production services
  • Strong cloud skills and familiarity with a hyperscaler cloud
  • Familiarity with auto scaling, fleet management, and capacity planning
  • Experience operating databases, monitoring, CI/CD, and infrastructure at scale
  • Ability to work in-person in New York City or Stockholm

Responsibilities

  • Identify architectural changes that improve reliability and performance
  • Foster reliability practices across engineering
  • Define and implement deployment and upgrade processes
  • Operate Kubernetes, Postgres, Redis, and other infrastructure
  • Participate in on-call rotations and respond to production incidents

Benefits

  • Equity
Member of Technical Staff Platform Engineering at Modal | JobStash