Search...

Site Reliability Engineer Intermediate to Senior Staff Infrastructure Platforms

GitLab Inc. logo
GitLab Inc.

GitLab is an AI-powered DevSecOps platform that unifies the entire software development lifecycle into a single application. It helps development, security, and operations teams to collaborate and deliver software more efficiently, with security integrated at every step. The platform is trusted by millions of users and a majority of the Fortune 100.

Distributed
About GitLab Inc.

GitLab is a comprehensive, AI-powered DevSecOps platform that streamlines the entire software delivery process by unifying the development lifecycle into a single application. It integrates source code management, CI/CD, security, and monitoring to help teams build, secure, and operate software more efficiently. Key features include automated security scans built into the development pipeline and AI-driven tools like GitLab Duo for code suggestions and chat, which enhance developer productivity. GitLab serves a diverse client base, from startups and open-source projects to large enterprises, including over half of the Fortune 100. The platform aims to reduce complexity, accelerate delivery cycles, and strengthen security and compliance for its users.

View jobs by GitLab Inc.

Skills

About the Role

You will keep user-facing services and production systems reliable, scalable, and efficient. You will build infrastructure automation, operate and troubleshoot Kubernetes systems, manage infrastructure as code, and safely deliver changes through CI/CD and GitOps. You will participate in on-call rotations, triage alerts, improve runbooks, support incident response and post-incident reviews, and document architecture decisions and repeatable practices.

Requirements

  • Experience maintaining reliable production systems using operations and software engineering practices
  • Experience building new infrastructure tooling and automation
  • Ability to read, debug, and reason about code behavior, performance, and failure modes
  • Experience with infrastructure as code and Kubernetes
  • Hands-on experience with GCP or AWS
  • Familiarity with metrics, logging, alerting, SLOs, and SLIs
  • Experience participating in on-call and incident response
  • Written communication skills for an asynchronous, distributed environment
  • Experience using automation and AI to reduce toil

Responsibilities

  • Keep user-facing services and production systems reliable, scalable, and efficient
  • Build automation and tooling that reduces toil and replaces manual work with infrastructure-as-code-driven workflows
  • Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling
  • Write and maintain infrastructure as code and ship changes through CI/CD and GitOps
  • Participate in on-call rotations, triage alerts, improve runbooks, and escalate appropriately
  • Contribute to observability using metrics, logs, and SLOs
  • Participate in incident response and post-incident reviews
  • Document runbooks, architecture decisions, and reviews

Benefits

  • Health, financial, and well-being benefits
  • Flexible Paid Time Off
  • Team Member Resource Groups
  • Equity compensation and Employee Stock Purchase Plan
  • Parental Leave
Site Reliability Engineer Intermediate to Senior Staff Infrastructure Platforms at GitLab Inc. | JobStash