Search...

Senior Site Reliability Engineer

Latitude logo
Latitude

Latitude.sh provides global bare metal cloud infrastructure with automated deployment and management through a unified API and dashboard.

Brazil
About Latitude

Latitude is a global Bare Metal Cloud provider with an automated platform that enables instant deployment and easy management of dedicated, high-performance, low-latency servers around the world, optimized for the digital economy.

View jobs by Latitude

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build reliable, observable, and self-healing infrastructure at scale. You will automate operational tasks and incident response, improve monitoring and alerting, collaborate on resilient system designs, participate in on-call rotations, lead post-incident reviews, and document operational processes and runbooks.

Requirements

  • Advanced knowledge of Linux/Unix systems in production environments
  • Experience with Kubernetes and container orchestration
  • Proficiency with Terraform and Ansible
  • Experience with Prometheus, Grafana, Loki, or ELK
  • Familiarity with Bash, Python, Go, or Ruby
  • Working knowledge of Git and CI/CD pipelines
  • Understanding of incident management and root cause analysis
  • Knowledge of cloud-native reliability and security best practices

Responsibilities

  • Improve platform reliability and performance
  • Design, build, and maintain tools that automate operational tasks and incident response
  • Implement and improve monitoring, alerting, and tracing solutions
  • Collaborate on scalable and resilient system designs
  • Participate in on-call rotations
  • Lead post-incident reviews
  • Develop and document operational processes and runbooks
  • Contribute to SLO, SLI, and reliability-metric adoption

Benefits

  • Paid Time Off
  • Wellhub
  • Annual bonus based on company and team performance
  • Flexible work hours