Operations Engineering Manager

Taiga Cloud is Northern Data Group’s AI-cloud division, providing on-demand NVIDIA GPU and bare-metal compute plus managed Kubernetes and SLURM services.

0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/2/2026

Distributed

Funding history

About Taiga Cloud

Taiga Cloud operates AI infrastructure for model training, prototyping, and scaling, including self-service GPU compute, bare metal, managed Kubernetes, SLURM, and NVIDIA AI Enterprise-related tooling. The 2024 annual report identifies Taiga Cloud Ltd. (Ireland) as the group company providing cloud services to external customers worldwide.

View jobs by Taiga Cloud

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead and develop Operations Engineers while setting priorities, managing resources, and supporting performance development. You will own the reliability, availability, and performance of GPU accelerated HPC infrastructure, oversee monitoring and incident analysis, define operational metrics, improve runbooks and change management, and drive automation. You will champion scaled Agile practices, coordinate cross functional delivery, maintain operational documentation, support incident escalations, and communicate status and risks to stakeholders.

Requirements

  • 5+ years of infrastructure or operations experience
  • 2+ years managing a technical team
  • Advanced Linux administration experience in production environments
  • Experience with incident and problem management
  • Experience working with third party or external support teams
  • Hands on automation experience with Ansible or equivalent
  • Experience with monitoring and observability tools such as Grafana and Prometheus
  • Experience with Agile ways of working and scaled Agile frameworks
  • Excellent communication and stakeholder management skills
  • Experience with HPC or GPU accelerated environments is desirable
  • Python and Bash scripting skills are desirable
  • Understanding of HPC and GPU performance tuning is desirable
  • Experience with CI/CD pipelines and modern DevOps tooling is desirable
  • Experience designing or improving on call rotations runbooks and incident readiness is desirable

Responsibilities

  • Lead coach and develop Operations Engineers
  • Set team goals priorities and expectations
  • Manage workload and resource allocation
  • Own GPU accelerated HPC infrastructure reliability performance and availability
  • Oversee monitoring incident trend analysis and root cause analysis
  • Define and report operational metrics
  • Improve processes runbooks change management automation and observability
  • Champion scaled Agile practices and support Agile ceremonies
  • Align priorities and manage backlogs with cross functional teams
  • Maintain operational documentation SOPs and troubleshooting guides
  • Coordinate critical incidents with Platform Network and third party support teams
  • Provide status updates and represent Operations in strategic discussions

Benefits

  • Flexible work from home
  • Hardware provided according to work needs
  • Regular wellbeing initiatives
  • Diversity and inclusion initiatives
Operations Engineering Manager at Taiga Cloud | JobStash