Operations Engineering Manager
Northern Data Group provides vertically integrated AI infrastructure, including Taiga Cloud AI compute services and Ardent Data Centers. It serves enterprises, AI companies, cloud providers, research organizations, and other customers with high-performance AI and HPC workloads. The website states that Northern Data is now part of Quake AI.
Maintainer signals as of 8/23/2026
Funding history
About Northern Data Group
Northern Data Group operates advanced full-stack AI infrastructure, combining high-density data centers with an AI cloud software layer. Its Taiga Cloud offering provides GPU and bare-metal compute, storage, networking, managed Kubernetes and SLURM services, while Ardent Data Centers provides colocation and modular data-center infrastructure. Its services target organizations requiring scalable, secure, high-performance infrastructure for AI, high-performance computing, cloud, research, and latency-sensitive workloads. The website states that Northern Data is now part of Quake AI, with accounts, services, and billing unchanged.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead and develop Operations Engineers while setting priorities, managing resources, and supporting performance development. You will own the reliability, availability, and performance of GPU accelerated HPC infrastructure, oversee monitoring and incident analysis, define operational metrics, improve runbooks and change management, and drive automation. You will champion scaled Agile practices, coordinate cross functional delivery, maintain operational documentation, support incident escalations, and communicate status and risks to stakeholders.
Requirements
- 5+ years of infrastructure or operations experience
- 2+ years managing a technical team
- Advanced Linux administration experience in production environments
- Experience with incident and problem management
- Experience working with third party or external support teams
- Hands on automation experience with Ansible or equivalent
- Experience with monitoring and observability tools such as Grafana and Prometheus
- Experience with Agile ways of working and scaled Agile frameworks
- Excellent communication and stakeholder management skills
- Experience with HPC or GPU accelerated environments is desirable
- Python and Bash scripting skills are desirable
- Understanding of HPC and GPU performance tuning is desirable
- Experience with CI/CD pipelines and modern DevOps tooling is desirable
- Experience designing or improving on call rotations runbooks and incident readiness is desirable
Responsibilities
- Lead coach and develop Operations Engineers
- Set team goals priorities and expectations
- Manage workload and resource allocation
- Own GPU accelerated HPC infrastructure reliability performance and availability
- Oversee monitoring incident trend analysis and root cause analysis
- Define and report operational metrics
- Improve processes runbooks change management automation and observability
- Champion scaled Agile practices and support Agile ceremonies
- Align priorities and manage backlogs with cross functional teams
- Maintain operational documentation SOPs and troubleshooting guides
- Coordinate critical incidents with Platform Network and third party support teams
- Provide status updates and represent Operations in strategic discussions
Benefits
- Flexible work from home
- Hardware provided according to work needs
- Regular wellbeing initiatives
- Diversity and inclusion initiatives
