Operational Engineer

Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.

Series CRecently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

London, United Kingdom
About Nscale

Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.

View jobs by Nscale

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will improve the reliability, security, measurability, and operability of engineering services. You will onboard services to operational tooling, build integrations and automation, coordinate incidents, improve readiness, and use operational data to reduce manual work.

Requirements

  • 2–5 years of operations engineering, SRE, cloud infrastructure, platform engineering, or similar experience
  • Experience operating or supporting production services
  • Experience with monitoring, alerting, on-call, runbooks, and engineering work management
  • Experience with operational tools such as Grafana, PagerDuty, Jira, Backstage, public cloud, dashboards, integrations, or workflow automation
  • Ability to use operational data to identify gaps and complete follow-up actions

Responsibilities

  • Onboard services to alerting, on-call, dashboards, runbooks, Jira, and service-catalog records
  • Build and maintain integrations, dashboards, reports, and automation
  • Identify and address gaps in service ownership, alerting, documentation, recovery procedures, and dependencies
  • Coordinate incident response, communications, timelines, post-incident reviews, and action tracking
  • Make routine changes safer and more repeatable through processes and automation
  • Contribute to service-health, SLA, cost, patching, and operational-risk reporting
  • Maintain operational data including ownership, on-call coverage, runbooks, and dependencies
  • Reduce manual toil and improve operational outcomes