Senior Operational Engineer
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Maintainer signals as of 9/25/2026
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design and deliver operational capabilities such as readiness checks, canaries, runbooks, reporting, alerting workflows, and automation. You will establish service standards, improve operational integrations, lead incident analysis, automate reporting, mentor engineers, and support continuity and recovery testing.
Requirements
- 6–10 years of experience in operations engineering, SRE, cloud infrastructure, platform engineering, or a similar production-focused engineering role
- Experience designing and operating monitoring, alerting, on-call, runbook, service-readiness, and operational-reporting capabilities
- Experience building integrations and automations across Grafana, PagerDuty, Jira, Backstage, and public cloud
- Experience using incident, change, service-health, continuity, patching, or cost data to drive operational improvement
- Ability to lead cross-service work, set standards, and support teams through incidents
Responsibilities
- Design and deliver shared operational capabilities, including readiness checks, canaries, runbooks, reporting, and automation
- Establish standards for service ownership, on-call readiness, alerts, dashboards, recovery procedures, and operational evidence
- Identify and resolve recurring operational issues and cross-team blockers
- Improve integrations and workflows across Grafana, PagerDuty, Jira, Backstage, public-cloud platforms, and reporting systems
- Turn incident findings, change failures, and near misses into engineering improvements
- Act as an incident commander and lead technical incident analysis
- Build observable and auditable automations
- Automate SLA, health, cost, patching, and action-closure reporting
- Mentor engineers and review operational designs
- Contribute to continuity testing and failure-and-recovery experiments
