Centre of Excellence Senior Engineer
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Maintainer signals as of 9/25/2026
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead complex infrastructure projects, improve operational processes across data centres, and coordinate major incident responses. You will create standards, runbooks, and scalable operating models; support capacity planning and deployment readiness; mentor operations teams; and travel occasionally to support site deployments and critical initiatives.
Requirements
- 5+ years of experience in technical infrastructure, data-centre, or operations environments
- Extensive experience in large-scale data-centre environments
- Ability to solve complex systemic infrastructure issues and deliver operational improvements
- Experience leading technical initiatives across multiple sites or operational teams
- Advanced technical certifications relevant to infrastructure operations
- Strong analytical troubleshooting and problem-solving skills
- Experience with hyperscale, HPC, AI, or NVIDIA GPU infrastructure
- Hands-on server deployment, maintenance, and troubleshooting experience
- Communication and stakeholder-management skills
Responsibilities
- Lead hardware installation, upgrade, and operational-improvement projects
- Develop repeatable operational processes and best practices
- Identify opportunities to improve efficiency, scalability, and service reliability
- Own complex infrastructure incidents from diagnosis through resolution
- Lead major incident response and coordinate technical teams
- Conduct root-cause analysis and implement corrective actions
- Provide technical leadership and mentorship to operations teams
- Support capacity planning, infrastructure optimisation, and deployment readiness
- Create and improve operational documentation, standards, and runbooks
- Maintain procedures for incidents, maintenance, and communications
- Travel to data-centre sites to support deployments and resolve technical issues
Benefits
- Competitive package with reviews every 12 months
- Flexible workplace
- Remote-first collaboration with an option for office-based work
