Reliability Engineer SME (M&E)
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Maintainer signals as of 9/23/2026
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will serve as the technical escalation point for mechanical and electrical systems across data centre sites. You will review designs and changes, support live incidents, lead root-cause investigations, maintain engineering standards, conduct audits, deliver training, and improve power, cooling, capacity, and energy performance.
Requirements
- Building-services engineering experience in critical environments.
- Hands-on mechanical or electrical systems expertise with working knowledge of the other discipline.
- Data centre cooling or power-systems knowledge.
- M&E design-review and engineering-standards compliance experience.
- Technical-audit experience.
- Change-control and live-system operational-risk knowledge.
- Root-cause analysis and corrective and preventive action experience.
- Experience supporting live incidents and training operational engineers.
- Technical communication skills.
Responsibilities
- Provide technical escalation support for mechanical and electrical systems.
- Set and maintain M&E maintenance standards.
- Review M&E designs, vendor submissions, drawings, schematics, and documentation.
- Support AI rack cooling and power-distribution systems.
- Review and approve changes to live systems.
- Conduct technical and colocation site audits.
- Support incident diagnosis and safe system recovery.
- Own root-cause analysis and corrective and preventive actions.
- Maintain engineering standards and deliver technical training.
- Support site due diligence, capacity planning, asset lifecycle management, and PUE improvement.
Benefits
- Bonus
- Equity
- Flexible workplace
