Site Reliability Engineer
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build automation and tooling for production systems, define service-level objectives and health dashboards, lead incident response, and perform root-cause analysis. You will investigate Linux, networking, and distributed-service issues, then improve availability, scalability, and efficiency through code.
Requirements
- 3 to 6 years of SRE, systems engineering, or software engineering experience
- Experience operating production systems in a data-centre or cloud environment
- Programming skills in Python, Go, or similar languages
- Linux, networking, and distributed-systems knowledge
- Experience troubleshooting live production issues
- Monitoring and observability experience with metrics, logs, dashboards, and alerting
- Ability to work in a fast-moving environment
Responsibilities
- Build and own platform automation and tooling
- Define and maintain SLOs, SLIs, and service-health dashboards
- Lead incident response, root-cause analysis, and post-incident reviews
- Investigate and fix performance and reliability issues
- Improve availability, scalability, and efficiency through code
- Partner with engineering, networking, and infrastructure stakeholders
- Participate in an on-call rotation
Benefits
- Equity
- Flexible work arrangements
