Infrastructure Operations Engineer
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will manage data center installations and upgrades, monitor critical systems, resolve escalated support issues, automate operational processes, investigate incidents, and support capacity planning. You will work shifts, join on-call coverage, and travel to other data centers when needed.
Requirements
- 4+ years of experience in data center environments as an operator, technician, or engineer
- Expertise in server administration, virtualization, and networking protocols
- Familiarity with Ansible, PowerShell, Python, or similar automation tools
- Problem-solving and analytical skills
- Advanced technical certifications
Responsibilities
- Manage infrastructure projects, installations, and upgrades
- Work shifts in AI data centers
- Monitor and maintain power, cooling, and network connectivity
- Handle escalated customer support issues against SLA requirements
- Design and implement automation scripts and tools
- Conduct root-cause analysis for major incidents and recommend long-term fixes
- Collaborate on capacity planning and infrastructure optimization
- Respond to critical incidents and participate in on-call coverage
- Provide specialist support outside core business hours
- Travel to other data centers to support deployments and operational tasks
