Senior Infrastructure Engineer (OpenStack Ironic Specialist)

Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.

London, United Kingdom
About Nscale

Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.

View jobs by Nscale

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design and operate scalable bare-metal provisioning platforms focused on OpenStack Ironic. You will manage hardware lifecycles, automate onboarding and deployment workflows, integrate Ironic with OpenStack services, troubleshoot complex provisioning issues, participate in on-call incident response, and contribute to upstream OpenStack communities.

Requirements

  • Experience operating Linux systems and troubleshooting production infrastructure
  • Specialist knowledge of OpenStack Ironic and bare-metal provisioning
  • Knowledge of PXE, iPXE, DHCP, TFTP, HTTP boot, BMC, RAID, firmware management, disk imaging, and node lifecycle states
  • Experience with Redfish, IPMI, or vendor management interfaces
  • Experience with Ansible and infrastructure automation
  • Python and Bash scripting skills
  • Experience troubleshooting provisioning and hardware integration issues
  • Experience operating infrastructure at scale
  • Experience with OpenStack, Ironic, Metal3, or related open-source communities is desirable

Responsibilities

  • Design scalable and resilient bare-metal provisioning platforms using OpenStack Ironic
  • Own physical infrastructure lifecycle management from discovery through deprovisioning
  • Build provisioning workflows for GPU-enabled and high-performance server platforms
  • Manage integrations between Ironic and Nova, Neutron, Glance, Keystone, and Placement
  • Automate hardware onboarding, firmware configuration, deployment, validation, and recovery
  • Troubleshoot PXE, iPXE, BMC, image deployment, network boot, and hardware compatibility issues
  • Act as a senior escalation point for provisioning incidents
  • Perform root cause analysis and implement long-term reliability fixes
  • Participate in on-call rotations and incident response
  • Contribute to upstream OpenStack bare-metal communities

Benefits

  • Bonus
  • Equity
  • Flexible workplace
  • Medical insurance
  • Dental insurance
  • Vision insurance
  • Flexible paid time off
  • Parental leave
  • Retirement plan participation