Director Deployment Engineering Systems Engineering
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
About the Role
You will lead systems engineering for deployment, validation, and compute operations. You will set validation standards, lead GPU testing, improve reliability and delivery velocity, align infrastructure stakeholders, establish monitoring and automation practices, and support deployments and site readiness.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related technical field
- At least 10 years of systems engineering, compute operations, infrastructure, or engineering management experience
- Experience in a large cloud provider, hyperscale data center, or similarly complex infrastructure environment
- Experience leading systems or compute operations teams in a high-availability production environment
- Knowledge of Linux or Unix systems administration, OS-level tuning, server architecture, and GPU hardware
- Experience with virtualization, containerization, distributed systems, Kubernetes, and Docker
- Experience with infrastructure as code and configuration management tools
- Experience with scripting, automation, and data center systems design
- Familiarity with system architecture, data synchronization, fault tolerance, state management, and distributed-system reliability
- Experience with storage, networking, compute, monitoring, observability, or telemetry platforms
Responsibilities
- Lead and scale the systems engineering team
- Define and deliver multi-quarter systems engineering initiatives
- Establish validation standards for servers, GPUs, networking, storage, and supporting infrastructure
- Lead GPU burn-in and validation testing at scale
- Partner with infrastructure, network engineering, deployment, product, and operations leaders
- Turn ambiguous technical challenges into plans, priorities, milestones, and accountable execution
- Drive alignment across interdependent teams
- Raise standards for engineering quality, automation, observability, documentation, and operational excellence
- Build approaches for systems monitoring, telemetry, incident learning, and platform improvement
- Travel up to 50% to support deployments, site readiness, vendor collaboration, and operations
Benefits
- Medical insurance
- Dental insurance
- Vision insurance
- Flexible paid time off
- Parental leave
- Retirement plan participation
- Potential bonus, equity, and/or commission programs
