Staff Infrastructure Software Engineer Fleet and Automation
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead the architecture and implementation of fleet-management and workflow-automation systems. You will build production services, APIs, and control-plane workflows for GPU and network infrastructure, investigate production issues, establish reliability standards, and mentor engineers.
Requirements
- 8+ years of experience building and operating large-scale infrastructure applications, platform services, cloud systems, or equivalent production systems
- Software engineering experience with Python, Go, Java, C++, or similar languages
- Linux
- Distributed systems
- Networking
- Reliable automation or control-plane systems
- Observability
- Monitoring
- Metrics
- Logs
- Tracing
- Alerting
- Incident response
- Capacity planning
- Performance analysis
Responsibilities
- Lead the architecture, roadmap, and implementation of workflow automation and fleet-management systems
- Build and operate production software, services, APIs, and automation for GPU compute and network infrastructure
- Own fleet inventory, provisioning, configuration, lifecycle management, validation, monitoring, remediation, capacity, and reliability workflows
- Investigate production issues and deliver durable software improvements
- Build safe, observable, and auditable control-plane workflows
- Establish standards for reliability, observability, testing, CI/CD, security, incident response, and operational readiness
- Partner with operational and engineering teams to deliver scalable automation
- Lead technical design reviews and incident deep-dives
- Mentor engineers
Benefits
- Equity
- Package reviews every 12 months
