Senior Engineer, Storage Services
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build, operate, and improve Block, File, Parallel, and Object storage services. You will integrate storage capabilities with compute, networking, and platform tooling; monitor health, performance, and capacity; troubleshoot incidents; improve observability and automation; and maintain operational documentation.
Requirements
- Experience as a Storage Engineer, Infrastructure Engineer, Systems Engineer, or Platform Engineer
- Knowledge of Block, File, Parallel, or Object storage
- Experience supporting production infrastructure or platform services
- Understanding of Linux systems, networking fundamentals, and infrastructure operations
- Experience with automation, scripting, or infrastructure as code
- Troubleshooting skills for technical, performance, and operational issues
- Understanding of monitoring, observability, and incident response
- Experience in cloud, AI, HPC, data-intensive, or performance-sensitive environments is advantageous
Responsibilities
- Build storage services across Block, File, Parallel, and Object storage technologies
- Operate and improve scalable, resilient, and secure storage platforms
- Implement designs and service improvements with senior engineers
- Integrate storage services with compute, networking, and platform tooling
- Monitor storage service health, performance, and capacity
- Improve observability, alerting, and operational readiness
- Troubleshoot incidents, performance issues, and service problems
- Support root cause analysis and long-term reliability improvements
- Contribute automation, tooling, documentation, and process optimisation
- Maintain documentation for storage systems and operational procedures
Benefits
- Bonus
- Equity
- Flexible workplace
