Staff Observability Platform Engineer
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design, build, and improve observability platforms for metrics, logs, traces, alerting, and telemetry. You will lead technical initiatives, improve monitoring and incident response, develop reusable standards, guide engineers, and turn operational learnings into durable platform improvements.
Requirements
- 6+ years of experience in SRE, platform engineering, infrastructure engineering, observability engineering, or related disciplines
- Experience building and operating observability platforms in cloud-native distributed environments
- Experience with Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic, or similar platforms
- Proficiency in Go, Python, or equivalent languages
- Experience operating and troubleshooting Kubernetes platforms at scale
- Knowledge of monitoring, logging, tracing, telemetry pipelines, and observability practices
- Experience designing scalable, reliable, and performant systems
- Proficiency with Terraform, Ansible, or equivalent infrastructure-as-code tools
- Ability to lead technical initiatives and influence engineering decisions
- Communication skills for explaining technical trade-offs and aligning stakeholders
Responsibilities
- Design, build, and evolve observability platforms across metrics, logs, traces, alerting, and telemetry pipelines
- Lead scalable observability solutions for GPU and AI infrastructure
- Improve monitoring coverage, alert quality, service health visibility, and incident response
- Develop reusable observability standards, frameworks, and patterns
- Identify reliability risks and operational blind spots
- Contribute to telemetry architecture and performance decisions
- Mentor engineers through reviews and knowledge sharing
- Participate in incident investigations and postmortems
- Evaluate observability technologies and practices
Benefits
- Bonus eligibility
- Equity eligibility
- Medical coverage
- Dental coverage
- Vision coverage
- Flexible paid time off
- Parental leave
- Retirement plan participation
