Senior Observability Platform Engineer
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design, build, and operate scalable observability systems for metrics, logs, traces, and alerting. You will improve signal quality, contribute to architecture and data-retention decisions, identify observability gaps, and integrate observability into services and platforms. You will also support incident response, evaluate tooling, develop reusable patterns, and mentor engineers through reviews and knowledge sharing.
Requirements
- 5+ years of experience in SRE, infrastructure engineering, platform engineering, or observability-focused roles
- Experience operating and scaling production observability systems
- Knowledge of metrics, logs, traces, alerting, and SLOs
- Experience with observability technologies including Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, or Elastic
- Programming skills in Python, Go, or similar languages
- Experience with Kubernetes-based infrastructure
- Familiarity with Infrastructure as Code tools such as Terraform or Ansible
Responsibilities
- Design, build, and operate scalable observability systems across metrics, logs, traces, and alerting
- Contribute to tooling, data-pipeline, storage, and retention architecture
- Improve signal quality by reducing noise, managing cardinality, and refining alerting
- Identify and address observability gaps before they affect reliability
- Integrate observability into services and platforms with SRE, infrastructure, and AI/ML stakeholders
- Develop reusable patterns, libraries, and best practices
- Participate in incident response and postmortems
- Evaluate and adopt tools that improve developer experience, scalability, and operational efficiency
- Mentor engineers through code reviews and knowledge sharing
Benefits
- Medical insurance
- Dental insurance
- Vision insurance
- Flexible paid time off
- Parental leave
- Retirement plan participation
