Staff Observability Platform Engineer

37 minutes agoLeadSalary: 190K - 260KUnited StatesDevopsJobs by Nscale

Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.

London, United Kingdom
About Nscale

Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.

View jobs by Nscale

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design, build, and improve observability platforms for metrics, logs, traces, alerting, and telemetry. You will lead technical initiatives, improve monitoring and incident response, develop reusable standards, guide engineers, and turn operational learnings into durable platform improvements.

Requirements

  • 6+ years of experience in SRE, platform engineering, infrastructure engineering, observability engineering, or related disciplines
  • Experience building and operating observability platforms in cloud-native distributed environments
  • Experience with Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic, or similar platforms
  • Proficiency in Go, Python, or equivalent languages
  • Experience operating and troubleshooting Kubernetes platforms at scale
  • Knowledge of monitoring, logging, tracing, telemetry pipelines, and observability practices
  • Experience designing scalable, reliable, and performant systems
  • Proficiency with Terraform, Ansible, or equivalent infrastructure-as-code tools
  • Ability to lead technical initiatives and influence engineering decisions
  • Communication skills for explaining technical trade-offs and aligning stakeholders

Responsibilities

  • Design, build, and evolve observability platforms across metrics, logs, traces, alerting, and telemetry pipelines
  • Lead scalable observability solutions for GPU and AI infrastructure
  • Improve monitoring coverage, alert quality, service health visibility, and incident response
  • Develop reusable observability standards, frameworks, and patterns
  • Identify reliability risks and operational blind spots
  • Contribute to telemetry architecture and performance decisions
  • Mentor engineers through reviews and knowledge sharing
  • Participate in incident investigations and postmortems
  • Evaluate observability technologies and practices

Benefits

  • Bonus eligibility
  • Equity eligibility
  • Medical coverage
  • Dental coverage
  • Vision coverage
  • Flexible paid time off
  • Parental leave
  • Retirement plan participation
Staff Observability Platform Engineer at Nscale | JobStash