Senior Observability Platform Engineer

49 minutes agoSeniorSalary: 160K - 230KUnited StatesDevopsJobs by Nscale

Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.

London, United Kingdom
About Nscale

Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.

View jobs by Nscale

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design, build, and operate scalable observability systems for metrics, logs, traces, and alerting. You will improve signal quality, contribute to architecture and data-retention decisions, identify observability gaps, and integrate observability into services and platforms. You will also support incident response, evaluate tooling, develop reusable patterns, and mentor engineers through reviews and knowledge sharing.

Requirements

  • 5+ years of experience in SRE, infrastructure engineering, platform engineering, or observability-focused roles
  • Experience operating and scaling production observability systems
  • Knowledge of metrics, logs, traces, alerting, and SLOs
  • Experience with observability technologies including Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, or Elastic
  • Programming skills in Python, Go, or similar languages
  • Experience with Kubernetes-based infrastructure
  • Familiarity with Infrastructure as Code tools such as Terraform or Ansible

Responsibilities

  • Design, build, and operate scalable observability systems across metrics, logs, traces, and alerting
  • Contribute to tooling, data-pipeline, storage, and retention architecture
  • Improve signal quality by reducing noise, managing cardinality, and refining alerting
  • Identify and address observability gaps before they affect reliability
  • Integrate observability into services and platforms with SRE, infrastructure, and AI/ML stakeholders
  • Develop reusable patterns, libraries, and best practices
  • Participate in incident response and postmortems
  • Evaluate and adopt tools that improve developer experience, scalability, and operational efficiency
  • Mentor engineers through code reviews and knowledge sharing

Benefits

  • Medical insurance
  • Dental insurance
  • Vision insurance
  • Flexible paid time off
  • Parental leave
  • Retirement plan participation
Senior Observability Platform Engineer at Nscale | JobStash