Principal Observability Platform Engineer

30 minutes agoLeadSalary: 190K - 300KUnited StatesDevopsJobs by Nscale

Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.

London, United Kingdom
About Nscale

Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.

View jobs by Nscale

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will define the technical strategy and architecture for observability across metrics, logs, traces, and alerting. You will make platform decisions, identify systemic gaps, establish engineering standards, mentor engineers, lead postmortems, and evaluate tools that improve signal quality, efficiency, and scalability.

Requirements

  • 8+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles
  • Experience operating observability infrastructure at scale
  • Hands-on experience with significant observability tooling
  • Python, Go, or similar engineering proficiency
  • Kubernetes experience at scale
  • Infrastructure-as-Code experience with Terraform, Ansible, or equivalent
  • Ability to architect systems, write code, review work, and communicate technical tradeoffs

Responsibilities

  • Own observability strategy and architecture across metrics, logs, traces, and alerting
  • Drive platform decisions on tooling, data models, ingestion, retention, and cardinality
  • Identify systemic gaps and design systems that expose failures
  • Partner with SRE, infrastructure, and AI/ML teams
  • Define observability standards and patterns
  • Mentor and technically develop observability engineers
  • Lead incident postmortems and durable platform improvements
  • Evaluate, introduce, and retire observability tooling

Benefits

  • Medical insurance
  • Dental insurance
  • Vision insurance
  • Flexible paid time off
  • Parental leave
  • Retirement plan participation