Operational Data and Observability Engineer
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Maintainer signals as of 9/23/2026
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design monitoring, logging, tracing, alerts, dashboards, and SLOs. You will manage telemetry and operational-data pipelines, troubleshoot production incidents, support on-call response, administer observability tools, automate deployments, and maintain documentation and runbooks.
Requirements
- 3+ years of experience in DevOps, SRE, operations engineering, platform engineering, or observability engineering
- Experience with Prometheus, Grafana, Datadog, New Relic, or equivalent monitoring platforms
- Experience with ELK/Elastic Stack, Splunk, CloudWatch, or similar logging platforms
- Proficiency in Python, Go, Bash, or equivalent
- Knowledge of metrics, logging, distributed tracing, and application performance monitoring
- Experience with AWS, Azure, or Google Cloud Platform and Kubernetes or container orchestration
- Knowledge of application, infrastructure, networking, database, and storage performance monitoring
Responsibilities
- Design and implement observability strategies across infrastructure, services, and applications
- Develop dashboards, alerts, and service-level objectives
- Build centralized logging and distributed-tracing capabilities
- Deploy and maintain metrics, logs, events, and telemetry collection systems
- Design operational-data pipelines, APIs, and integrations
- Troubleshoot production issues and participate in on-call incident response
- Create runbooks, troubleshooting guides, and operational documentation
- Administer and enhance observability platforms
- Automate monitoring deployments, instrumentation, and platform configuration
- Maintain and upgrade observability infrastructure
Benefits
- Medical insurance
- Dental insurance
- Vision insurance
- Flexible paid time off
- Parental leave
- Retirement plan participation
- Hybrid or remote work arrangements
