Senior Observability Engineer

Verda is a Helsinki-headquartered full-stack AI cloud provider offering on-demand GPU compute, clusters, serverless containers, storage, APIs, and related AI infrastructure.

Helsinki, Finland
About Verda

Formerly DataCrunch, Verda operates European AI cloud infrastructure for the full AI lifecycle, from prototyping and distributed model training to scalable inference. It manages data-center infrastructure, hardware, cloud-platform services, and an in-house AI Lab.

View jobs by Verda

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design, scale, and operate shared metrics, logs, and tracing services across infrastructure and workloads. You will improve telemetry for servers, networks, Kubernetes, GPUs, and applications; automate deployments and configuration; build dashboards and alerts; support incident response; and document shared observability standards.

Requirements

  • Hands-on experience with Grafana and production telemetry systems
  • Experience with VictoriaMetrics, VictoriaLogs, or comparable metrics and logging backends
  • Strong Linux fundamentals and production Kubernetes experience
  • Experience monitoring cloud-native applications and distributed systems
  • Practical OpenTelemetry experience, including instrumentation, metrics, traces, structured logs, trace context propagation, and Collector pipelines
  • Experience with GitOps, CI/CD, and configuration management using Argo CD, GitLab CI/CD, Ansible, and SaltStack
  • Experience investigating production incidents and maintaining alerts
  • Ability to explain technical decisions and support adoption of observability practices

Responsibilities

  • Design, scale, and operate shared metrics, logs, and tracing services
  • Manage declarative Kubernetes deployments and automate host and VM configuration
  • Improve telemetry for hardware, Linux, networks, Kubernetes, GPU clusters, and AI workloads
  • Define OpenTelemetry instrumentation, context propagation, and Collector standards
  • Build dashboards and alerts, improve PagerDuty workflows, and participate in on-call rotations
  • Document observability standards and support their adoption
  • Evaluate AI-assisted monitoring and incident-investigation approaches

Benefits

  • Equity compensation
  • Local benefits