Senior Observability Engineer
Verda is a Helsinki-headquartered full-stack AI cloud provider offering on-demand GPU compute, clusters, serverless containers, storage, APIs, and related AI infrastructure.
Funding history
About Verda
Formerly DataCrunch, Verda operates European AI cloud infrastructure for the full AI lifecycle, from prototyping and distributed model training to scalable inference. It manages data-center infrastructure, hardware, cloud-platform services, and an in-house AI Lab.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design, scale, and operate shared metrics, logs, and tracing services across infrastructure and workloads. You will improve telemetry for servers, networks, Kubernetes, GPUs, and applications; automate deployments and configuration; build dashboards and alerts; support incident response; and document shared observability standards.
Requirements
- Hands-on experience with Grafana and production telemetry systems
- Experience with VictoriaMetrics, VictoriaLogs, or comparable metrics and logging backends
- Strong Linux fundamentals and production Kubernetes experience
- Experience monitoring cloud-native applications and distributed systems
- Practical OpenTelemetry experience, including instrumentation, metrics, traces, structured logs, trace context propagation, and Collector pipelines
- Experience with GitOps, CI/CD, and configuration management using Argo CD, GitLab CI/CD, Ansible, and SaltStack
- Experience investigating production incidents and maintaining alerts
- Ability to explain technical decisions and support adoption of observability practices
Responsibilities
- Design, scale, and operate shared metrics, logs, and tracing services
- Manage declarative Kubernetes deployments and automate host and VM configuration
- Improve telemetry for hardware, Linux, networks, Kubernetes, GPU clusters, and AI workloads
- Define OpenTelemetry instrumentation, context propagation, and Collector standards
- Build dashboards and alerts, improve PagerDuty workflows, and participate in on-call rotations
- Document observability standards and support their adoption
- Evaluate AI-assisted monitoring and incident-investigation approaches
Benefits
- Equity compensation
- Local benefits
