Performance Engineer Containers Serverless
Verda (formerly DataCrunch) is a Helsinki-headquartered full-stack AI cloud provider offering on-demand GPU compute, self-service clusters, serverless containers, storage, and inference infrastructure.
Funding history
About Verda
Founded in Helsinki in 2020, Verda operates European AI cloud infrastructure across physical data centers, hardware, a cloud platform, and AI research. DataCrunch renamed to Verda in November 2025 without changing its service offering.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will profile and optimize containerized ML and AI workloads from image distribution and runtime startup through model loading and inference. You will tune storage and caching layers, benchmark real workloads, remove cross-stack bottlenecks, and communicate performance trade-offs through technical write-ups.
Requirements
- Production experience making ML/AI workloads measurably faster
- Knowledge of Linux storage and container runtime internals
- Experience with distributed or network filesystems and S3-compatible object storage
- Knowledge of small-object overhead, range-request patterns, eventual consistency, and multipart tuning
- Experience with model-serving runtimes such as vLLM and SGLang
- Knowledge of safetensors, GGUF, and sharded checkpoints
- Ability to reason across networking, filesystems, caches, container runtimes, and GPUs
- Systems programming experience with Go, Rust, or Python
- Experience with checkpoint and restore of CPU and GPU workloads
- Experience with eStargz, SOCI, Nydus, or ML weight and dataset caching layers
- Familiarity with RDMA, GPUDirect Storage, or NVMe-oF
- Experience with serverless GPU platforms, model registries, or Kubernetes-based ML infrastructure
- Contributions to open-source storage, ML runtime, container, or kernel projects
- Performance engineering experience on bare metal
Responsibilities
- Profile and optimize the end-to-end path for containerized ML/AI workloads, including image distribution, runtime startup, weight loading, and the inference hot path
- Design and tune storage layers between S3-compatible object stores and GPU nodes
- Drive improvements in time-to-first-token, training step time, and cold-start latency
- Benchmark and characterize real workloads and turn findings into platform changes
- Work across compute, networking, and platform teams to remove end-to-end bottlenecks
- Publish internal and occasional external write-ups about performance trade-offs
- Keep up to date with the evolving ML/AI ecosystem
Benefits
- Cash and equity compensation
- Fringe benefits
- Access to GPUs for testing
