Senior AI Platform Engineer
Bitdeer Technologies Group is a technology company providing Bitcoin mining solutions, mining hardware, data-center infrastructure, and AI cloud services. It serves individual, institutional, and enterprise customers globally.
Funding history
Investors
Projects
About Bitdeer Technologies Group
Bitdeer provides vertically integrated Bitcoin mining and high-performance computing services. Its operations include mining equipment procurement and manufacturing, datacenter design and construction, equipment management, daily mining operations, cloud mining, and mining-related services. The company also offers AI cloud infrastructure and high-performance computing powered by NVIDIA GPUs for AI and machine-learning workloads, serving customers across global markets.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Operate and harden Kubernetes-based MaaS production environments across CPU nodes, edge ingress, and regional GPU tiers. Own SLOs, observability, runbooks, incident response, rollout safety, capacity planning, automation, and incident debugging from the public API edge to model workers.
Requirements
- 6+ years in SRE, platform engineering, or infrastructure engineering for production cloud services
- Deep Kubernetes experience
- Experience with Helm, Argo CD, GitOps, CNI, ingress, secrets, storage, and workload scheduling
- GPU, AI infrastructure, or HPC workload experience strongly preferred
- Observability experience with Prometheus, VictoriaMetrics, OpenTelemetry, logs, and traces
- Go, Python, Bash, Linux networking, and production automation experience
- Ability to design reliable systems with SLOs and operational ownership
Responsibilities
- Operate and harden Kubernetes-based MaaS production environments
- Define and own SLOs, alerting, dashboards, runbooks, and incident response
- Improve rollout safety with canaries, fallback, health-aware routing, maintenance mode, and rollback
- Plan capacity for GPU utilization, burst traffic, quotas, rate limits, latency, and customer growth
- Automate operations with Helm, Argo CD, operators, scripts, and self-healing workflows
- Debug incidents from the public API edge to model workers
- Partner with runtime and performance engineers on incident resolution
Benefits
- Welfare benefits
- Inclusive work environment
- Training and mentoring
- Developmental opportunities
- Autonomy and fast growth opportunities
