K8 Site Reliability SME
Bitdeer Technologies Group is a technology company providing Bitcoin mining solutions, mining hardware, data-center infrastructure, and AI cloud services. It serves individual, institutional, and enterprise customers globally.
Funding history
Investors
Projects
About Bitdeer Technologies Group
Bitdeer provides vertically integrated Bitcoin mining and high-performance computing services. Its operations include mining equipment procurement and manufacturing, datacenter design and construction, equipment management, daily mining operations, cloud mining, and mining-related services. The company also offers AI cloud infrastructure and high-performance computing powered by NVIDIA GPUs for AI and machine-learning workloads, serving customers across global markets.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Design, deploy, and operate production Kubernetes control planes for GPU workloads, including GPU operators, topology-aware scheduling, workload lifecycle resources, tenant isolation, bare-metal provisioning, infrastructure as code, SLOs, monitoring, incident automation, and automated node failure recovery.
Requirements
- 5+ years in Kubernetes operations
- At least 2 years managing GPU workloads on Kubernetes
- Deep understanding of the Nvidia GPU operator, device plugin, and GPU scheduling
- Experience with topology-aware scheduling and GPU resource management
- Experience building multi-tenant Kubernetes platforms with strong isolation
- Experience with bare-metal provisioning and lifecycle automation using Ironic, MAAS, or custom systems
- Proficiency in Terraform, Helm, and GitOps workflows such as ArgoCD or Flux
- Strong SRE background in SLI/SLO frameworks, incident management, and capacity planning
- Experience with Prometheus, Grafana, and alerting at scale
- Strong Go or Python programming skills for operator and CRD development
- Experience designing or implementing automated Kubernetes remediation
- Runbook-as-code mindset
Responsibilities
- Operate production Kubernetes clusters for GPU workloads
- Configure the Nvidia GPU operator, device plugin, MIG, and GPU time slicing
- Implement topology-aware scheduling for GPU locality, NVLink domains, and network rails
- Develop CRDs for GPU workload lifecycle management
- Integrate Slurm, Ray, and Kubeflow with Kubernetes
- Implement multi-tenant isolation with namespaces, network policies, quotas, RBAC, and pod security
- Automate bare-metal provisioning, onboarding, lifecycle, and reclamation
- Build Terraform providers and modules for GPU cluster infrastructure
- Define and publish cluster availability, job completion, and provisioning latency SLOs
- Automate incident management and runbook execution
- Operate Prometheus, Grafana, Alertmanager, and PagerDuty monitoring
- Detect GPU node failures and automate drain, cordon, taint, and workload rescheduling
