K8 Site Reliability SME
Bitdeer is a NASDAQ-listed (BTDR) high-performance computing and Bitcoin mining company headquartered in Singapore. It provides end-to-end Bitcoin mining solutions (mining hardware like SEALMINER, Minerbase containers, cloud mining, hosting, and mining farm/data center operations) as well as AI Cloud services offering GPU compute (NVIDIA GB200 NVL72, B200, H200, H100) for AI training and deployment. Its customers range from individual and institutional Bitcoin miners to enterprises and developers needing scalable AI/ML compute infrastructure.
Funding history
Investors
About Bitdeer
Bitdeer Technologies Group is a NASDAQ-listed (ticker: BTDR) technology company headquartered in Singapore that describes itself as a "world-leading" high-performance computing platform and Bitcoin mining services provider. The company is vertically integrated across the value chain, spanning IC design and hardware manufacturing (its own SEALMINER ASIC miners and Minerbase mobile cooling containers), infrastructure construction and cloud mining, and artificial intelligence/high-performance computing. Bitdeer handles the full range of mining-related processes for its customers, including equipment procurement, transport logistics, datacenter design and construction, equipment management, and daily operations, and offers institutional services, a hash rate market, and a miner rights trading marketplace via its mobile apps (Bitdeer App and Minerplus App). Since 2013 Bitdeer has built more than 30 data centers globally and currently operates 9 large-scale data centers (including one of North America's largest) with roughly 3GW of diversified energy capacity and tens of exahashes of managed hash rate, with major operations in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia. Since 2023 it has also been expanding a global AI infrastructure business (Bitdeer AI Cloud), powered by thousands of NVIDIA GPUs (including GB200 NVL72 and B200, with GB300 NVL72 and B300 planned), offering turnkey AI datacenter solutions and GPU cloud compute for AI training and deployment starting at around $2/hour. Bitdeer serves both individual/retail Bitcoin miners and institutional clients, as well as AI developers and enterprises seeking scalable, energy-efficient compute.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You design, deploy, and operate production Kubernetes control planes for GPU workloads. You manage GPU operators, topology-aware scheduling, workload lifecycle resources, tenant isolation, bare-metal provisioning, infrastructure as code, SLOs, monitoring, incident automation, and automated node failure recovery.
Requirements
- 5+ years in Kubernetes operations
- At least 2 years managing GPU workloads on Kubernetes
- Deep understanding of the Nvidia GPU operator, device plugin, and GPU scheduling
- Experience with topology-aware scheduling and GPU resource management
- Experience building multi-tenant Kubernetes platforms with strong isolation
- Experience with bare-metal provisioning and lifecycle automation using Ironic, MAAS, or custom systems
- Proficiency in Terraform, Helm, and GitOps workflows such as ArgoCD or Flux
- Strong SRE background in SLI/SLO frameworks, incident management, and capacity planning
- Experience with Prometheus, Grafana, and alerting at scale
- Strong Go or Python programming skills for operator and CRD development
- Experience designing or implementing automated Kubernetes remediation
- Runbook-as-code mindset
Responsibilities
- Operate production Kubernetes clusters for GPU workloads
- Configure the Nvidia GPU operator, device plugin, MIG, and GPU time slicing
- Implement topology-aware scheduling for GPU locality, NVLink domains, and network rails
- Develop CRDs for GPU workload lifecycle management
- Integrate Slurm, Ray, and Kubeflow with Kubernetes
- Implement multi-tenant isolation with namespaces, network policies, quotas, RBAC, and pod security
- Automate bare-metal provisioning, onboarding, lifecycle, and reclamation
- Build Terraform providers and modules for GPU cluster infrastructure
- Define and publish cluster availability, job completion, and provisioning latency SLOs
- Automate incident management and runbook execution
- Operate Prometheus, Grafana, Alertmanager, and PagerDuty monitoring
- Detect GPU node failures and automate drain, cordon, taint, and workload rescheduling
