Senior GPU Cloud Kubernetes Expert SRE SME
Bitdeer is a technology company providing Bitcoin mining solutions.
Funding history
Investors
Projects
About Bitdeer
Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design, deploy, and operate production Kubernetes control planes for GPU workloads. You will manage GPU scheduling, tenant isolation, bare-metal provisioning, infrastructure as code, monitoring, incident response, and automated remediation. You will build workflows that detect faults, reschedule workloads, and improve reliability.
Requirements
- At least 5 years of Kubernetes operations experience
- At least 2 years managing GPU workloads on Kubernetes
- Knowledge of NVIDIA GPU Operator, device plugins, and GPU scheduling
- Experience with topology-aware scheduling and GPU resource management
- Experience building multi-tenant Kubernetes platforms
- Experience with bare-metal provisioning and lifecycle automation
- Proficiency in Terraform, Helm, GitOps, ArgoCD, or Flux
- SRE expertise in SLI/SLO frameworks, incident management, and capacity planning
- Experience with Prometheus, Grafana, and alerting at scale
- Go or Python programming skills for operator and CRD development
- Experience designing or implementing automated remediation
- Runbook-as-code mindset
Responsibilities
- Operate production Kubernetes clusters for GPU workloads
- Manage NVIDIA GPU Operator, device plugins, MIG configuration, and GPU time-slicing
- Implement topology-aware GPU scheduling
- Develop CRDs for GPU workload lifecycle management
- Integrate Slurm, Ray, and Kubeflow with Kubernetes
- Implement multi-tenant isolation through namespaces, network policies, quotas, RBAC, and pod security standards
- Automate bare-metal provisioning, tenant onboarding, lifecycle management, and reclamation
- Develop Terraform providers and modules
- Define and manage SLIs and SLOs
- Automate runbooks, escalations, and post-incident reviews
- Operate Prometheus, Grafana, Alertmanager, and PagerDuty
- Automate GPU-node fault detection, draining, cordoning, tainting, and workload rescheduling
- Enable safe automated remediation workflows
