Senior GPU Cloud K8S Expert SRE SME
Bitdeer is a technology company providing Bitcoin mining solutions.
Funding history
Investors
Projects
About Bitdeer
Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.
Skills
About the Role
You will own production Kubernetes clusters for GPU workloads. You will manage GPU scheduling, tenant isolation, bare-metal provisioning, infrastructure as code, observability, incident automation, and automated remediation workflows.
Requirements
- 5+ years of Kubernetes operations experience
- At least 2 years managing GPU workloads on Kubernetes
- Knowledge of Nvidia GPU Operator, device plugins, and GPU scheduling
- Experience with topology-aware scheduling and GPU resource management
- Experience building multi-tenant Kubernetes platforms
- Experience with bare-metal provisioning and lifecycle automation
- Proficiency in Terraform, Helm, and GitOps workflows
- SRE experience with SLI/SLO frameworks, incident management, and capacity planning
- Experience with Prometheus, Grafana, and alerting at scale
- Programming skills in Go or Python for operator and CRD development
Responsibilities
- Operate production Kubernetes clusters optimized for GPU workloads
- Manage Nvidia GPU Operator, device plugins, MIG configuration, and GPU time-slicing
- Implement topology-aware GPU scheduling
- Develop CRDs for GPU workload lifecycle management
- Integrate Slurm, Ray, and Kubeflow with Kubernetes
- Implement multi-tenant isolation and security controls
- Automate bare-metal provisioning, tenant onboarding, lifecycle management, and reclamation
- Build Terraform providers and modules
- Define and manage cluster availability, job completion, and provisioning SLIs and SLOs
- Automate incident management, runbooks, and remediation workflows
- Operate Prometheus, Grafana, Alertmanager, and PagerDuty monitoring
- Automate GPU node failure detection and workload rescheduling
