Senior GPU Cloud K8S Expert SRE SME

Bitdeer is a technology company providing Bitcoin mining solutions.

Singapore, SG
About Bitdeer

Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.

View jobs by Bitdeer

Skills

About the Role

You will own production Kubernetes clusters for GPU workloads. You will manage GPU scheduling, tenant isolation, bare-metal provisioning, infrastructure as code, observability, incident automation, and automated remediation workflows.

Requirements

  • 5+ years of Kubernetes operations experience
  • At least 2 years managing GPU workloads on Kubernetes
  • Knowledge of Nvidia GPU Operator, device plugins, and GPU scheduling
  • Experience with topology-aware scheduling and GPU resource management
  • Experience building multi-tenant Kubernetes platforms
  • Experience with bare-metal provisioning and lifecycle automation
  • Proficiency in Terraform, Helm, and GitOps workflows
  • SRE experience with SLI/SLO frameworks, incident management, and capacity planning
  • Experience with Prometheus, Grafana, and alerting at scale
  • Programming skills in Go or Python for operator and CRD development

Responsibilities

  • Operate production Kubernetes clusters optimized for GPU workloads
  • Manage Nvidia GPU Operator, device plugins, MIG configuration, and GPU time-slicing
  • Implement topology-aware GPU scheduling
  • Develop CRDs for GPU workload lifecycle management
  • Integrate Slurm, Ray, and Kubeflow with Kubernetes
  • Implement multi-tenant isolation and security controls
  • Automate bare-metal provisioning, tenant onboarding, lifecycle management, and reclamation
  • Build Terraform providers and modules
  • Define and manage cluster availability, job completion, and provisioning SLIs and SLOs
  • Automate incident management, runbooks, and remediation workflows
  • Operate Prometheus, Grafana, Alertmanager, and PagerDuty monitoring
  • Automate GPU node failure detection and workload rescheduling