Senior SRE & Automation Engineer Customer Facing

Bitdeer is a technology company providing Bitcoin mining solutions.

Singapore, SG
About Bitdeer

Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own reliability for a customer facing GPU cloud, operate large Kubernetes clusters, manage GPU allocation and topology aware scheduling, automate tenant onboarding and bare metal lifecycle, define SLIs SLOs and SLAs, handle incidents and communications, build observability and alerting, automate GPU failure recovery, manage infrastructure as code, and deliver customer self service tools.

Requirements

  • 5+ years of SRE or cloud operations experience
  • At least 2 years operating GPU workloads at scale
  • Deep understanding of Kubernetes operations and GPU workload management
  • Experience with topology aware scheduling and GPU resource management
  • Experience building multi tenant cloud platforms with strong isolation
  • Customer facing cloud service experience with SLAs SLOs tenant incidents and communications
  • Experience with bare metal provisioning and lifecycle automation
  • Proficiency in Terraform Helm and ArgoCD or Flux
  • Strong SRE background in incident management and capacity planning
  • Experience with Prometheus Grafana and alerting at scale
  • Strong programming skills in Go or Python
  • AIOps aptitude and a runbook as code mindset

Responsibilities

  • Own availability job completion provisioning latency and tenant experience
  • Operate production Kubernetes clusters for GPU workloads
  • Manage NVIDIA GPU Operator device plugins MIG time slicing and GPU allocation
  • Implement topology aware GPU scheduling
  • Automate tenant onboarding quota management isolation and offboarding
  • Automate bare metal provisioning lifecycle and reclamation
  • Define and publish SLIs SLOs SLAs and error budget priorities
  • Manage incident response customer updates and post incident reviews
  • Build tenant aware monitoring dashboards and alerting
  • Automate GPU node failure handling and workload rescheduling
  • Manage Terraform Helm and GitOps workflows
  • Deliver service documentation capacity planning and tenant self service observability