Senior SRE and Automation Engineer Customer Facing

Bitdeer is a technology company providing Bitcoin mining solutions.

Singapore, SG
About Bitdeer

Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own the reliability of customer-facing GPU cloud services. You will operate Kubernetes GPU clusters, manage tenant lifecycle and bare-metal provisioning, define service objectives, automate remediation, maintain observability, and communicate during customer incidents.

Requirements

  • 5+ years of SRE or cloud operations experience
  • At least 2 years operating GPU workloads at scale
  • Kubernetes operations and GPU workload management expertise
  • Experience with Nvidia GPU Operator, device plugins, MIG, time-slicing, and GPU scheduling
  • Experience with topology-aware scheduling and GPU resource management
  • Experience building multi-tenant cloud platforms with isolation guarantees
  • Customer-facing cloud service and SLA or SLO experience
  • Bare-metal provisioning and lifecycle automation experience
  • Terraform, Helm, and GitOps proficiency
  • Knowledge of SLI, SLO, SLA, error budgets, incident management, and capacity planning
  • Experience with Prometheus, Grafana, and alerting
  • Go or Python programming skills
  • AIOps and runbook-as-code experience

Responsibilities

  • Operate customer-facing GPU cloud services and production Kubernetes clusters
  • Manage Nvidia GPU Operator, device plugins, MIG, GPU time-slicing, and allocation policies
  • Implement topology-aware GPU scheduling
  • Manage tenant onboarding, quotas, isolation, offboarding, and reclamation
  • Automate bare-metal provisioning and tenant lifecycle management
  • Define and operate SLIs, SLOs, and SLAs
  • Manage incidents, customer communications, and post-incident reviews
  • Build tenant-aware monitoring, dashboards, and alerting
  • Automate GPU node failure detection, workload draining, and rescheduling
  • Manage infrastructure as code with Terraform, Helm, and GitOps
  • Maintain service documentation, tenant runbooks, capacity plans, and support tiering
  • Build self-service customer observability