Senior AI Scheduling and Orchestration Engineer

Bitdeer is a technology company providing Bitcoin mining solutions.

Singapore, SG
About Bitdeer

Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design batch scheduling and admission-control systems for AI workloads. You will build scheduler integrations, optimize GPU-aware placement, implement GPU sharing and tenant isolation, improve scheduling reliability, and mentor junior engineers through design reviews.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
  • 6+ years of distributed systems engineering experience.
  • Hands-on expertise in Kubernetes scheduling frameworks and orchestrators.
  • Experience with AI workload execution patterns and distributed training frameworks.
  • Experience operating and scaling scheduling stacks in HPC or large-scale production cloud environments.
  • Knowledge of GPU hardware architectures and scheduling for distributed AI training and inference.
  • Experience with infrastructure automation and infrastructure as code.
  • Technical communication and leadership skills.

Responsibilities

  • Design and implement batch scheduling architectures supporting multi-node gang scheduling.
  • Develop cluster-wide admission control and job-queueing mechanisms.
  • Use Kubernetes Dynamic Resource Allocation and custom scheduler plugins for accelerator requests.
  • Architect topology-aware pod placement for NVLink and InfiniBand fabrics.
  • Implement GPU-sharing technologies and multi-tenancy isolation policies.
  • Integrate scheduling with bare-metal hardware and storage I/O patterns.
  • Improve reliability and scalability by resolving contention and deadlock scenarios.
  • Mentor junior engineers and conduct design reviews.