Staff AI Scheduling & Orchestration Engineer
Bitdeer Technologies Group is a technology company providing Bitcoin mining solutions, mining hardware, data-center infrastructure, and AI cloud services. It serves individual, institutional, and enterprise customers globally.
Funding history
Investors
Projects
About Bitdeer Technologies Group
Bitdeer provides vertically integrated Bitcoin mining and high-performance computing services. Its operations include mining equipment procurement and manufacturing, datacenter design and construction, equipment management, daily mining operations, cloud mining, and mining-related services. The company also offers AI cloud infrastructure and high-performance computing powered by NVIDIA GPUs for AI and machine-learning workloads, serving customers across global markets.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Design and implement advanced batch scheduling architectures for multi-node AI workloads, develop admission control and job queueing mechanisms, manage accelerator requests and topology-aware placement, implement GPU sharing and multi-tenancy isolation, integrate scheduling with hardware and storage systems, improve reliability and scalability, resolve resource contention, and mentor engineers.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field
- 6+ years of distributed systems engineering experience
- Hands-on expertise with Kubernetes scheduling frameworks and orchestrators
- Experience with AI workload execution patterns and distributed training frameworks such as PyTorch Distributed, Ray, or MPI
- Experience operating, debugging, and scaling scheduling stacks in HPC or large-scale production cloud environments
- Knowledge of GPU hardware architectures and distributed AI training and inference scheduling
- Experience with Terraform and Go-based Operators
- Strong technical communication and leadership skills
- Ability to translate complex, ambiguous requirements into scalable engineering solutions
Responsibilities
- Design and implement batch scheduling architectures using Volcano or YuniKorn
- Develop cluster-wide admission control and job queueing mechanisms with Kueue
- Use Kubernetes Dynamic Resource Allocation and custom scheduler plugins for accelerator requests
- Architect topology-aware pod placement for NVLink and InfiniBand
- Implement GPU sharing and multi-tenancy isolation policies
- Integrate the scheduling layer with bare-metal hardware and storage systems
- Improve scheduling-stack reliability and scalability
- Resolve resource contention and deadlock scenarios
- Mentor junior engineers and conduct design reviews
Benefits
- Welfare benefits
- Inclusive and respectful work environment
- Opportunity to network with industry pioneers
- Direct opportunity to impact the digital asset industry
- Personal accountability, autonomy, fast growth, and learning opportunities
- Training and mentoring opportunities
