Staff AI Scheduling & Orchestration Engineer
Bitdeer Technologies Group is a technology company providing Bitcoin mining solutions, mining hardware, data-center infrastructure, and AI cloud services. It serves individual, institutional, and enterprise customers globally.
Funding history
Investors
Projects
About Bitdeer Technologies Group
Bitdeer provides vertically integrated Bitcoin mining and high-performance computing services. Its operations include mining equipment procurement and manufacturing, datacenter design and construction, equipment management, daily mining operations, cloud mining, and mining-related services. The company also offers AI cloud infrastructure and high-performance computing powered by NVIDIA GPUs for AI and machine-learning workloads, serving customers across global markets.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead workload placement and scheduling logic for large-scale AI training and inference. You will build batch scheduling and queueing systems, develop accelerator-aware placement strategies, implement GPU sharing and tenant isolation, integrate scheduling with hardware and storage, resolve contention and deadlocks, and guide architecture through reviews and mentoring.
Requirements
- Bachelor’s or Master’s degree in Computer Science Electrical Engineering or a related field
- 6+ years of distributed systems engineering experience
- Deep expertise in Kubernetes scheduling frameworks and orchestrators
- Experience with AI workloads and PyTorch Distributed Ray or MPI
- Experience operating debugging and scaling production or HPC scheduling stacks
- Strong knowledge of GPU hardware architectures
- Experience with infrastructure automation and infrastructure-as-code
- Excellent technical communication and leadership skills
- Ability to translate ambiguous requirements into scalable engineering solutions
Responsibilities
- Design batch scheduling architectures with Volcano or YuniKorn
- Develop admission control and job queueing with Kueue
- Use Kubernetes DRA and custom scheduler plugins for accelerator requests
- Architect topology-aware placement for NVLink and InfiniBand
- Implement GPU sharing with MIG and time-slicing
- Define multi-tenancy isolation policies
- Integrate scheduling with GPU hardware and storage systems
- Resolve resource contention and deadlocks in HPC environments
- Mentor engineers and conduct design reviews
