AI Platform Development Engineer
Beijing-based enterprise AI company building foundation-model, agentic-AI, and AI-transformation products, including the TrueNorth enterprise decision hub.
Funding history
Investors
About 01.AI
01.AI develops full-stack enterprise AI solutions, industry-agent applications, and sovereign/industry model-training capabilities.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build a multi-tenant GPU compute scheduling platform on Kubernetes, supporting elastic scaling and task-level management. You will develop topology-aware GPU scheduling, explore GPU partitioning approaches, and build high-concurrency online inference platforms. You will also improve end-to-end observability, reduce compute costs through Spot instances and autoscaling, and establish incident diagnosis, response, and root-cause analysis practices.
Requirements
- Proficiency in Golang and Python
- Knowledge of Kubernetes and Docker
- Familiarity with at least one training or inference framework, including PyTorch, Megatron, DeepSpeed, vLLM, or SGLang
- Understanding of GPU architecture, CUDA, and heterogeneous computing
- Knowledge of RDMA, NVLink, or InfiniBand is a plus
- Experience with Kubernetes Scheduler Framework, Volcano, Kueue, PD separation, KV Cache, or AI Agent infrastructure is a plus
Responsibilities
- Build a multi-tenant GPU compute scheduling platform on Kubernetes
- Implement topology-aware GPU scheduling for NVLink and PCIe
- Explore MIG and time-slicing GPU partitioning approaches
- Build a high-concurrency large-model online inference platform
- Develop observability across GPU, network, storage, and inference
- Implement Spot instances and automatic scaling to reduce compute costs
- Establish fault diagnosis, incident response, and root-cause analysis processes
