AI Platform Development Engineer

Beijing-based enterprise AI company building foundation-model, agentic-AI, and AI-transformation products, including the TrueNorth enterprise decision hub.

Beijing, China

Funding history

About 01.AI

01.AI develops full-stack enterprise AI solutions, industry-agent applications, and sovereign/industry model-training capabilities.

View jobs by 01.AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build a multi-tenant GPU compute scheduling platform on Kubernetes, supporting elastic scaling and task-level management. You will develop topology-aware GPU scheduling, explore GPU partitioning approaches, and build high-concurrency online inference platforms. You will also improve end-to-end observability, reduce compute costs through Spot instances and autoscaling, and establish incident diagnosis, response, and root-cause analysis practices.

Requirements

  • Proficiency in Golang and Python
  • Knowledge of Kubernetes and Docker
  • Familiarity with at least one training or inference framework, including PyTorch, Megatron, DeepSpeed, vLLM, or SGLang
  • Understanding of GPU architecture, CUDA, and heterogeneous computing
  • Knowledge of RDMA, NVLink, or InfiniBand is a plus
  • Experience with Kubernetes Scheduler Framework, Volcano, Kueue, PD separation, KV Cache, or AI Agent infrastructure is a plus

Responsibilities

  • Build a multi-tenant GPU compute scheduling platform on Kubernetes
  • Implement topology-aware GPU scheduling for NVLink and PCIe
  • Explore MIG and time-slicing GPU partitioning approaches
  • Build a high-concurrency large-model online inference platform
  • Develop observability across GPU, network, storage, and inference
  • Implement Spot instances and automatic scaling to reduce compute costs
  • Establish fault diagnosis, incident response, and root-cause analysis processes