Large Language Model Training Framework Engineer
Active AI foundation-model company offering multimodal models, consumer AI products, and an enterprise/developer API platform.
Funding history
About MiniMax Group Inc.
MiniMax develops proprietary multimodal foundation models spanning language, video, speech, and music, and operates AI-native products including MiniMax Agent, MiniMax Design, MiniMax Audio, Talkie, Hailuo AI, and an Open Platform.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and optimize large-scale model training frameworks. You will improve parallelism, communication, memory use, checkpointing, fault recovery, numerical correctness, stability, throughput, and resource utilization. You will also evaluate training-framework advances and apply them to production training workloads.
Requirements
- Strong programming and system-design skills
- Deep understanding of large language model training systems, distributed systems, parallel computing, GPU performance optimization, or numerical correctness
- Ability to independently diagnose system problems and deliver practical optimizations
- Ability to translate research papers and open-source framework ideas into engineering decisions
- Experience with PyTorch, Megatron-LM, DeepSpeed, Colossal-AI, or TorchTitan is preferred
- Experience with NCCL, RDMA, CUDA, Triton, communication optimization, memory optimization, or checkpoint optimization is preferred
Responsibilities
- Build and optimize large-scale training frameworks
- Optimize data, tensor, pipeline, and expert parallelism strategies
- Solve communication, memory, checkpointing, fault-recovery, numerical-correctness, and stability problems
- Analyze relationships among model architecture, training strategy, and system efficiency
- Apply advances from PyTorch, Megatron-LM, DeepSpeed, and TorchTitan to training workloads
