Machine Learning Researcher
Inference is a distributed GPU network for running AI models efficiently across a global infrastructure.
Funding
About Inference
Inference is a distributed GPU network for running AI models. Users can connect their devices to contribute computing power, access APIs to perform model inference, and monitor workloads through a dashboard. The platform offers an open infrastructure that helps scale machine learning workloads without relying on centralized servers.
Skills
About the Role
You will conduct research into experimental models, training systems, and modalities to create novel products for customers. Your work spans exploring new architectures and learning methods to optimizing latency and efficiency, all with the goal of delivering better models to customers. Your north star is pushing the frontier of what's possible in LLM post-training. You'll explore new techniques, run rigorous experiments, and when something works, help bring it into production with your teammates. This includes training models for customers and running evaluations as part of validating your research. You'll report directly to the founding team and will have autonomy, a large compute budget/GPU reservation, and technical support to explore ambitious ideas and ship the ones that work.
Requirements
- 3+ years of experience training AI models using PyTorch
- Deep understanding of transformer architectures, attention mechanisms, and model internals
- Hands-on experience with post-training LLMs using SFT, RLHF, DPO, or other alignment techniques
- Experience with LLM-specific training frameworks (e.g., Hugging Face Transformers, DeepSpeed, Megatron, TRL, or similar)
- Strong experimental methodology, including ability to design, run, and analyze rigorous experiments
- Track record of implementing ideas from recent ML papers
- Experience training on NVIDIA GPUs at scale
- Strong foundation in ML fundamentals: optimization, loss functions, regularization, generalization
Responsibilities
- Research and experiment with new model architectures to improve quality, efficiency, or capability
- Explore methods to decrease inference latency and improve serving efficiency
- Run experiments with new learning methods, including novel approaches to SFT, RLHF, DPO, and other post-training techniques
- Perform reinforcement learning research to improve model alignment and capability
- Develop and improve the distillation pipeline for training high-quality models from frontier teachers
- Train models for clients and run evaluations to validate research findings in production settings
- Create robust benchmarks and evaluation frameworks that ensure custom models match or exceed frontier performance
- Stay current with ML research and identify techniques that can improve the platform
- Collaborate with applied engineers to bring successful research into production systems
- Document findings and share knowledge with the team
Benefits
- Equity
- Comprehensive benefits
