Post-Training Applied Researcher
BasetenVisit Baseten website
Baseten is an AI inference platform for deploying, optimizing, and scaling custom, open-source, and fine-tuned models in production.
San Francisco, United States
Funding history
About Baseten
Baseten provides model runtimes, inference infrastructure, developer workflows, and deployment options including managed cloud, self-hosted, and hybrid environments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will create post-training pipelines, reward functions, training environments, and evaluation harnesses for domain-specific LLMs. You will translate production data into training signals, analyze training experiments, address training instability, and publish research findings.
Requirements
- Hands-on experience training LLMs with reinforcement learning
- Understanding of GRPO or PPO, group advantage computation, clipped objectives, and KL penalty design
- Experience with reward engineering
- Experience building multi-turn agent environments with tool use
- Experience across dataset construction, training, evaluation, and deployment
- Experience with production ML systems
Responsibilities
- Design and run SFT, GRPO, DPO, RLVR, reward engineering, and synthetic-data pipelines
- Build task-specific training environments and evaluations
- Translate production data into training signals and reward loops
- Run and analyze end-to-end training experiments
- Diagnose reward hacking, importance-sampling drift, and advantage-estimation instabilities
- Publish findings and contribute to open-source training libraries
Benefits
- Equity
- Medical, dental, and vision insurance for U.S. employees and dependents
- Flexible PTO
- Company-wide Winter Break
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k) for U.S. employees
