Post-Training Applied Researcher

Baseten is an AI inference platform for deploying, optimizing, and scaling custom, open-source, and fine-tuned models in production.

San Francisco, United States
About Baseten

Baseten provides model runtimes, inference infrastructure, developer workflows, and deployment options including managed cloud, self-hosted, and hybrid environments.

View jobs by Baseten

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will create post-training pipelines, reward functions, training environments, and evaluation harnesses for domain-specific LLMs. You will translate production data into training signals, analyze training experiments, address training instability, and publish research findings.

Requirements

  • Hands-on experience training LLMs with reinforcement learning
  • Understanding of GRPO or PPO, group advantage computation, clipped objectives, and KL penalty design
  • Experience with reward engineering
  • Experience building multi-turn agent environments with tool use
  • Experience across dataset construction, training, evaluation, and deployment
  • Experience with production ML systems

Responsibilities

  • Design and run SFT, GRPO, DPO, RLVR, reward engineering, and synthetic-data pipelines
  • Build task-specific training environments and evaluations
  • Translate production data into training signals and reward loops
  • Run and analyze end-to-end training experiments
  • Diagnose reward hacking, importance-sampling drift, and advantage-estimation instabilities
  • Publish findings and contribute to open-source training libraries

Benefits

  • Equity
  • Medical, dental, and vision insurance for U.S. employees and dependents
  • Flexible PTO
  • Company-wide Winter Break
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Company-facilitated 401(k) for U.S. employees