Researcher Post Training
AI app builder that turns natural-language prompts into full-stack web applications.
Maintainer signals as of 8/14/2026
About Lovable
Lovable (lovable.dev) is a Swedish AI software-builder platform: users describe an app in natural language and Lovable generates production full-stack code — one of the fastest-growing AI products.
Skills
About the Role
You will own the full post-training lifecycle, from curating data and running training jobs through evaluation and deployment. You will apply reinforcement learning, preference optimization, and supervised fine-tuning to improve code generation, user-intent reasoning, and agent reliability. You will build evaluation and experimentation infrastructure for helpfulness, safety, latency, and reliability; operate production-scale training systems; investigate failures across training, data, and serving; and turn promising research into production quickly.
Requirements
- Have personally run post-training jobs on large language models using RFT/RLVR, preference optimization, or similar methods
- Write solid production code
- Be fluent in PyTorch or JAX
- Be comfortable with distributed training setups and GPU clusters
- Understand preference optimization, reward modeling, and alignment techniques
- Have built or significantly contributed to evaluation systems that capture real-world quality
- Be able to trace model-quality regressions through serving, inference, and training
Responsibilities
- Own the full post-training lifecycle from data curation and training runs through evaluation and deployment
- Apply and adapt reinforcement learning, preference optimization, and supervised fine-tuning methods
- Build evaluation and experimentation infrastructure covering helpfulness, safety, latency, and reliability
- Develop and operate production systems for large-scale training jobs, including GPU orchestration and data pipelines
- Work with agent, product, and infrastructure engineers to turn model gains into product improvements
- Investigate and resolve failures across training recipes, data, and serving
- Read papers, run experiments, and move promising research into production quickly
