Researcher, Post Training

AI platform for creating, deploying, and managing full-stack software through natural-language interaction.

Series CRecently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

Stockholm, Sweden
About Lovable

Lovable lets people describe an idea in plain language and collaboratively build production-grade software. Its platform includes hosting, authentication, payments, integrations, security features, and deployment infrastructure.

View jobs by Lovable

Skills

About the Role

Own the full post-training lifecycle, from data curation and training jobs through evaluation and deployment. Apply reinforcement learning, preference optimization, and supervised fine-tuning to improve code generation, user-intent reasoning, and agent reliability while building evaluation infrastructure and operating production-scale training systems.

Requirements

  • Have personally run post-training jobs on large language models using RFT/RLVR, preference optimization, or similar methods
  • Write solid production code
  • Be fluent in PyTorch or JAX
  • Be comfortable with distributed training setups and GPU clusters
  • Understand preference optimization, reward modeling, and alignment techniques
  • Have built or significantly contributed to evaluation systems that capture real-world quality
  • Be able to trace model-quality regressions through serving, inference, and training

Responsibilities

  • Own the full post-training lifecycle from data curation and training runs through evaluation and deployment
  • Apply and adapt reinforcement learning, preference optimization, and supervised fine-tuning methods
  • Build evaluation and experimentation infrastructure covering helpfulness, safety, latency, and reliability
  • Develop and operate production systems for large-scale training jobs, including GPU orchestration and data pipelines
  • Work with agent, product, and infrastructure engineers to turn model gains into product improvements
  • Investigate and resolve failures across training recipes, data, and serving
  • Read papers, run experiments, and move promising research into production quickly
Researcher, Post Training at Lovable | JobStash