Agent Post-Training Artifacts Research

1 day agoSalary: 380K - 500KSan Francisco, USAResearchJobs by OpenAI

AI research and deployment company building and deploying frontier AI products, including ChatGPT, Codex, and its API platform.

Recently funded151 current maintainers88 active leads20 new active leads99 lead step-downs12 early lead departuresTeam intelligence

Maintainer signals as of 9/25/2026

San Francisco, United States
About OpenAI

OpenAI’s mission is to ensure artificial general intelligence benefits all of humanity. It consists of the nonprofit OpenAI Foundation and OpenAI Group, a public benefit corporation governed by the Foundation.

View jobs by OpenAI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will train frontier models to create documents, spreadsheets, slide decks, dashboards, reports, analyses, and other editable artifacts. You will own post-training improvements across reinforcement learning, data pipelines, graders, reward signals, evaluations, diagnostics, behavioral analysis, and product integration.

Requirements

  • Technical fundamentals in machine learning, software engineering, systems, statistics, or a related field
  • Experience with LLMs, reinforcement learning, RLHF, RLAIF, post-training, evaluations, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems
  • Ability to define hypotheses, build pipelines, run models, and analyze results
  • Ability to work across research, product, infrastructure, data, evaluation, and safety functions
  • Background in consulting, finance, marketing, operations, or data science

Responsibilities

  • Design and run experiments to improve agentic model behavior for complex software and plugins
  • Own post-training improvements across reinforcement learning, data pipelines, graders, reward signals, and evaluations
  • Build evaluations and environments that expose model failures
  • Turn model failures into training data, product fixes, or research directions
  • Translate product signals into model improvements
  • Develop early-training and alignment interventions
  • Improve training and launch reliability, observability, reproducibility, cost, and latency
  • Debug model failures and develop experiments and fixes

Benefits

  • Medical, dental, and vision insurance with employer Health Savings Account contributions
  • Pre-tax health, dependent-care, parking, and transit accounts
  • 401(k) retirement plan with employer match
  • Paid parental, medical, and caregiver leave
  • Flexible PTO or up to 15 days of annual PTO depending on employee classification
  • Paid holidays, company office closures, and sick or safe time
  • Mental health and wellness support
  • Employer-paid basic life and disability coverage
  • Daily office meals and eligible meal-delivery credits
  • Relocation support for eligible employees
  • Charitable donation matching and wellness stipends may be provided