Research Engineer Scientist Post Training and Reinforcement Learning

H Company develops computer-use AI agents, models, and enterprise automation products.

Paris, France

Funding history

About H Company

H Company is a Paris-founded AI company offering a full-stack platform for deploying action-oriented agents across enterprise systems, alongside Holo models, APIs, and browser/desktop agent tools.

View jobs by H Company

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will develop and train advanced language and vision-language models, research training methods for instruction following and tool use, and optimize data pipelines and distributed training systems. You will integrate models into agentic AI systems, evaluate performance, communicate findings, and stay current with relevant research.

Requirements

  • Python, Rust, or similar programming skills
  • Software engineering
  • PyTorch, JAX, or TensorFlow
  • LLM training
  • VLM training
  • Supervised fine-tuning
  • Direct preference optimization
  • RLHF or RLVR
  • Reward modeling
  • Offline reinforcement learning
  • Online reinforcement learning

Responsibilities

  • Develop and train advanced LLMs and VLMs
  • Research and implement training methods for instruction following and tool use
  • Design and optimize data pipelines and training systems for large-scale distributed training
  • Integrate models into agentic AI systems
  • Evaluate model performance and communicate findings
  • Stay current with LLM, VLM, and related research
Research Engineer Scientist Post Training and Reinforcement Learning at H Company | JobStash