Research Engineer Scientist Post Training and Reinforcement Learning
H CompanyVisit H Company website
H Company develops computer-use AI agents, models, and enterprise automation products.
Paris, France
Funding history
About H Company
H Company is a Paris-founded AI company offering a full-stack platform for deploying action-oriented agents across enterprise systems, alongside Holo models, APIs, and browser/desktop agent tools.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will develop and train advanced language and vision-language models, research training methods for instruction following and tool use, and optimize data pipelines and distributed training systems. You will integrate models into agentic AI systems, evaluate performance, communicate findings, and stay current with relevant research.
Requirements
- Python, Rust, or similar programming skills
- Software engineering
- PyTorch, JAX, or TensorFlow
- LLM training
- VLM training
- Supervised fine-tuning
- Direct preference optimization
- RLHF or RLVR
- Reward modeling
- Offline reinforcement learning
- Online reinforcement learning
Responsibilities
- Develop and train advanced LLMs and VLMs
- Research and implement training methods for instruction following and tool use
- Design and optimize data pipelines and training systems for large-scale distributed training
- Integrate models into agentic AI systems
- Evaluate model performance and communicate findings
- Stay current with LLM, VLM, and related research
