Research Engineer Model Inference and Serving
H CompanyVisit H Company website
H Company develops computer-use AI agents, models, and enterprise automation products.
Paris, France
Funding history
About H Company
H Company is a Paris-founded AI company offering a full-stack platform for deploying action-oriented agents across enterprise systems, alongside Holo models, APIs, and browser/desktop agent tools.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and operate the inference stack for multimodal agentic models. You will improve latency, throughput, and serving cost; research and implement agent-focused inference techniques; evaluate serving and hardware platforms; and share findings with stakeholders.
Requirements
- Software engineering
- Python
- Rust, C++, or Go
- PyTorch or JAX
- Distributed systems
- Cloud environments
- Production ML infrastructure
- Kubernetes
- Transformers
- Multimodal architectures
- Research output, publications, research internships, or substantive open-source contributions
- Communication
- Presentation
- Collaboration
- Teamwork
Responsibilities
- Build and operate the inference stack for multimodal agentic models
- Improve model-serving latency, throughput, and cost
- Research and implement inference techniques for agent workloads
- Co-design training-time decisions that affect inference
- Integrate inference into agentic AI products with cross-functional partners
- Evaluate inference, serving, and hardware platforms and communicate findings
- Stay current with inference, model-serving, and accelerator technology
