Research Engineer Model Inference and Serving

H Company develops computer-use AI agents, models, and enterprise automation products.

Paris, France

Funding history

About H Company

H Company is a Paris-founded AI company offering a full-stack platform for deploying action-oriented agents across enterprise systems, alongside Holo models, APIs, and browser/desktop agent tools.

View jobs by H Company

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build and operate the inference stack for multimodal agentic models. You will improve latency, throughput, and serving cost; research and implement agent-focused inference techniques; evaluate serving and hardware platforms; and share findings with stakeholders.

Requirements

  • Software engineering
  • Python
  • Rust, C++, or Go
  • PyTorch or JAX
  • Distributed systems
  • Cloud environments
  • Production ML infrastructure
  • Kubernetes
  • Transformers
  • Multimodal architectures
  • Research output, publications, research internships, or substantive open-source contributions
  • Communication
  • Presentation
  • Collaboration
  • Teamwork

Responsibilities

  • Build and operate the inference stack for multimodal agentic models
  • Improve model-serving latency, throughput, and cost
  • Research and implement inference techniques for agent workloads
  • Co-design training-time decisions that affect inference
  • Integrate inference into agentic AI products with cross-functional partners
  • Evaluate inference, serving, and hardware platforms and communicate findings
  • Stay current with inference, model-serving, and accelerator technology
Research Engineer Model Inference and Serving at H Company | JobStash