ML Inference Performance Engineer Intern

European AI-semiconductor company building AI inference accelerators and software for embedded, edge, server, and data-center deployments.

Eindhoven, Netherlands
About Axelera AI

Axelera AI develops purpose-built AI acceleration hardware and the Voyager SDK for computer-vision, generative-AI, and other inference workloads, emphasizing power- and thermally-constrained deployments.

View jobs by Axelera AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will benchmark ML inference systems, improve evaluation tooling, analyse hardware and software bottlenecks, evaluate accelerator platforms, maintain lab infrastructure, and turn performance measurements into structured reports and dashboards.

Requirements

  • Enrollment in the final years of a Bachelor's programme or a Master's programme in a related field
  • Python development experience
  • C/C++ knowledge
  • Experience with end-to-end computer vision pipelines
  • Familiarity with benchmarking concepts
  • Experience with inference tools, APIs, or SDKs
  • Familiarity with deep learning model concepts including quantization, ONNX, and PyTorch
  • Development experience using agentic AI
  • Git knowledge
  • Familiarity with LLM benchmarking
  • Linux, Bash, and Docker proficiency
  • Hands-on experience with embedded hosts
  • English communication skills

Responsibilities

  • Develop and improve benchmarking tools and reproducible procedures
  • Maintain a performance-visualisation dashboard
  • Evaluate AI accelerator products and their SDKs and toolchains
  • Track model support across platforms
  • Characterise end-to-end inference pipelines
  • Maintain equivalent pipeline configurations across platforms
  • Set up and maintain lab hosts and evaluation hardware
  • Produce structured reports for engineering and roadmap decisions

Benefits

  • Pension plan
  • Employee insurances
  • Company share option