Algorithm AI System Engineer

FuriosaAI is a South Korean AI semiconductor company building energy-efficient inference accelerators, servers, and software for enterprise and cloud AI deployments.

Seoul, South Korea
About FuriosaAI

FuriosaAI develops the RNGD AI inference accelerator and NXT RNGD Server, alongside a software toolchain for compiling, optimizing, and deploying LLM and agentic-AI workloads.

View jobs by FuriosaAI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will prove core components of next-generation serving systems through NPU-based proof-of-concept implementations. You will explore and validate compression techniques, such as quantization and KV-cache compression, on NPUs. You will use implementation insights to propose research topics and optimization ideas, collaborating on productization after validation.

Requirements

  • At least 3 years of relevant professional experience or equivalent research or project experience
  • Understanding of LLM inference, including attention, KV cache, prefill, decode, and batching
  • Knowledge of or experience with AI inference systems such as vLLM, SGLang, or TensorRT-LLM
  • Experience with CUDA, Triton, or accelerator programming
  • Experience solving undocumented problems through experimentation and debugging
  • Ability to communicate technical complexity clearly and lead collaboration with other teams

Responsibilities

  • Prove core next-generation serving-system concepts through NPU-based proof-of-concept implementations
  • Explore, implement, and validate compression techniques in NPU environments
  • Identify and propose research topics and optimization ideas from implementation insights