Algorithm AI System Engineer
FuriosaAI is a South Korean AI semiconductor company building energy-efficient inference accelerators, servers, and software for enterprise and cloud AI deployments.
Funding history
Investors
About FuriosaAI
FuriosaAI develops the RNGD AI inference accelerator and NXT RNGD Server, alongside a software toolchain for compiling, optimizing, and deploying LLM and agentic-AI workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will prove core components of next-generation serving systems through NPU-based proof-of-concept implementations. You will explore and validate compression techniques, such as quantization and KV-cache compression, on NPUs. You will use implementation insights to propose research topics and optimization ideas, collaborating on productization after validation.
Requirements
- At least 3 years of relevant professional experience or equivalent research or project experience
- Understanding of LLM inference, including attention, KV cache, prefill, decode, and batching
- Knowledge of or experience with AI inference systems such as vLLM, SGLang, or TensorRT-LLM
- Experience with CUDA, Triton, or accelerator programming
- Experience solving undocumented problems through experimentation and debugging
- Ability to communicate technical complexity clearly and lead collaboration with other teams
Responsibilities
- Prove core next-generation serving-system concepts through NPU-based proof-of-concept implementations
- Explore, implement, and validate compression techniques in NPU environments
- Identify and propose research topics and optimization ideas from implementation insights
