Software Engineer AI Model Enablement and Inference

FuriosaAI is a South Korean AI semiconductor company building energy-efficient inference accelerators, servers, and software for enterprise and cloud AI deployments.

Seoul, South Korea
About FuriosaAI

FuriosaAI develops the RNGD AI inference accelerator and NXT RNGD Server, alongside a software toolchain for compiling, optimizing, and deploying LLM and agentic-AI workloads.

View jobs by FuriosaAI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will analyze model architectures, implement and optimize TCL kernels, and integrate models into Furiosa-LLM. You will validate correctness on NPUs, automate benchmarking and regression testing, evaluate inference-framework techniques, and provide technical guidance on model and inference trade-offs.

Requirements

  • Deep understanding of transformer-based LLMs, including attention variants, mixture-of-experts architectures, and KV-cache behavior
  • Strong Python skills and experience reading, modifying, and debugging model implementations in PyTorch or a comparable framework
  • Experience implementing, debugging, and optimizing tensor operations or accelerator kernels
  • Understanding of LLM inference performance, including prefill, decode, batching, and latency-throughput trade-offs
  • Ability to read and reason about Rust or C++ code
  • Technical communication and cross-team collaboration skills

Responsibilities

  • Analyze model architectures, algorithms, and reference implementations to identify implementation requirements and trade-offs
  • Design, implement, and optimize model-specific TCL kernels
  • Integrate models and TCL kernels into Furiosa-LLM
  • Build reusable analysis and integration tools and automate validation and benchmarking
  • Validate kernel and model correctness on NPUs and build regression tests
  • Evaluate and apply relevant techniques from vLLM and SGLang
  • Study state-of-the-art models and provide technical guidance on model capabilities and inference requirements
Software Engineer AI Model Enablement and Inference at FuriosaAI | JobStash