Senior Software Engineer Inference Engine Platform Software
FuriosaAIVisit FuriosaAI website
FuriosaAI is a South Korean AI semiconductor company building energy-efficient inference accelerators, servers, and software for enterprise and cloud AI deployments.
Seoul, South Korea
Funding history
Investors
About FuriosaAI
FuriosaAI develops the RNGD AI inference accelerator and NXT RNGD Server, alongside a software toolchain for compiling, optimizing, and deploying LLM and agentic-AI workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design and implement a production inference engine for large and multimodal language models. You will optimize throughput, latency, memory use, scheduling, caching, distributed inference, and NPU execution while researching and integrating state-of-the-art inference techniques.
Requirements
- BS degree in Computer Science, Engineering, or a related field, with at least 3 years of relevant industry experience, or equivalent practical experience
- Proficiency in Rust or C++
- Knowledge of deep learning, LLMs, or generative AI models
- Excellent problem-solving and data analysis skills
- Strong communication and collaboration skills
- Experience building inference serving systems for large models preferred
- Deep understanding of performance optimization systems preferred
- Proficiency in C++, CUDA, or Triton kernel development preferred
- Contributions to vLLM, SGLang, or TensorRT-LLM preferred
Responsibilities
- Design and implement a next-generation inference engine for large and multimodal language models
- Design and implement advanced inference optimizations
- Develop distributed and scalable inference capabilities
- Collaborate with the Compiler team to optimize execution for NPUs
- Research, evaluate, and integrate inference optimization techniques and LLM-serving framework features
