Senior Software Engineer Runtime
FuriosaAIVisit FuriosaAI website
FuriosaAI is a South Korean AI semiconductor company building energy-efficient inference accelerators, servers, and software for enterprise and cloud AI deployments.
Seoul, South Korea
Funding history
Investors
About FuriosaAI
FuriosaAI develops the RNGD AI inference accelerator and NXT RNGD Server, alongside a software toolchain for compiling, optimizing, and deploying LLM and agentic-AI workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will develop the low-level runtime stack for NPU hardware, including DMA I/O, execution scheduling, asynchronous pipelines, multi-node communication, and embedded firmware. You will profile and tune the complete runtime stack to remove bottlenecks in inference workloads.
Requirements
- BS degree in Computer Science, Engineering, or a related field, or equivalent practical experience
- 5+ years of relevant systems-programming experience using Rust, C, or C++
- Understanding of computer architecture, memory hierarchy, cache coherency, operating systems, DMA, interrupts, and MMIO
- Strong communication skills and ability to gather requirements and drive technical alignment
- Expertise in low-latency runtime systems, embedded firmware development, or high-performance I/O preferred
- Experience with asynchronous execution models and scheduling systems preferred
- Experience with DMA engines, scatter-gather I/O, or zero-copy data transfer preferred
- Experience developing embedded firmware for ARM-based processors preferred
- Familiarity with RDMA and high-performance networking preferred
- Experience with CUDA runtime internals preferred
- Experience with Linux kernel modules, eBPF, perf, or ftrace preferred
- Understanding of deep learning inference workloads preferred
- Experience profiling and tuning system software on accelerator or SoC platforms preferred
Responsibilities
- Develop low-level runtime software for DMA-based I/O and kernel execution scheduling
- Build and optimize asynchronous execution pipelines
- Enable multi-node inference through communication primitives and RDMA-based data transfer
- Develop embedded firmware for the NPU integrated ARM core
- Profile and tune system-level performance across the runtime stack
