Software Engineer Compiler Kernel Optimization
FuriosaAI is a South Korean AI semiconductor company building energy-efficient inference accelerators, servers, and software for enterprise and cloud AI deployments.
Funding history
Investors
About FuriosaAI
FuriosaAI develops the RNGD AI inference accelerator and NXT RNGD Server, alongside a software toolchain for compiling, optimizing, and deploying LLM and agentic-AI workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own the performance of critical AI kernels in Furiosa-LLM. You will analyze serving workloads, identify bottlenecks, implement optimized TCL kernels, and develop scheduling and algorithmic techniques for dynamic workloads. You will integrate, benchmark, and validate kernels across models and serving scenarios.
Requirements
- Experience developing low-level or performance-critical software
- Experience analyzing performance bottlenecks using profiling, benchmarking, and hardware performance characteristics
- Understanding of parallel computation, memory hierarchies, and data movement on modern architectures
Responsibilities
- Analyze end-to-end LLM serving workloads and identify kernel-level performance bottlenecks
- Design, implement, and optimize high-performance TCL kernels
- Develop algorithmic techniques for dynamic serving workloads
- Integrate, benchmark, and validate optimized kernels across models and serving scenarios
- Collaborate with compiler and serving teams to improve compiler capabilities
