Software Engineer LLM Performance and Evaluation
FuriosaAIVisit FuriosaAI website
FuriosaAI is a South Korean AI semiconductor company building energy-efficient inference accelerators, servers, and software for enterprise and cloud AI deployments.
Seoul, South Korea
Funding history
Investors
About FuriosaAI
FuriosaAI develops the RNGD AI inference accelerator and NXT RNGD Server, alongside a software toolchain for compiling, optimizing, and deploying LLM and agentic-AI workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and maintain LLM benchmarks, evaluation pipelines, dashboards, and reports. You will measure latency, throughput, memory use, utilization, and model accuracy; automate reproducible evaluations; investigate regressions; and communicate performance and accuracy trade-offs.
Requirements
- Python
- Automation
- Data pipeline
- Performance analysis
- Profiling
- Latency analysis
- Throughput analysis
- Transformer
- LLM inference
- PyTorch
- Hugging Face Transformers
- Experimental design
- Data analysis
- Reproducibility
- vLLM
- SGLang
- TensorRT-LLM
- lm-evaluation-harness
- Quantization
- Mixed precision
- Speculative decoding
- Distributed inference
- GPU
- NPU
- TPU
- Linux
- Container
- Cluster
- CI
- Rust
- C++
Responsibilities
- Design and maintain LLM inference benchmarks
- Measure latency, throughput, memory use, and accelerator utilization
- Evaluate model accuracy and compare results with reference implementations
- Build automated and reproducible benchmark and accuracy evaluation pipelines
- Maintain performance and accuracy dashboards and reports
- Define regression thresholds and release acceptance criteria
- Investigate regressions through profiling and controlled experiments
- Improve measurement quality through statistical analysis and validation
- Communicate performance and accuracy findings and recommendations
