NPU Software Engineer Runtime

South Korean AI semiconductor and infrastructure company building inference accelerators, servers, racks, and software for production-scale AI deployment.

Seongnam-si, Gyeonggi-do, South Korea
About Rebellions

Rebellions develops purpose-built AI inference hardware and an accompanying software stack. Its current offerings include the Rebel100 accelerator and deployable RebelServer, RebelRack, and RebelPOD systems, designed for energy-efficient data-center inference.

View jobs by Rebellions

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build runtime modules that connect compilers, drivers, and ML frameworks for model deployment. You will support PyTorch execution, develop profiling capabilities, extend vLLM, optimize multi-NPU distributed inference, benchmark performance, and help deploy scalable inference services.

Requirements

  • Over 5 years of software-engineering experience with ML frameworks, inference runtimes, or AI accelerator toolchains
  • Bachelor’s degree or higher in Computer Science, Electrical Engineering, or a related field
  • Strong proficiency in C++ and Python
  • Understanding of deep learning, LLM architectures, generative AI, and inference optimization
  • Experience with LLM serving frameworks such as vLLM or TensorRT-LLM
  • Understanding of tensor parallelism, KV-cache optimization, and memory-efficient execution
  • Familiarity with compilers, runtimes, drivers, firmware, and hardware acceleration
  • Debugging and performance-profiling skills for high-throughput inference
  • Written and verbal communication skills

Responsibilities

  • Design and implement runtime modules that interface with compilers and drivers
  • Maintain native PyTorch execution support, torch.compile integration, and compiler toolchains
  • Develop a user-facing performance profiler for the SDK
  • Extend vLLM to improve NPU inference performance
  • Design and optimize distributed multi-NPU inference and collective communication
  • Benchmark, profile, and optimize runtime-system performance
  • Deploy and scale inference services with ML and infrastructure engineers

Hiring Process

Document screening > online interview > onsite interview including an assignment > culture-fit interview > compensation discussion > final offer

NPU Software Engineer Runtime at Rebellions | JobStash