Production Engineer LLM Serving

FuriosaAI is a South Korean AI semiconductor company building energy-efficient inference accelerators, servers, and software for enterprise and cloud AI deployments.

Seoul, South Korea
About FuriosaAI

FuriosaAI develops the RNGD AI inference accelerator and NXT RNGD Server, alongside a software toolchain for compiling, optimizing, and deploying LLM and agentic-AI workloads.

View jobs by FuriosaAI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own reliability, observability, and service-level performance for production LLM serving services. You will define SLIs and SLOs, build telemetry and alerts, benchmark workloads, diagnose bottlenecks, improve deployment and traffic-management configurations, automate releases, and participate in on-call incident response and preventive improvements.

Requirements

  • Bachelor's degree in Computer Science or a related field, or equivalent practical experience
  • 3+ years developing and operating production services
  • Programming skills in Rust, C++, Python, or Go
  • Experience operating services in Kubernetes
  • Understanding of Linux, containers, networking, concurrency, and distributed systems
  • Understanding of LLM serving, including prefill/decode, batching, KV-cache management, and request scheduling
  • Experience profiling production services and analyzing telemetry
  • Experience building observability systems using metrics, logs, and traces
  • Ability to communicate quantitative findings and collaborate across teams

Responsibilities

  • Operate and improve LLM serving services and define SLIs, SLOs, and production-readiness criteria
  • Build metrics, logs, traces, dashboards, and actionable alerts across the serving request path
  • Design production-representative benchmarks and load tests
  • Establish performance baselines, detect regressions, and plan serving capacity
  • Diagnose bottlenecks in routing, queueing, batching, caching, networking, and inference execution
  • Improve service components and configurations to increase throughput and resource efficiency while meeting latency SLOs
  • Improve deployment topologies, autoscaling, health checks, graceful degradation, and Istio traffic policies
  • Automate deployments and validate changes through performance tests, canary rollouts, and rollback procedures
  • Participate in on-call and incident response, root-cause analysis, and preventive improvements
  • Build runbooks and recovery automation
Production Engineer LLM Serving at FuriosaAI | JobStash