Production Engineer LLM Serving
FuriosaAI is a South Korean AI semiconductor company building energy-efficient inference accelerators, servers, and software for enterprise and cloud AI deployments.
Funding history
Investors
About FuriosaAI
FuriosaAI develops the RNGD AI inference accelerator and NXT RNGD Server, alongside a software toolchain for compiling, optimizing, and deploying LLM and agentic-AI workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own reliability, observability, and service-level performance for production LLM serving services. You will define SLIs and SLOs, build telemetry and alerts, benchmark workloads, diagnose bottlenecks, improve deployment and traffic-management configurations, automate releases, and participate in on-call incident response and preventive improvements.
Requirements
- Bachelor's degree in Computer Science or a related field, or equivalent practical experience
- 3+ years developing and operating production services
- Programming skills in Rust, C++, Python, or Go
- Experience operating services in Kubernetes
- Understanding of Linux, containers, networking, concurrency, and distributed systems
- Understanding of LLM serving, including prefill/decode, batching, KV-cache management, and request scheduling
- Experience profiling production services and analyzing telemetry
- Experience building observability systems using metrics, logs, and traces
- Ability to communicate quantitative findings and collaborate across teams
Responsibilities
- Operate and improve LLM serving services and define SLIs, SLOs, and production-readiness criteria
- Build metrics, logs, traces, dashboards, and actionable alerts across the serving request path
- Design production-representative benchmarks and load tests
- Establish performance baselines, detect regressions, and plan serving capacity
- Diagnose bottlenecks in routing, queueing, batching, caching, networking, and inference execution
- Improve service components and configurations to increase throughput and resource efficiency while meeting latency SLOs
- Improve deployment topologies, autoscaling, health checks, graceful degradation, and Istio traffic policies
- Automate deployments and validate changes through performance tests, canary rollouts, and rollback procedures
- Participate in on-call and incident response, root-cause analysis, and preventive improvements
- Build runbooks and recovery automation
