Senior Inference Runtime Engineer
Bitdeer Technologies Group is a technology company providing Bitcoin mining solutions, mining hardware, data-center infrastructure, and AI cloud services. It serves individual, institutional, and enterprise customers globally.
Funding history
Investors
Projects
About Bitdeer Technologies Group
Bitdeer provides vertically integrated Bitcoin mining and high-performance computing services. Its operations include mining equipment procurement and manufacturing, datacenter design and construction, equipment management, daily mining operations, cloud mining, and mining-related services. The company also offers AI cloud infrastructure and high-performance computing powered by NVIDIA GPUs for AI and machine-learning workloads, serving customers across global markets.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Own the performance-critical serving layer for self-hosted large language models by optimizing scheduling, batching, KV cache behavior, decoding, streaming, and inference runtimes; profiling bottlenecks; leading model onboarding; defining runtime playbooks; and translating benchmark findings into production improvements.
Requirements
- 6+ years of systems, ML infrastructure, or high-performance backend engineering experience.
- Hands-on experience with LLM serving runtimes such as vLLM, Dynamo, SGLang, TensorRT-LLM, TGI, or Triton.
- Strong understanding of GPU memory, CUDA/NCCL, KV cache, batching, streaming, and distributed inference tradeoffs.
- Proficiency in Go or Python and ability to read runtime source code, profiling traces, and production metrics.
- Experience operating production inference services with strict latency, availability, and cost targets.
- Ability to translate low-level performance work into customer-visible reliability, latency, and margin improvements.
Responsibilities
- Optimize prefill/decode scheduling, continuous batching, KV cache behavior, speculative decoding, long-context serving, and streaming.
- Tune and operate vLLM, Dynamo, SGLang, TensorRT-LLM-style runtimes for latency, throughput, GPU utilization, and cost efficiency.
- Profile bottlenecks across GPU memory, HBM bandwidth, NCCL/network, tokenizers, proxies, and model workers.
- Lead model onboarding, including runtime selection, tensor/pipeline parallelism, quantization, context length, and rollback strategy.
- Define runtime playbooks and safe defaults for reasoning, tool calling, multimodal workloads, prompt caching, and provider-specific parameters.
- Partner with SRE and performance/evaluation engineers on production runtime improvements.
Benefits
- Welfare benefits
- Inclusive and respectful work environment
- Training and mentoring opportunities
- Autonomy, personal accountability, and growth opportunities
