Senior Inference Runtime Engineer

Bitdeer is a technology company providing Bitcoin mining solutions.

0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

Singapore, SG
About Bitdeer

Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Own the performance-critical serving layer for self-hosted large language models by optimizing scheduling, batching, KV cache behavior, decoding, streaming, and inference runtimes; profiling bottlenecks; leading model onboarding; defining runtime playbooks; and translating benchmark findings into production improvements.

Requirements

  • 6+ years of systems, ML infrastructure, or high-performance backend engineering experience.
  • Hands-on experience with LLM serving runtimes such as vLLM, Dynamo, SGLang, TensorRT-LLM, TGI, or Triton.
  • Strong understanding of GPU memory, CUDA/NCCL, KV cache, batching, streaming, and distributed inference tradeoffs.
  • Proficiency in Go or Python and ability to read runtime source code, profiling traces, and production metrics.
  • Experience operating production inference services with strict latency, availability, and cost targets.
  • Ability to translate low-level performance work into customer-visible reliability, latency, and margin improvements.

Responsibilities

  • Optimize prefill/decode scheduling, continuous batching, KV cache behavior, speculative decoding, long-context serving, and streaming.
  • Tune and operate vLLM, Dynamo, SGLang, TensorRT-LLM-style runtimes for latency, throughput, GPU utilization, and cost efficiency.
  • Profile bottlenecks across GPU memory, HBM bandwidth, NCCL/network, tokenizers, proxies, and model workers.
  • Lead model onboarding, including runtime selection, tensor/pipeline parallelism, quantization, context length, and rollback strategy.
  • Define runtime playbooks and safe defaults for reasoning, tool calling, multimodal workloads, prompt caching, and provider-specific parameters.
  • Partner with SRE and performance/evaluation engineers on production runtime improvements.

Benefits

  • Welfare benefits
  • Inclusive and respectful work environment
  • Training and mentoring opportunities
  • Autonomy, personal accountability, and growth opportunities