Search...

Senior Inference Runtime Engineer

Bitdeer logo
Bitdeer

Bitdeer is a NASDAQ-listed (BTDR) high-performance computing and Bitcoin mining company headquartered in Singapore. It provides end-to-end Bitcoin mining solutions (mining hardware like SEALMINER, Minerbase containers, cloud mining, hosting, and mining farm/data center operations) as well as AI Cloud services offering GPU compute (NVIDIA GB200 NVL72, B200, H200, H100) for AI training and deployment. Its customers range from individual and institutional Bitcoin miners to enterprises and developers needing scalable AI/ML compute infrastructure.

Singapore, SG
About Bitdeer

Bitdeer Technologies Group is a NASDAQ-listed (ticker: BTDR) technology company headquartered in Singapore that describes itself as a "world-leading" high-performance computing platform and Bitcoin mining services provider. The company is vertically integrated across the value chain, spanning IC design and hardware manufacturing (its own SEALMINER ASIC miners and Minerbase mobile cooling containers), infrastructure construction and cloud mining, and artificial intelligence/high-performance computing. Bitdeer handles the full range of mining-related processes for its customers, including equipment procurement, transport logistics, datacenter design and construction, equipment management, and daily operations, and offers institutional services, a hash rate market, and a miner rights trading marketplace via its mobile apps (Bitdeer App and Minerplus App). Since 2013 Bitdeer has built more than 30 data centers globally and currently operates 9 large-scale data centers (including one of North America's largest) with roughly 3GW of diversified energy capacity and tens of exahashes of managed hash rate, with major operations in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia. Since 2023 it has also been expanding a global AI infrastructure business (Bitdeer AI Cloud), powered by thousands of NVIDIA GPUs (including GB200 NVL72 and B200, with GB300 NVL72 and B300 planned), offering turnkey AI datacenter solutions and GPU cloud compute for AI training and deployment starting at around $2/hour. Bitdeer serves both individual/retail Bitcoin miners and institutional clients, as well as AI developers and enterprises seeking scalable, energy-efficient compute.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own the performance-critical serving layer for self-hosted large language models. You will optimize scheduling, batching, KV cache behavior, decoding, and streaming; tune inference runtimes; profile GPU, network, tokenizer, proxy, and worker bottlenecks; lead model onboarding; define runtime playbooks and safe defaults; and turn benchmark findings into production improvements.

Requirements

  • 6+ years in systems, ML infrastructure, or high-performance backend engineering
  • Experience with LLM serving runtimes such as vLLM, Dynamo, SGLang, TensorRT-LLM, TGI, or Triton
  • Strong understanding of GPU memory, CUDA, NCCL, KV cache, batching, streaming, and distributed inference
  • Go or Python proficiency
  • Ability to read runtime source code and profiling traces
  • Experience operating production inference services with latency, availability, and cost targets
  • Ability to translate performance work into reliability, latency, and margin improvements

Responsibilities

  • Optimize prefill and decode scheduling, continuous batching, KV cache behavior, speculative decoding, long-context serving, and streaming
  • Tune and operate LLM inference runtimes for latency, throughput, GPU utilization, and cost efficiency
  • Profile bottlenecks across GPU memory, HBM bandwidth, NCCL, networking, tokenizers, proxies, and model workers
  • Lead model onboarding and select runtime, parallelism, quantization, context length, and rollback strategies
  • Define runtime playbooks and safe defaults for reasoning, tool calling, multimodal workloads, prompt caching, and provider parameters
  • Partner with SRE and performance and evaluation engineers on production runtime improvements

Benefits

  • Welfare benefits
Senior Inference Runtime Engineer at Bitdeer | JobStash