Senior Inference Runtime Engineer
Bitdeer is a technology company providing Bitcoin mining solutions.
Maintainer signals as of 9/25/2026
Funding history
Investors
Projects
About Bitdeer
Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Own the performance-critical serving layer for self-hosted large language models by optimizing scheduling, batching, KV cache behavior, decoding, streaming, and inference runtimes; profiling bottlenecks; leading model onboarding; defining runtime playbooks; and translating benchmark findings into production improvements.
Requirements
- 6+ years of systems, ML infrastructure, or high-performance backend engineering experience.
- Hands-on experience with LLM serving runtimes such as vLLM, Dynamo, SGLang, TensorRT-LLM, TGI, or Triton.
- Strong understanding of GPU memory, CUDA/NCCL, KV cache, batching, streaming, and distributed inference tradeoffs.
- Proficiency in Go or Python and ability to read runtime source code, profiling traces, and production metrics.
- Experience operating production inference services with strict latency, availability, and cost targets.
- Ability to translate low-level performance work into customer-visible reliability, latency, and margin improvements.
Responsibilities
- Optimize prefill/decode scheduling, continuous batching, KV cache behavior, speculative decoding, long-context serving, and streaming.
- Tune and operate vLLM, Dynamo, SGLang, TensorRT-LLM-style runtimes for latency, throughput, GPU utilization, and cost efficiency.
- Profile bottlenecks across GPU memory, HBM bandwidth, NCCL/network, tokenizers, proxies, and model workers.
- Lead model onboarding, including runtime selection, tensor/pipeline parallelism, quantization, context length, and rollback strategy.
- Define runtime playbooks and safe defaults for reasoning, tool calling, multimodal workloads, prompt caching, and provider-specific parameters.
- Partner with SRE and performance/evaluation engineers on production runtime improvements.
Benefits
- Welfare benefits
- Inclusive and respectful work environment
- Training and mentoring opportunities
- Autonomy, personal accountability, and growth opportunities
