Applied Scientist Agent Evaluation and Adaptive Model Routing

Bitdeer is a technology company providing Bitcoin mining solutions.

Singapore, SG
About Bitdeer

Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build and extend LLM and agent evaluation pipelines, develop representative task suites, and measure task success, tool use, quality, cost, latency, and token consumption at the trajectory level. You will research adaptive model-routing strategies, prototype routing policies, and evaluate methods through traffic replay, shadow testing, and internal pilots. You will also support production integration and analyze quality-cost trade-offs using rigorous experimental methods.

Requirements

  • Bachelor’s, Master’s, or PhD in Computer Science, Machine Learning, Statistics, Electrical Engineering, or a related field
  • Substantial hands-on experience in LLM evaluation, agentic systems, applied machine learning, or adaptive inference
  • Strong programming ability in Python
  • Practical experience with PyTorch and modern data and evaluation tooling
  • Experience designing or operating LLM or agent evaluation pipelines
  • Experience with task-level and trajectory-level metrics, dataset construction, automated scoring, regression testing, and failure analysis
  • Implementation-level depth in model selection and routing, uncertainty estimation and calibration, cascading and escalation, or stage-aware agent inference
  • Experience evaluating multi-turn or tool-using agents
  • Rigorous experimental practice with controlled comparisons, statistical analysis, and honest baselines
  • Experience analyzing per-request and per-trajectory quality, cost, latency, and token usage
  • Experience with multi-model APIs, agent harnesses, traffic replay, shadow evaluation, A/B testing, or production model monitoring is preferred
  • Familiarity with tool calling, context windows, prompt caching, reasoning controls, vLLM, or SGLang is a plus
  • Strong ownership and product judgment

Responsibilities

  • Own and extend LLM and agent evaluation pipelines
  • Develop representative task suites and trajectory-level metrics
  • Measure task success, tool use, quality, cost, latency, and token consumption
  • Research and prototype adaptive model-routing strategies
  • Evaluate routing methods through offline evaluation, traffic replay, shadow testing, and internal pilots
  • Analyze quality-cost trade-offs and construct cost-quality Pareto frontiers
  • Support production integration of evaluation and routing systems