Staff AI Observability & Telemetry Engineer

Bitdeer is a technology company providing Bitcoin mining solutions.

Singapore, SG
About Bitdeer

Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will architect high cardinality telemetry infrastructure, integrate GPU and network exporters with Kubernetes observability, build eBPF diagnostic tools, create proactive dashboards and alerting, design usage metrics for multi tenant billing, establish observability standards for AI workloads, and lead technical reviews while mentoring peers.

Requirements

  • Bachelor’s or Master’s degree in Computer Science Electrical Engineering or a related field
  • 6+ years of software or site reliability engineering experience
  • Deep hands on expertise in Prometheus and OpenTelemetry
  • Advanced proficiency in Go
  • Extensive experience writing custom Kubernetes metric exporters and Operators
  • Experience with eBPF BCC and Linux performance tuning
  • Familiarity with AI hardware metrics and high performance network telemetry
  • Experience operating debugging and scaling large telemetry stacks
  • Strong technical leadership and communication skills
  • Experience in high velocity engineering environments preferred

Responsibilities

  • Architect scalable high cardinality telemetry infrastructure
  • Integrate NVIDIA DCGM network switch telemetry IPMI and Redfish exporters
  • Build eBPF based diagnostic tools for network and kernel performance
  • Develop dashboards and alerting for degraded hardware
  • Design metric pipelines for multi tenant consumption billing
  • Create observability standards for AI workloads
  • Lead observability architecture reviews and mentor team members