Staff AI Observability & Telemetry Engineer
Bitdeer is a technology company providing Bitcoin mining solutions.
Funding history
Investors
Projects
About Bitdeer
Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will architect high cardinality telemetry infrastructure, integrate GPU and network exporters with Kubernetes observability, build eBPF diagnostic tools, create proactive dashboards and alerting, design usage metrics for multi tenant billing, establish observability standards for AI workloads, and lead technical reviews while mentoring peers.
Requirements
- Bachelor’s or Master’s degree in Computer Science Electrical Engineering or a related field
- 6+ years of software or site reliability engineering experience
- Deep hands on expertise in Prometheus and OpenTelemetry
- Advanced proficiency in Go
- Extensive experience writing custom Kubernetes metric exporters and Operators
- Experience with eBPF BCC and Linux performance tuning
- Familiarity with AI hardware metrics and high performance network telemetry
- Experience operating debugging and scaling large telemetry stacks
- Strong technical leadership and communication skills
- Experience in high velocity engineering environments preferred
Responsibilities
- Architect scalable high cardinality telemetry infrastructure
- Integrate NVIDIA DCGM network switch telemetry IPMI and Redfish exporters
- Build eBPF based diagnostic tools for network and kernel performance
- Develop dashboards and alerting for degraded hardware
- Design metric pipelines for multi tenant consumption billing
- Create observability standards for AI workloads
- Lead observability architecture reviews and mentor team members
