Senior AI Observability and Telemetry Engineer
Bitdeer is a technology company providing Bitcoin mining solutions.
Funding history
Investors
Projects
About Bitdeer
Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will architect and scale high-cardinality telemetry infrastructure for AI clusters. You will integrate hardware exporters into Kubernetes observability, build eBPF diagnostic tools, create automated dashboards and alerts, develop metering pipelines, and lead observability design reviews.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
- 6+ years of software or site reliability engineering experience.
- Hands-on expertise in the Prometheus and OpenTelemetry ecosystem.
- Advanced proficiency in Go.
- Experience writing custom Kubernetes metric exporters and operators.
- Experience with eBPF, BCC, and Linux performance tuning.
- Familiarity with AI hardware metrics and high-performance network telemetry.
- Experience operating and scaling large-scale telemetry stacks in HPC or cloud environments.
- Technical leadership and communication skills.
- Experience in high-velocity, high-growth engineering environments is preferred.
Responsibilities
- Architect and scale high-cardinality telemetry infrastructure using highly available time-series databases.
- Integrate hardware-level exporters into the Kubernetes observability stack.
- Build eBPF-based diagnostic tools for network, kernel I/O, and distributed-training bottlenecks.
- Develop automated dashboards and alerting pipelines for degraded hardware.
- Design metric pipelines for multi-tenant consumption billing.
- Create observability standards for AI-native workloads.
- Lead technical design reviews and mentor team members.
