Staff AI Observability and Telemetry Engineer
Bitdeer Technologies Group is a technology company providing Bitcoin mining solutions, mining hardware, data-center infrastructure, and AI cloud services. It serves individual, institutional, and enterprise customers globally.
Funding history
Investors
Projects
About Bitdeer Technologies Group
Bitdeer provides vertically integrated Bitcoin mining and high-performance computing services. Its operations include mining equipment procurement and manufacturing, datacenter design and construction, equipment management, daily mining operations, cloud mining, and mining-related services. The company also offers AI cloud infrastructure and high-performance computing powered by NVIDIA GPUs for AI and machine-learning workloads, serving customers across global markets.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Architect and scale high-cardinality telemetry infrastructure for AI Cloud, integrate hardware-level exporters into Kubernetes observability, build eBPF diagnostic tools, automate dashboards and alerting, and design metrics for multi-tenant billing. Define observability standards for AI workloads, lead architecture reviews, and mentor engineers on high-performance telemetry collection and analysis.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
- 6+ years of software or site reliability engineering experience.
- Deep hands-on expertise in the Prometheus and OpenTelemetry ecosystem.
- Advanced proficiency in Go.
- Experience writing custom Kubernetes metric exporters and operators.
- Experience with eBPF and BCC kernel-level tracing tools.
- Linux performance tuning experience.
- Familiarity with GPU power states, SM utilization, memory bandwidth, and high-performance network telemetry.
- Experience operating, debugging, and scaling large-scale telemetry stacks in HPC or cloud environments.
- Technical leadership and architectural decision-making skills.
- Communication skills for translating system requirements into engineering milestones.
Responsibilities
- Architect and scale high-cardinality telemetry infrastructure using highly available time-series databases.
- Integrate NVIDIA DCGM, network switch telemetry, and IPMI or Redfish exporters into Kubernetes observability.
- Build eBPF diagnostic tools for network congestion, kernel-level I/O latency, and distributed training bottlenecks.
- Develop automated dashboards and alerting pipelines for degraded hardware.
- Design metric pipelines for multi-tenant consumption billing.
- Create observability standards for AI-native workloads.
- Lead technical design reviews for observability architecture.
- Mentor team members on high-performance telemetry collection and analysis.
Benefits
- Attractive welfare benefits
- Training and mentoring opportunities
- Autonomy, personal accountability, and growth opportunities
- Inclusive and diverse work environment
