AI Cloud Senior DevOps Engineer
Bitdeer Technologies Group is a technology company providing Bitcoin mining solutions, mining hardware, data-center infrastructure, and AI cloud services. It serves individual, institutional, and enterprise customers globally.
Funding history
Investors
Projects
About Bitdeer Technologies Group
Bitdeer provides vertically integrated Bitcoin mining and high-performance computing services. Its operations include mining equipment procurement and manufacturing, datacenter design and construction, equipment management, daily mining operations, cloud mining, and mining-related services. The company also offers AI cloud infrastructure and high-performance computing powered by NVIDIA GPUs for AI and machine-learning workloads, serving customers across global markets.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Operate as the backbone of deployment and infrastructure operations for AI products and platforms. Automate CI/CD and MLOps workflows, build scalable cloud-native infrastructure, manage GPU resources, implement high availability and disaster recovery, establish observability, enforce security standards, and lead incident resolution.
Requirements
- Bachelor's degree or above in Computer Science, Engineering, or a related technical field
- 5+ years of experience in DevOps, SRE, or cloud infrastructure roles
- Expert knowledge of Linux and networking principles
- Mastery of Docker and Kubernetes
- Experience with AWS, GCP, Azure, Alibaba Cloud, or other public or hybrid cloud platforms
- Strong coding or scripting skills in Go, Python, Shell, or another major language
- Knowledge of CI/CD, Infrastructure as Code, observability, and SRE
- Experience with MLOps, model serving, GPU clusters, or large-scale distributed systems is preferred
Responsibilities
- Design and maintain CI/CD pipelines for applications and machine learning models
- Automate build, testing, deployment, and rollback processes
- Build and scale Kubernetes- and Docker-based cloud infrastructure
- Manage GPU clusters and specialized computing resources
- Design high-availability and disaster recovery strategies
- Provision infrastructure with Terraform, Ansible, and Helm
- Build monitoring, logging, and alerting systems
- Establish security, release, secrets management, and compliance standards
- Lead troubleshooting, root cause analysis, and remediation during incidents
