Production Engineer Kubernetes
Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Crusoe, Inc
Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and scale tooling and features for managed Kubernetes and virtual machine platforms. You will monitor alerts, performance metrics, and logs; develop monitoring tools; and maintain service-level indicators and objectives. You will participate in incident response, post-mortems, root-cause analysis, and proactive automated remediation. You will work with software engineers on resilient code, deployment reviews, and data-center action plans. You will document your work and share operational insights.
Requirements
- 3–6 years of professional Production Engineer experience.
- Experience building Kubernetes platforms or Kubernetes controllers.
- Exposure to server-class hardware and provisioning.
- Understanding of distributed-systems architecture, design patterns, reliability, and scaling.
- Understanding of infrastructure design and operational trade-offs in network, storage, and RPC serving designs.
- Proficiency in at least one programming language, such as Python or Go.
- Exposure to observability tooling, including logging, monitoring, and alerting.
- Experience with Unix/Linux environments.
- Understanding of TCP/IP and network programming fundamentals.
- Awareness of basic information security best practices.
- Bachelor's degree in Computer Science or a related field, or self-education in computer science fundamentals.
- Strong communication skills.
Responsibilities
- Build and scale tooling and features for managed Kubernetes and managed virtual machine platforms.
- Collaborate on daily priorities, incidents, and action plans for new or retrofitted data centers.
- Advise software engineers on resilient code and review changes before deployment.
- Review alerts and system performance metrics.
- Analyze system logs and develop monitoring tools.
- Participate in incident response drills, post-mortems, and root-cause analysis.
- Automate remediation for common errors.
- Maintain high service-level indicators and service-level objectives.
- Document work and share insights.
Benefits
- Pension contributions.
- Private health insurance.
- Private dental insurance.
- Income protection.
- Life assurance.
