Search...

Production Engineer Kubernetes

Crusoe logo
Crusoe

Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.

Distributed
About Crusoe

Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.

View jobs by Crusoe

Skills

About the Role

You will build and scale tooling and features for managed Kubernetes and virtual machine platforms. You will monitor alerts, performance metrics, and logs; develop monitoring tools; and maintain service-level indicators and objectives. You will participate in incident response, post-mortems, root-cause analysis, and proactive automated remediation. You will work with software engineers on resilient code, deployment reviews, and data-center action plans. You will document your work and share operational insights.

Requirements

  • 3–6 years of professional Production Engineer experience.
  • Experience building Kubernetes platforms or Kubernetes controllers.
  • Exposure to server-class hardware and provisioning.
  • Understanding of distributed-systems architecture, design patterns, reliability, and scaling.
  • Understanding of infrastructure design and operational trade-offs in network, storage, and RPC serving designs.
  • Proficiency in at least one programming language, such as Python or Go.
  • Exposure to observability tooling, including logging, monitoring, and alerting.
  • Experience with Unix/Linux environments.
  • Understanding of TCP/IP and network programming fundamentals.
  • Awareness of basic information security best practices.
  • Bachelor's degree in Computer Science or a related field, or self-education in computer science fundamentals.
  • Strong communication skills.

Responsibilities

  • Build and scale tooling and features for managed Kubernetes and managed virtual machine platforms.
  • Collaborate on daily priorities, incidents, and action plans for new or retrofitted data centers.
  • Advise software engineers on resilient code and review changes before deployment.
  • Review alerts and system performance metrics.
  • Analyze system logs and develop monitoring tools.
  • Participate in incident response drills, post-mortems, and root-cause analysis.
  • Automate remediation for common errors.
  • Maintain high service-level indicators and service-level objectives.
  • Document work and share insights.

Benefits

  • Pension contributions.
  • Private health insurance.
  • Private dental insurance.
  • Income protection.
  • Life assurance.