Staff Software Engineer Cloud Monitoring Service Distributed Systems
Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.
Maintainer signals as of 8/20/2026
Funding history
Projects
About Crusoe
Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead the design and development of Crusoe Cloud’s distributed monitoring systems. You will own high-throughput, multi-tenant telemetry pipelines spanning edge collection, ingestion, stream processing, storage, and query. You will improve reliability, participate in customer-facing on-call, guide architecture, collaborate across engineering teams, and mentor engineers.
Requirements
- 8+ years of software development experience
- Deep hands-on experience designing and operating distributed systems at scale
- Experience with sharding, replication, consistency, load balancing, and concurrency
- Experience with time-series databases, log aggregation, streaming pipelines, or distributed tracing backends
- Strong programming fundamentals in Go or another modern compiled language
- Experience with Kubernetes, microservices, and CI/CD
- Customer-facing on-call experience
- Experience driving technical outcomes across team boundaries
Responsibilities
- Own the architecture and evolution of large-scale telemetry pipelines
- Design scalable and multi-tenant services
- Improve pipeline reliability and data freshness
- Participate in a customer-facing on-call rotation
- Set technical direction for distributed systems work
- Drive design reviews and architecture decisions
- Collaborate with product, compute, networking, and platform teams
- Coach engineers through design work, code review, and incident response
Benefits
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off
- Paid holidays
- Leave of absence programs
- Comprehensive health, dental and vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance
- Short-term and long-term disability insurance
- Professional development and tuition reimbursement
- Mental health and wellness support
- Commuter benefits
- Cell phone stipend
- 401(k) retirement plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance and emergency assistance
- Daily meals allowance
- Additional location-specific perks and programs
