Senior Software Engineer, Streaming
Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Crusoe, Inc
Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Design, build, and operate high-throughput streaming systems for observability data across an AI cloud, processing logs, metrics, traces, and events while improving reliability and developer experience.
Requirements
- Experience building and operating distributed systems, streaming systems, or real-time data platforms
- Hands-on experience with Kafka or similar distributed streaming technologies
- Proficiency in Java, Scala, Go, or Python
- Experience operating services in cloud or large-scale infrastructure environments
- Understanding of metrics, logging, tracing, and alerting
- Ability to debug production issues across distributed systems
- Ability to own features from design through production operations
- Experience with observability platforms, stream processing frameworks, Kubernetes, schema management, data contracts, serialization formats, or bare-metal infrastructure is beneficial
Responsibilities
- Design, build, and maintain streaming services and pipelines
- Ingest and process logs, metrics, traces, and operational events
- Implement real-time data processing systems using Kafka, Kinesis, Pub/Sub, Flink, or similar platforms
- Scale streaming infrastructure for high-throughput telemetry and high-cardinality workloads
- Ensure streaming systems are reliable and observable
- Build instrumentation, dashboards, and alerting
- Integrate streaming data into observability tools and operational workflows
- Participate in on-call rotations, incident response, and post-incident reviews
- Improve reliability and developer experience through automation, CI/CD, and infrastructure as code
- Contribute to technical design discussions and reviews
Benefits
- Restricted Stock Units
- Health insurance
- Vision insurance
- Dental insurance
- HSA contributions
- Paid parental leave
- Life insurance
- Short-term disability insurance
- Long-term disability insurance
- Teladoc
- 401(k) with 100% match up to 4% of salary
- Paid time off
- Paid holidays
- Cell phone reimbursement
- Tuition reimbursement
- Calm app subscription
- Legal services
- Commuter benefit of $300 per month
