Staff Software Engineer Observability
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design and implement observability instrumentation, telemetry pipelines, internal platforms, and tooling. You will operationalize service-level indicators, objectives, and alerting; improve incident root-cause analysis; create actionable dashboards; and balance telemetry signal, cost, noise, and performance impact.
Requirements
- Backend or systems software engineering
- Go, C++, Rust, Java, or Python
- Distributed systems
- Networking
- Concurrency
- Performance optimization
- Metrics
- Logging
- Distributed tracing
- Production monitoring
- Alerting
- OpenTelemetry
- Prometheus
- Grafana
- Datadog, Elastic, Jaeger, Tempo, or similar tools
- Telemetry pipeline
- Service-level indicator
- Service-level objective
Responsibilities
- Design and implement observability instrumentation across services and platforms
- Build and maintain telemetry pipelines for metrics, logs, and traces
- Develop internal observability platforms, libraries, and tooling
- Define and operationalize service-level indicators, objectives, and alerting strategies
- Make systems debuggable by design
- Reduce incident resolution time through root-cause analysis
- Create actionable dashboards and alerts
- Balance telemetry signal against cost, noise, and performance impact
- Improve the developer experience for observability and debugging
