Senior Software Engineer Observability

Together AI operates an AI-native cloud platform for open and custom AI models.

San Francisco, United States
About Together AI

Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.

View jobs by Together AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design and implement observability platforms for metrics, logs, and traces. You will build telemetry pipelines, alerting, anomaly detection, and infrastructure automation. You will work with engineering teams to improve tracing and monitoring, lead incident response and post-mortem analysis, and define observability best practices.

Requirements

  • Expertise with Prometheus, Grafana, ClickStack, OpenTelemetry, and cloud-native monitoring services
  • Programming experience with Go, Python, or similar languages
  • Proficiency with Terraform, Ansible, and Helm
  • Experience designing and operating large-scale distributed systems and high-volume data pipelines
  • Understanding of Docker and Kubernetes
  • Knowledge of microservices, service mesh, CI/CD, and GitOps
  • Experience managing PostgreSQL, MongoDB, Redis, and time-series databases

Responsibilities

  • Design and implement observability platforms for metrics, logs, and traces
  • Build telemetry data pipelines and log aggregation workflows
  • Develop automated monitoring, alerting, anomaly detection, SLIs, SLOs, runbooks, and predictive analytics
  • Build observability tools and infrastructure as code
  • Enhance distributed tracing and application monitoring
  • Lead incident response and post-mortem analysis
  • Define observability best practices

Benefits

  • Startup equity
  • Health insurance
  • Remote-work flexibility
Senior Software Engineer Observability at Together AI | JobStash