Platform Support Engineer
Braintrust is an active AI observability platform for teams building and operating production AI agents.
Funding history
About Braintrust
Braintrust provides tracing, evaluation, experimentation, and pattern-discovery tooling for AI agents. Teams can inspect agent traces, evaluate outputs with automated or human scoring, build datasets from production behavior, and gate releases against quality regressions.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will support hybrid and self-hosted deployments from installation through ongoing operation. You will diagnose infrastructure, performance, and reliability issues; lead customer-impacting incidents; contribute fixes to services and deployment tooling; build diagnostic tools; maintain runbooks; and participate in on-call coverage.
Requirements
- Customer-facing technical support, SRE, DevOps, solutions architecture, infrastructure engineering, or backend infrastructure engineering experience
- Kubernetes
- Terraform
- AWS
- Python, TypeScript, or Go
- Observability tooling
- Communication
Responsibilities
- Support hybrid and self-hosted deployments across AWS, Azure, and GCP
- Debug Kubernetes, Terraform, networking, VPC, IAM, TLS, and cloud-provider issues
- Diagnose backend performance and reliability issues using logs, metrics, and traces
- Lead incident response and communicate clearly with customers
- Submit fixes to backend services, Terraform modules, and deployment tooling
- Build diagnostics, health checks, preflight validation, and self-service tooling
- Write and maintain runbooks and deployment documentation
- Share recurring failure patterns with Engineering and Product
- Participate in an on-call rotation for critical customer issues
Benefits
- Medical, dental, and vision insurance
- Daily lunch, snacks, and beverages
- Flexible time off
- Equity
- Wifi and cellphone stipend
