Senior DevOps Engineer
Alpaca provides developer APIs for stock, options, and crypto trading, enabling businesses and developers to build trading applications and automate investment strategies.
Projects
About Alpaca
Alpaca helps developers and businesses access financial markets through API-first infrastructure for stock, options, and cryptocurrency trading. Users can build algorithmic trading strategies, create investing applications, and integrate brokerage services into their platforms.
Skills
About the Role
You will design, build and operate the infrastructure that lets Alpaca scale globally and run trading-critical systems with confidence. You will have the autonomy to design and implement solutions against clearly defined goals, and a real voice in shaping those goals with the team. You will design and evolve cloud architecture on GCP expressed entirely as code with Terraform, build and own CI/CD pipelines for IaC changes, create self-serve platform capabilities, strengthen the observability stack, operate GKE clusters and the services running on them, participate in a Follow-The-Sun on-call model, and embed SRE practices into how Core Infrastructure builds and operates.
Requirements
- 5+ years in a DevOps, Platform/Infrastructure, or SRE role with a proven track record operating large-scale, high-availability, high-performance systems in production
- Deep hands-on experience designing cloud architecture on Google Cloud Platform (GCP) as the primary cloud
- Strong Infrastructure-as-Code skills with Terraform across multiple environments, with GitOps as a first principle
- Proven experience building CI/CD pipelines for IaC with automated plan/apply, Policy-as-Code, drift detection and safe rollout
- Significant production experience with Kubernetes (ideally GKE) and packaging/deploying workloads with Helm
- Solid cloud and L3/L4-L7 networking fundamentals (VPCs, routing, load balancing, DNS, TLS, interconnects)
- Hands-on experience with a modern observability stack: Prometheus, Thanos, Grafana, Loki, Tempo and Alertmanager
- Operator-level familiarity with data stores such as PostgreSQL and message brokers (e.g. RabbitMQ, RedPanda)
- Good understanding of SRE practices such as SLOs, error budgets and capacity planning
- Strong grasp of incident management end to end, including structured debugging, escalation, documentation and post-mortems
- Able and willing to take part in a Follow-The-Sun on-call rotation from APAC hours and work effectively in a distributed, async-first team
Responsibilities
- Design and evolve cloud architecture on GCP, including networking, interconnects, IAM and high-availability topology, expressed as code with Terraform following GitOps principles
- Build and own CI/CD pipelines that plan, review, test and safely apply IaC changes with Policy-as-Code guardrails, drift detection and progressive rollout
- Advance Platform-as-a-Product by building self-serve capabilities and paved paths so engineers can provision what they need through a golden path
- Strengthen the observability stack across metrics, logs, traces and alerting using Prometheus, Thanos, Grafana, Loki, Tempo and Alertmanager
- Operate GKE clusters and the infrastructure services running on them, including Helm-packaged workloads, message brokers and data stores
- Participate in the Follow-The-Sun on-call model, triaging alerts, joining and declaring incidents, leading structured debugging and escalation, and driving blameless post-mortems
- Embed SRE practices such as SLIs/SLOs, error budgets and capacity planning into how Core Infrastructure builds and operates
Benefits
- Stock options
- Health benefits
- One-time USD $500 new hire home-office setup
- Monthly stipend of USD $150 via a Brex Card
