Lead SRE (Site Reliability Engineer)

Subsquid (SQD) provides the onchain data layer.

Maintainer signals as of 9/25/2026

Distributed
About Subsquid

Subsquid operates SQD, a data infrastructure stack for blockchain data comprising four integrated products: Portal (an HTTP API for raw blockchain data with streaming and finality handling), Network (a decentralized data lake with 2,000+ worker nodes and 225+ supported networks), SDK (TypeScript libraries — Squid SDK and Pipes SDK — for data transformation and persistence), and Cloud (a managed deployment platform). Every block is validated with six cryptographic checks at ingestion, replicated across the network, and delivered via a single streaming API to destinations like Snowflake, BigQuery, PostgreSQL, Kafka, and S3/GCS. The company serves institutional and enterprise clients in DeFi, compliance, trading, security, analytics, and prediction markets, with over $20B in TVL served by SQD-powered protocols, and is repositioning toward institutional clients under new CEO Wanja S. Oberhof. The stack is open-source (Portal software is AGPL-3.0 licensed Rust) and self-hostable, with an SQD token used for network staking and delegation.

View jobs by Subsquid

Skills

About the Role

Design, build, and optimize high-availability blockchain data infrastructure while owning CI/CD, orchestration, infrastructure as code, monitoring, alerting, incident management, on-call operations, reliability, observability, automation, fault tolerance, and cost efficiency.

Requirements

  • 3+ years of experience as an SRE, DevOps Engineer, or similar role.
  • Experience running production services against real SLAs, including on-call, incident response, and post-mortems.
  • Experience defining and implementing metrics, logging, and alerting.
  • Proficiency in Kubernetes, Terraform, Prometheus, Grafana, or equivalent monitoring tools.
  • Strong understanding of distributed systems and streaming data pipelines.
  • Deep knowledge of AWS, GCP, or bare-metal infrastructure and cost optimization.
  • Experience balancing performance, reliability, and cost.
  • Programming skills in Python, Go, Rust, or Bash.
  • Experience monitoring and running blockchain nodes or working with node providers.
  • Willingness to learn blockchain node internals, EVM/SVM data, and new chain integration.
  • Previous Web3 experience preferred.

Responsibilities

  • Design, build, and optimize a high-availability blockchain data ingestion pipeline.
  • Own CI/CD, orchestration, and infrastructure-as-code layers.
  • Identify and implement tools for running and monitoring blockchain nodes.
  • Assess infrastructure trade-offs to optimize performance, reliability, and cost.
  • Build and maintain a public status page and incident management process.
  • Define and maintain SRE metrics, logging, and alerting.
  • Improve observability, automation, and fault tolerance.
  • Contribute to incident response, troubleshooting, and on-call rotations.

Benefits

  • Competitive salary plus token incentives
  • Fully remote work
  • Flexible hours
  • High-impact role with ownership
  • Opportunity to build operational foundations for an AI/Web3 company