Search...

Senior Site Reliability Engineer

ZetaChain logo
ZetaChain

ZetaChain is a Layer 1 blockchain enabling cross-chain interoperability between Bitcoin, Ethereum, Solana, and other networks without bridges or wrapped tokens.

United States
About ZetaChain

ZetaChain is a Proof of Stake Layer 1 blockchain built on Cosmos SDK and Tendermint consensus that enables native cross-chain smart contract functionality across multiple blockchain networks. The platform allows developers to deploy Universal Apps that can read from and write to any connected blockchain, including non-smart contract chains like Bitcoin and Dogecoin, through its omnichain smart contracts architecture. ZetaChain uses a distributed network of validators, observers, and signers with Threshold Signature Scheme technology to facilitate secure cross-chain transactions without requiring bridges or wrapped tokens.

View jobs by ZetaChain

Skills

About the Role

You will operate and maintain production blockchain infrastructure including validators, RPC services, indexers, and supporting services. You will build and maintain monitoring, alerting, and dashboards, write automation and infrastructure code to reduce toil, and participate in on-call rotations and incident response. You will partner with engineering teams to embed reliability, scalability, and security best practices, improve Kubernetes reliability across cloud and bare-metal environments, and refine deployment, rollback, and recovery strategies.

Requirements

  • 4+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or Platform Engineering
  • Strong software engineering background with production experience in Go and/or Python
  • Deep experience operating Linux systems in production
  • Proven experience running Kubernetes at scale
  • Experience supporting high-availability distributed systems
  • Comfortable working in fast-moving startup environments
  • Strong security mindset for infrastructure running on public or adversarial networks
  • Excellent collaboration and communication skills

Responsibilities

  • Operate and maintain production blockchain infrastructure including validators, RPC services, and indexers
  • Ensure high availability and performance for AI-enabled developer platforms and internal tooling
  • Build and maintain monitoring, alerting, and dashboards for protocol, infrastructure, and application health
  • Write automation and infrastructure code to reduce operational overhead
  • Participate in on-call rotations, incident response, and post-incident reviews
  • Partner with engineering teams to embed reliability, scalability, and security best practices into system design
  • Improve Kubernetes reliability across cloud and bare-metal environments
  • Continuously refine deployment, rollback, and recovery strategies

Benefits

  • Fully remote work with quarterly in-person team meetups
  • Additional 10% to 25% in liquid benefits
  • Meaningful ownership