Site Reliability Engineer (SRE)

NodeOps Network provides a chain-agnostic, AI-powered DePIN orchestration layer for general-purpose compute, secured by Autonomous Verifiable Services (AVS).

Seed0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/2/2026

Distributed

Funding history

About NodeOps Network

NodeOps Network helps protocols and developers simplify blockchain node operations through permissionless infrastructure and node orchestration. The network enables one-click node deployment, containerized infrastructure, and multi-region compute distribution, bridging the gap between compute providers, node operators, and protocols.

View jobs by NodeOps Network

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own the reliability, scalability, availability, and performance of production deployment infrastructure. You will build and maintain observability and CI/CD systems, define SLOs and SLIs, scale Kubernetes workloads, participate in on-call rotations, lead incident learning, harden security, and automate operational work.

Requirements

  • 3 to 6 years of SRE, DevOps, or platform engineering experience
  • Expertise in Kubernetes and container orchestration at scale
  • Proficiency in Go, Rust, or Python
  • Experience with AWS, GCP, or bare-metal cloud-native infrastructure
  • Experience operating production systems at significant scale
  • Incident command experience

Responsibilities

  • Own and improve production-system reliability, availability, and performance
  • Design, implement, and maintain metrics, logging, and distributed tracing
  • Build and refine CI/CD pipelines
  • Conduct blameless post-mortems and eliminate recurring incident classes
  • Define SLOs and SLIs for new features with product engineering
  • Scale Kubernetes clusters supporting GPU workloads and AI inference
  • Automate recurring operational work
  • Participate in a 24/7 on-call rotation
  • Harden secrets management, network policies, and runtime security
  • Deploy agents for automated runbooks, anomaly detection, and incident triage

Benefits

  • Equity stake
  • Incident bonuses
  • Health benefits
  • Flexible PTO
Site Reliability Engineer (SRE) at NodeOps Network | JobStash