Senior Site Reliability Engineer
SSV Network is a decentralized, open-source Ethereum staking network built on Secret Shared Validator (SSV) technology. It distributes validator duties and key shares across independent operators to provide fault-tolerant, non-custodial staking infrastructure for staking providers, institutions, pools, and solo stakers.
Funding
Projects
About SSV Network
SSV Network provides distributed validator technology (DVT) for Ethereum staking. Validator keys are split into shares and operated by multiple non-trusting nodes, improving redundancy, key security, decentralization, and resilience. The network offers an application, SDK, APIs, subgraph, smart contracts, node-operator tooling, and documentation for onboarding and operating validators. It also provides SSV staking, through which SSV holders can stake tokens, receive cSSV, support the network’s oracle infrastructure, and earn network-fee rewards.
Skills
About the Role
You will work at the intersection of cloud infrastructure and blockchain, building the platform that product teams deploy to. You will design and implement infrastructure and tools that let teams iterate rapidly and securely, with a strong focus on reliability and automation. You will take a proactive role in resolving production issues and learning from incidents in a blameless way, and you will collaborate closely with product teams on production deployments, release management, and incident handling. You will bring AI into engineering workflows by building and deploying autonomous agents and LLM-powered tooling — from sandboxed coding agents to LLM-assisted CI/CD and automated incident triage. This role comes with real ownership and room to shape how the organization operates as it grows.
Requirements
- Kubernetes expertise with strong understanding of core concepts and ability to manage and maintain clusters
- Expertise with modern cloud native tools such as ArgoCD for GitOps, Terraform/Crossplane for IaC, and the Grafana LGTM stack for observability
- 3-5 years of experience with Infrastructure as Code and cloud provisioning tools
- 3-5 years of development and scripting experience in languages like Go or Python
- Proficient written and spoken English with exceptional communication abilities
- Expertise in Linux environments, containerization, and cloud technologies
- Comprehensive knowledge of production management concepts for distributed systems
- 3-5 years in operational roles overseeing production settings
- AI fluency with daily use of AI coding tools and ability to build LLM-powered developer tooling and autonomous agents
- Bonus: service mesh experience, platform engineering, and cross-cloud networking
- Advantage: familiarity with the Ethereum ecosystem, staking, and blockchain technologies
Responsibilities
- Design and implement infrastructure and tools that empower product teams to rapidly and securely iterate, emphasizing reliability and automation
- Influence the strategic direction of infrastructure and operational practices to support a growing organization
- Take a proactive role in resolving production issues and learning from incidents in a blameless manner
- Work closely with product teams on production deployments, release management, and incident handling
- Offer technical expertise to support continual adoption and modernization of the platform and infrastructure
- Build and deploy AI-powered tooling such as autonomous coding agents, LLM-assisted CI/CD, and automated incident triage
- Foster a culture of continuous learning and improvement through constructive review and adaptation processes
