Senior Site Reliability Engineer, Node Platform
Chainlink Labs helped pioneer decentralized finance and now works with the world's largest financial institutions to enable transactions across the tokenized asset economy. It builds and operates the Chainlink platform, serving financial market infrastructures, asset managers, and top DeFi protocols.
About Chainlink Labs
Chainlink Labs is a company at the center of the shift of global finance onchain, working with financial market infrastructures, asset managers, and leading DeFi protocols. The company employs a world-class team of over 600 developers, researchers, and capital markets experts with deep experience in cryptography and decentralized systems, working toward building Chainlink into the global standard for onchain finance. Chainlink Labs operates a unified platform powering a global system of onchain finance across DeFi and TradFi, enabling significant transaction value across a multi-chain ecosystem, and has been battle-tested on mainnet for years with a strong security track record. Its clients and partners span capital markets, DeFi, gaming, and other major verticals, including partnerships with institutions such as ANZ Bank, DTCC, Swift, Sygnum, Fidelity International, and Fireblocks.
Skills
About the Role
You will design and build Kubernetes-based infrastructure primitives and the CRE control plane that enable deterministic horizontal scaling of decentralized oracle networks. You will develop Kubernetes Operators, scaling automation, and reusable platform components, codify scaling logic, and implement safe, repeatable infrastructure expansion. You will improve operational efficiency, diagnosability, and the scalability of stateful distributed systems.
Requirements
- 6–9+ years in SRE / Platform / Infrastructure Engineering
- Proven experience scaling Kubernetes in high-throughput production environments
- Deep knowledge of scheduler behavior
- Deep knowledge of StatefulSets & persistent workloads
- Deep knowledge of autoscaling strategies (HPA, VPA, KEDA, custom scaling)
- Resource management & performance tuning
- Multi-cluster and multi-region architectures
- Experience in diagnosing production failures at the cluster scale
- Strong Terraform or Crossplane experience
- GitOps workflows (ArgoCD / Flux) experience
- CI/CD reliability experience
- Automation-first mindset
- AWS production experience
- Proficiency in Go or equivalent systems language
Responsibilities
- Design infrastructure primitives for decentralized oracle networks
- Build Kubernetes-based control plane components
- Develop Kubernetes Operators and scaling automation
- Codify scaling logic into reusable operators and automation
- Ensure deterministic horizontal scaling of networks
- Implement safe and repeatable infrastructure expansion
- Improve operational efficiency and scalability
- Enhance diagnosability and observability for production systems
