Search...

Senior Site Reliability Engineer CCIP

Chainlink Labs logo
Chainlink Labs

Chainlink Labs helped pioneer decentralized finance and now works with the world's largest financial institutions to enable transactions across the tokenized asset economy. It builds and operates the Chainlink platform, serving financial market infrastructures, asset managers, and top DeFi protocols.

Distributed
About Chainlink Labs

Chainlink Labs is a company at the center of the shift of global finance onchain, working with financial market infrastructures, asset managers, and leading DeFi protocols. The company employs a world-class team of over 600 developers, researchers, and capital markets experts with deep experience in cryptography and decentralized systems, working toward building Chainlink into the global standard for onchain finance. Chainlink Labs operates a unified platform powering a global system of onchain finance across DeFi and TradFi, enabling significant transaction value across a multi-chain ecosystem, and has been battle-tested on mainnet for years with a strong security track record. Its clients and partners span capital markets, DeFi, gaming, and other major verticals, including partnerships with institutions such as ANZ Bank, DTCC, Swift, Sygnum, Fidelity International, and Fireblocks.

View jobs by Chainlink Labs

Skills

About the Role

You will ensure the reliability, scalability, and operational excellence of the CCIP platform. You will strengthen production resilience by improving deployment safety, establishing distributed tracing for observability, eliminating operational toil through automation, driving adoption of meaningful SLOs and SLIs and error budgets, and increasing platform scalability and readiness as CCIP grows. You will be responsible for maintaining highly available production systems and reducing operational overhead.

Requirements

  • Demonstrated experience in Site Reliability Engineering, Production Engineering, or a similar role operating large-scale distributed systems.
  • Deep expertise defining, implementing, and driving adoption of SLOs, SLIs, and error budgets across engineering organizations.
  • Built and operated production Kubernetes environments supporting critical services.
  • Applied OpenTelemetry to improve observability across distributed systems.
  • Experience improving the reliability, scalability, and operability of production infrastructure.
  • Demonstrated technical leadership influencing reliability practices across engineering teams.
  • Experience performing capacity planning and performance tuning for high-throughput distributed services.
  • Previous experience working on Web3 infrastructure or within a crypto-native engineering organization.
  • Applied chaos engineering or fault-injection techniques to improve production resilience.
  • Partnered with software engineering teams to conduct production-readiness reviews before service launches.
  • Experience leading on-call operations, including defining rotations, escalation policies, and improving alert quality.

Responsibilities

  • Improve deployment safety and increase delivery velocity by advancing production engineering practices.
  • Establish distributed tracing across the platform to improve observability and accelerate incident investigation.
  • Eliminate operational toil through automation that increases engineering efficiency and platform reliability.
  • Drive adoption of meaningful SLOs, SLIs, and error budgets that guide engineering decisions and improve service health.
  • Increase platform scalability and operational readiness as CCIP continues to grow.
  • Strengthen Chainlink's reputation through highly available production systems while reducing operational overhead.

Benefits

  • Long term incentives
  • Comprehensive benefits
Senior Site Reliability Engineer CCIP at Chainlink Labs | JobStash