Senior Site Reliability Engineer, CCIP

Chainlink Labs develops and supports Chainlink.

Maintainer signals as of 9/2/2026

London, United Kingdom
About Chainlink Labs

Chainlink Labs is a technology company focused on building Chainlink as infrastructure for onchain finance. Its current platform includes decentralized data feeds and streams, cross-chain messaging, automation, verifiable randomness, serverless functions, compliance, privacy, data publishing, and runtime-orchestration services for financial institutions, developers, and DeFi protocols.

View jobs by Chainlink Labs

Skills

About the Role

As a Senior Site Reliability Engineer on the CCIP Platform team, you will ensure the reliability, scalability, and operational excellence of systems powering Chainlink's Cross-Chain Interoperability Protocol. You will strengthen production resilience, reduce operational toil, enable safe delivery, influence reliability practices, and establish scalable operational standards.

Requirements

  • Demonstrated experience in Site Reliability Engineering, Production Engineering, or a similar role operating large-scale distributed systems.
  • Deep expertise defining, implementing, and driving adoption of SLOs, SLIs, and error budgets across engineering organizations.
  • Built and operated production Kubernetes environments supporting critical services.
  • Applied OpenTelemetry to improve observability across distributed systems.
  • Experience improving the reliability, scalability, and operability of production infrastructure.
  • Demonstrated technical leadership influencing reliability practices across engineering teams.
  • Experience performing capacity planning and performance tuning for high-throughput distributed services.
  • Previous experience working on Web3 infrastructure or within a crypto-native engineering organization.
  • Applied chaos engineering or fault-injection techniques to improve production resilience.
  • Partnered with software engineering teams to conduct production-readiness reviews before service launches.
  • Experience leading on-call operations, including defining rotations, escalation policies, and improving alert quality.

Responsibilities

  • Improve deployment safety and increase delivery velocity by advancing production engineering practices.
  • Establish distributed tracing across the platform to improve observability and accelerate incident investigation.
  • Eliminate operational toil through automation that increases engineering efficiency and platform reliability.
  • Drive adoption of meaningful SLOs, SLIs, and error budgets that guide engineering decisions and improve service health.
  • Increase platform scalability and operational readiness as CCIP continues to grow.
  • Maintain highly available production systems while reducing operational overhead.

Benefits

  • Long term incentives
  • Comprehensive benefits

Hiring Process

Applications are carefully reviewed, with candidates expected to receive a status response within two weeks after the job posting closes. The closing date is listed on the job advert.