Senior Site Reliability Engineer, CCIP
Chainlink Labs develops and supports Chainlink.
Maintainer signals as of 9/2/2026
Projects
About Chainlink Labs
Chainlink Labs is a technology company focused on building Chainlink as infrastructure for onchain finance. Its current platform includes decentralized data feeds and streams, cross-chain messaging, automation, verifiable randomness, serverless functions, compliance, privacy, data publishing, and runtime-orchestration services for financial institutions, developers, and DeFi protocols.
Skills
About the Role
As a Senior Site Reliability Engineer on the CCIP Platform team, you will ensure the reliability, scalability, and operational excellence of systems powering Chainlink's Cross-Chain Interoperability Protocol. You will strengthen production resilience, reduce operational toil, enable safe delivery, influence reliability practices, and establish scalable operational standards.
Requirements
- Demonstrated experience in Site Reliability Engineering, Production Engineering, or a similar role operating large-scale distributed systems.
- Deep expertise defining, implementing, and driving adoption of SLOs, SLIs, and error budgets across engineering organizations.
- Built and operated production Kubernetes environments supporting critical services.
- Applied OpenTelemetry to improve observability across distributed systems.
- Experience improving the reliability, scalability, and operability of production infrastructure.
- Demonstrated technical leadership influencing reliability practices across engineering teams.
- Experience performing capacity planning and performance tuning for high-throughput distributed services.
- Previous experience working on Web3 infrastructure or within a crypto-native engineering organization.
- Applied chaos engineering or fault-injection techniques to improve production resilience.
- Partnered with software engineering teams to conduct production-readiness reviews before service launches.
- Experience leading on-call operations, including defining rotations, escalation policies, and improving alert quality.
Responsibilities
- Improve deployment safety and increase delivery velocity by advancing production engineering practices.
- Establish distributed tracing across the platform to improve observability and accelerate incident investigation.
- Eliminate operational toil through automation that increases engineering efficiency and platform reliability.
- Drive adoption of meaningful SLOs, SLIs, and error budgets that guide engineering decisions and improve service health.
- Increase platform scalability and operational readiness as CCIP continues to grow.
- Maintain highly available production systems while reducing operational overhead.
Benefits
- Long term incentives
- Comprehensive benefits
Hiring Process
Applications are carefully reviewed, with candidates expected to receive a status response within two weeks after the job posting closes. The closing date is listed on the job advert.
