Search...

Senior Infrastructure and Operations Engineer Kubernetes Platform Reliability

Range logo
Range

Range provides blockchain security and intelligence for monitoring, compliance, and investigations across chains.

Distributed
About Range

Range helps blockchain teams secure protocols by providing comprehensive monitoring, threat detection, compliance tools, and forensic capabilities. Users can detect security breaches, perform due diligence, and respond to incidents. Range enables proactive security for cross-chain infrastructure.

View jobs by Range

Skills

About the Role

You will design, deploy, and maintain production Kubernetes clusters and own their reliability, security, upgrades, and performance. You will build monitoring, logging, alerting, observability, CI/CD, and infrastructure automation systems. You will investigate incidents, strengthen resilience and recovery, optimize Cloudflare, define reliability targets, and improve operational practices.

Requirements

  • Senior-level experience operating production infrastructure
  • Deep hands-on Kubernetes expertise, including cluster internals, networking, storage, and security
  • Networking fundamentals, including TCP/IP, routing, DNS, TLS, and load balancing
  • Experience debugging distributed systems and network-related issues
  • Experience optimizing CDN and edge setups, including Cloudflare
  • Experience building monitoring and observability systems
  • Experience with metrics, logs, traces, and alerting pipelines
  • Experience designing reliable CI/CD pipelines
  • Linux fundamentals
  • Experience with infrastructure as code and automation
  • Experience debugging issues across the entire stack
  • Experience handling incidents and conducting postmortems
  • Experience with multi-cluster or multi-region setups
  • Experience with high-throughput or data-heavy systems
  • Experience with Elasticsearch or large-scale data infrastructure
  • Experience with service meshes
  • Experience with cost optimization and capacity planning
  • Experience in regulated or reliability-focused environments

Responsibilities

  • Design, deploy, and maintain production Kubernetes clusters
  • Own cluster reliability, upgrades, security, and performance
  • Build and operate monitoring, logging, and alerting pipelines
  • Ensure full-stack observability across infrastructure and services
  • Design and maintain reliable CI/CD pipelines
  • Improve deployment strategies, including rollouts, canaries, and rollbacks
  • Automate infrastructure provisioning and configuration
  • Investigate and resolve production incidents
  • Improve system resilience, redundancy, and recovery strategies
  • Define SLOs and SLIs and track reliability targets
  • Optimize and maintain Cloudflare caching, routing, security, and edge behavior
  • Improve operational practices with engineering teams
  • Identify and remove single points of failure

Benefits

  • Meaningful equity upside
  • Remote-first culture
  • Bi-yearly international off-sites
  • Global conference travel and ecosystem engagement
  • Health and wellness benefits