Search...

Senior Site Reliability Engineer, Core AI Infrastructure

Blockchain Association logo
Blockchain Association

The Blockchain Association is the collective voice of the crypto industry in the United States. It represents over 100 members from the sector's leading investors, companies, and projects to advocate for a pro-innovation national policy and regulatory framework for the crypto economy.

Washington, USA
About Blockchain Association

Blockchain Association is the collective voice of the crypto industry in the United States. With over 100 members, including the sector’s leading investors, companies, and projects, it works to support a future-forward, pro-innovation national policy and regulatory framework for the crypto economy. The association's mission is to advance the future of crypto in the U.S. by promoting the potential of blockchain technology and shaping policy that ensures its success.

View jobs by Blockchain Association

Skills

About the Role

You will own the reliability and automation of AI infrastructure. You will monitor services respond to incidents perform root cause analysis with blameless retros and provide on call support for AWS deployment pipelines. You will build automation to streamline operational IT workflows across CI CD frameworks and Kubernetes environments. You will partner with the infrastructure team to extend CI CD frameworks supporting IT services and enterprise network platforms and with security and compliance to integrate surveillance tooling into deployment pipelines. You will strengthen observability and documentation standards and develop full stack applications powering internal AI products and infrastructure with Go or Python.

Requirements

  • 5+ years of experience automating and supporting cloud infrastructure (AWS) and network environments, with hands-on use of infrastructure-as-code tools (Terraform, Ansible, Chef, Puppet, or Salt)
  • Proven experience deploying, managing, and troubleshooting containerized workloads using Docker and Kubernetes in production environments
  • Proficiency in at least one scripting or programming language (Python, Bash, Ruby, or Go) and version control workflows using Git-based CI/CD pipelines
  • Track record of leading incident response in environments with strict SLAs, including root cause analysis, blameless retros, and measurable reliability improvements
  • Utilizes generative AI responsibly, maintaining human oversight to deliver business-ready outputs and drive measurable improvements in workflow efficiency, cost, and quality

Responsibilities

  • Own the reliability monitoring and incident response lifecycle for AI infrastructure services including on-call support for AWS deployment pipelines root cause analysis and blameless retros
  • Build automation and tooling to streamline operational IT workflows eliminate manual tasks and improve deployment velocity across CI CD frameworks and Kubernetes environments
  • Partner with the infrastructure team to extend CI CD frameworks supporting IT services and enterprise network platforms and with security and compliance to integrate surveillance tooling into deployment pipelines
  • Strengthen observability and documentation standards across IT engineering by defining metrics implementing monitoring solutions and maintaining technical documentation that sets a standard of excellence
  • Develop full stack applications that power internal AI products and infrastructure with Go or Python

Benefits

  • Equity and bonus eligibility
  • Medical dental and vision insurance
  • 401(k)
  • Remote first work arrangement