Site Reliability Engineer

A 501(c)(3) science-and-technology research institute/foundation operating direct Neuro-AI research, life-sciences programs through Radial, and an open-science residency program.

Emeryville, CA, United States
About Astera Institute

Astera Institute supports and operates open public-goods work for science and technology. Its current activities include neuroscience-informed AGI research, AI-enabled life-sciences programs, open datasets and publishing infrastructure, and residencies for early-stage projects.

View jobs by Astera Institute

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will operate the infrastructure supporting research compute, container registries, and dashboards. You will improve compute access and resource visibility, enable autoscaling, manage access controls, create reproducible deployments, and automate operational processes.

Requirements

  • Take accountability for cluster health and capacity
  • Understand interactions among schedulers, containers, networking, storage, and hardware
  • Design systems with predictable failure modes
  • Apply observability, reproducibility, and clear operational boundaries
  • Support experimental research workloads pragmatically

Responsibilities

  • Ensure efficient access to compute resources
  • Provide visibility into resource utilization and cluster health
  • Enable automatic scaling of compute resources
  • Manage access to infrastructure resources
  • Drive deterministic deployments and reproducible research environments
  • Automate operational processes
  • Operate infrastructure using Ansible, Kubernetes, Docker, Tailscale, Python, Grafana, Prometheus, and Talos Linux