Senior SRE Engineer

Kronos Research is a quantitative trading firm providing technology-driven trading and liquidity solutions across global digital-asset markets. It serves exchanges, protocols, institutional clients, investors, allocators, and founders.

Distributed
About Kronos Research

Kronos Research delivers quantitative trading and liquidity services for digital-asset markets. Its offerings include market liquidity provision for protocol tokens and exchanges, capital deployment through investment solutions and structured products, systematic OTC liquidity for major digital assets, and strategic venture investments in digital-asset companies. The firm positions its services for exchanges and protocols, investors and allocators, institutions, and founders, and states that its products and services are offered exclusively to professional or qualified investors.

View jobs by Kronos Research

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Manage large-scale Linux environments, HPC clusters, storage, multi-cloud infrastructure, CI/CD systems, and internal AI platforms while supporting production operations.

Requirements

  • 5+ years of hands-on Linux systems administration and infrastructure operations experience.
  • Strong Linux internals knowledge across processes, memory, filesystems, networking, systemd, and cgroups.
  • Strong Bash or shell scripting skills.
  • Programming ability for data processing, CLI tools, and API services; Python preferred.
  • Storage fundamentals including RAID, filesystems, snapshots, backups, NFS, and SMB.
  • Experience with a major public cloud and IaC tooling such as Terraform, CDK, or Ansible.
  • Familiarity with Docker and Kubernetes.
  • CI/CD pipeline design and operations experience with GitLab CI, Jenkins, or Airflow.
  • Ability to own cross-service subsystems end-to-end.
  • Strong autonomy and problem-solving ability.
  • Self-directed approach to identifying and prioritizing problems.

Responsibilities

  • Manage large-scale Linux environments, troubleshooting, and root-cause analysis.
  • Write maintainable Bash, Ansible, and Python automation.
  • Participate in on-call support for infrastructure, CI/CD, and production incidents.
  • Operate HPC clusters using Slurm and related analytics, auditing, and monitoring tools.
  • Maintain and plan Lustre and NAS storage for compute environments.
  • Manage AWS, Alibaba Cloud, and GCP infrastructure using Terraform and AWS CDK.
  • Build and operate Docker/ECS and Kubernetes/EKS environments and deployment workflows.
  • Operate self-hosted GitLab servers and Runner fleets.
  • Design and operate CI/CD systems and deployment pipelines.
  • Build internal AI platforms using LangChain, LangGraph, Bedrock, and Elasticsearch RAG.
  • Develop MCP servers, chatbots, AI agents, and similar services.
Senior SRE Engineer at Kronos Research | JobStash