Software Engineer Infrastructure

Open-source AI company building models, the Hermes Agent, a hosted agent platform, and distributed-training infrastructure.

Series A14 current maintainers11 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

Austin, United States
About Nous Research

Nous Research, Inc. trains models, builds AI agents, and develops infrastructure for open intelligence. Its current offerings include the installable Hermes Agent and Nous Portal’s paid model, tool, and cloud-agent services.

View jobs by Nous Research

Skills

About the Role

You will design, build, and operate the infrastructure and platform supporting Hermes Agent, Nous Portal, and managed AI services. You will own backend services, platform infrastructure, deployment pipelines, and engineering tooling end to end. You will improve CI/CD, infrastructure as code, observability, incident response, performance, scalability, and cost efficiency. You will also evaluate third-party services and support production launches for AI capabilities.

Requirements

  • 3+ years of software engineering experience with significant ownership of infrastructure or platform systems
  • Strong programming skills in TypeScript, Node.js, Python, Go, or Rust
  • Hands-on experience with AWS, Azure, or GCP
  • Experience with Docker and Kubernetes
  • Familiarity with CI/CD systems, infrastructure as code, observability, monitoring, and production incident response
  • Strong software engineering fundamentals across backend, infrastructure, and platform systems
  • Problem-solving skills
  • Curiosity about AI infrastructure

Responsibilities

  • Design, build, and operate infrastructure and platforms for AI services
  • Own backend services, platform infrastructure, deployment pipelines, and supporting tooling end to end
  • Architect scalable and reliable infrastructure for inference, agent execution, model serving, and enterprise deployments
  • Build and improve CI/CD pipelines, deployment workflows, and infrastructure as code
  • Implement observability, monitoring, alerting, incident response, and reliability improvements
  • Optimize infrastructure performance, scalability, and cost
  • Evaluate and integrate third-party infrastructure and managed services
  • Support new AI capabilities and production launches