Software Engineer Site Reliability

Hebbia is an AI platform for institutional finance that lets teams use AI agents to research, analyze, and produce work across complex document and data sets.

New York, United States
About Hebbia

Hebbia provides institutional-intelligence software for financial organizations. Its Max and Matrix products support agentic research, analysis, and traceable workflows across firm data and external financial sources.

View jobs by Hebbia

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own critical production systems from design through incident response. You will write production code, improve service performance and observability, define reliability objectives, build deployment tooling, and turn incident learnings into durable architectural improvements.

Requirements

  • 5+ years of software development experience writing, shipping, and maintaining production services
  • Proficiency in Go, Python, C++, or Rust
  • Experience as a Production Engineer, SRE, or infrastructure-focused software engineer
  • Deep understanding of distributed systems
  • Container orchestration expertise
  • Experience debugging distributed production failures
  • Knowledge of operating-system concepts
  • Cloud platform experience, preferably AWS
  • Experience building and maintaining observability stacks
  • CI/CD pipeline expertise

Responsibilities

  • Own critical production services from design and code review through deployment, operation, and incident response
  • Profile, benchmark, and rewrite hot paths to eliminate performance bottlenecks
  • Lead incident response and turn post-mortem findings into code and architecture improvements
  • Build observability frameworks, instrumentation, alerting logic, and debugging tooling
  • Define and enforce SLOs for platform services
  • Own capacity planning and cost-efficiency automation
  • Build internal platforms and deployment tooling
  • Improve CI/CD systems
  • Embed with product engineering teams to co-design reliable systems
  • Partner on infrastructure security through threat modeling, hardening, and compliance tooling

Benefits

  • Unlimited PTO
  • Medical, dental, vision, and 401K
  • Daily catered lunch
  • DoorDash dinner credit for late work
  • 3 months of parental leave for non-birthing parents and 4 months for birthing parents
  • $15k lifetime fertility benefit
  • New-hire equity grant
Software Engineer Site Reliability at Hebbia | JobStash