Search...

Senior Software Engineer

Robinhood logo
Robinhood

Robinhood helps users invest in stocks, ETFs, options, and cryptocurrencies through commission-free trading with no minimum account requirements.

United States
About Robinhood

Robinhood helps retail investors access financial markets through commission-free trading of stocks, ETFs, options, and cryptocurrencies. With it, users can invest with no account minimums, earn rewards through retirement accounts with matching contributions, and access advanced trading tools. Robinhood democratizes investing by making financial markets accessible to everyone, not just wealthy investors.

View jobs by Robinhood

Skills

About the Role

You will join a highly visible role focused on incident leadership, operational excellence, and reliability tooling. You will own the processes and tools that enable fast, high quality incident response. You will define and maintain incident management processes, dashboards and alerts, and you will mentor others to improve monitoring and observability. You will design and implement next generation failure mitigation strategies and contribute to postmortem practices and learning. You will collaborate with engineers across domains to raise the bar for reliability and service quality.

Requirements

  • 5+ years of software engineering experience including operating production systems
  • 2+ years focused on reliability engineering infrastructure distributed systems or production operations
  • Hands-on experience in incident leadership roles
  • Strong communication and cross-functional collaboration skills during high-severity incidents
  • Deep knowledge of systems reliability observability frameworks and fault-tolerant design
  • Experience with multi-region or multi-cluster architectures capacity planning and failover strategies
  • Familiarity with modern observability stacks such as OpenTelemetry Prometheus Grafana
  • Proven ability to drive measurable improvements in MTTD MTTR availability or customer impact

Responsibilities

  • Lead incident response efforts and coordinate cross-functional teams during outages
  • Define and maintain incident management processes and postmortems
  • Develop and maintain dashboards and alerts tied to critical user journeys and availability
  • Drive MTTD MTTR improvements and evolve incident response tooling
  • Own incident discovery and maintain a clear source of truth during active incidents
  • Mentor teams on monitoring and observability to improve reliability
  • Deliver executive level reporting on service quality and reliability
  • Support capacity planning and multi-region reliability strategies

Benefits

  • Challenging high-impact work
  • Equity ownership and performance-based compensation
  • 100% paid health insurance with 90% coverage for dependents
  • Lifestyle wallet for wellness and learning
  • Employer-paid life and disability insurance fertility benefits and mental health benefits
  • Paid time off including holidays sick time parental leave and more
  • Exceptional office experience with catered meals events and comfortable workspaces