Search...

Senior Observability Engineer

FanDuel logo
FanDuel

FanDuel is a premier US-based gaming destination, offering daily fantasy sports, sports betting, and online casino games. Founded in 2009 to innovate the fantasy sports industry, it has become a leading mobile sports betting operator, allowing users to bet on sports, play online games, and build fantasy teams across a wide array of professional sports leagues.

New York, USA
About FanDuel

FanDuel is the premier gaming destination and #1 Sportsbook in the United States. What started as a backyard brainstorm in Texas in 2009 has grown into a major force in the gaming industry. The company revolutionized fantasy sports by simplifying the season-long game into daily contests, allowing fans to compete and win daily. Today, with over 12 million registered users, FanDuel offers a wide range of products including sports betting on leagues like the NFL, NBA, MLB, and NHL; an online casino with blackjack, slots, and live dealer games; daily fantasy sports for various sports like football, baseball, and basketball; and horse racing betting for major events. Additionally, FanDuel provides content through FanDuel TV and insights via FanDuel Research, making every moment of the game more engaging for its users.

View jobs by FanDuel

Skills

About the Role

You will design, build, and mature observability capabilities for platform services. You will combine system telemetry and user signals to improve performance, reliability, and user experience. You will partner with engineering and product stakeholders, improve incident response and on-call practices, automate root cause analysis, and help teams use self-service insights and tooling.

Requirements

  • Experience in observability engineering, SRE, platform engineering, or related roles
  • Expertise in monitoring and observability practices using Datadog
  • Experience contributing to observability or reliability initiatives across teams or services
  • Proficiency with Kubernetes, AWS, and Terraform
  • Understanding of distributed systems principles and trade-offs
  • Experience defining and implementing SLOs, SLIs, and alerting strategies
  • Software engineering fundamentals and proficiency in Go, Java, Python, or TypeScript
  • Experience building tooling, automation, and scalable systems
  • Experience reducing operational toil and recurring issues through automation
  • Analytical problem-solving skills
  • Communication and collaboration skills

Responsibilities

  • Contribute to the observability strategy and roadmap
  • Design and enhance scalable observability solutions
  • Establish and promote monitoring, alerting, incident management, and postmortem practices
  • Improve incident response processes, on-call practices, and post-incident reviews
  • Collaborate on initiatives to improve system reliability and resolve risks
  • Apply automation and AI-assisted workflows to improve root cause analysis and reduce operational toil
  • Surface observability insights that inform technical decisions and prioritization
  • Analyze system and user signals to detect, prevent, and mitigate reliability issues
  • Optimize observability platforms for performance, scalability, and cost-efficiency
  • Mentor peers and raise observability and reliability standards

Benefits

  • Health plans
  • Fertility and family-planning programs
  • Mental health support
  • Fitness benefits
  • Paid time off
  • Sick leave
  • Annual bonus opportunities
  • Long-term incentive opportunities
  • 401(k) matching up to 5%
  • Commuter benefits
  • Pet insurance
  • Medical insurance
  • Vision insurance
  • Dental insurance
  • Life insurance
  • Disability insurance
  • Cash bonuses
  • Stock program participation
  • 14 paid company holidays
Senior Observability Engineer at FanDuel | JobStash