Staff Observability Engineer
FanDuel is a premier US-based gaming destination, offering daily fantasy sports, sports betting, and online casino games. Founded in 2009 to innovate the fantasy sports industry, it has become a leading mobile sports betting operator, allowing users to bet on sports, play online games, and build fantasy teams across a wide array of professional sports leagues.
Funding
Projects
About FanDuel
FanDuel is the premier gaming destination and #1 Sportsbook in the United States. What started as a backyard brainstorm in Texas in 2009 has grown into a major force in the gaming industry. The company revolutionized fantasy sports by simplifying the season-long game into daily contests, allowing fans to compete and win daily. Today, with over 12 million registered users, FanDuel offers a wide range of products including sports betting on leagues like the NFL, NBA, MLB, and NHL; an online casino with blackjack, slots, and live dealer games; daily fantasy sports for various sports like football, baseball, and basketball; and horse racing betting for major events. Additionally, FanDuel provides content through FanDuel TV and insights via FanDuel Research, making every moment of the game more engaging for its users.
Skills
About the Role
You will design, build, and mature observability capabilities that give teams actionable insight into system health, performance, reliability, and user experience. You will standardize monitoring, alerting, incident management, and postmortem practices; lead reliability initiatives; automate root-cause analysis; optimize platforms; and mentor engineers.
Requirements
- Significant hands-on experience in observability engineering, SRE, platform engineering, or related roles.
- Expertise in monitoring and observability with hands-on Datadog experience.
- Experience driving observability or reliability strategy across teams or domains.
- Proficiency with Kubernetes, AWS, and Terraform.
- Understanding of distributed systems principles and trade-offs.
- Experience implementing SLOs, SLIs, alerting strategies, and user-centric metrics.
- Software engineering proficiency in Go, Java, Python, or TypeScript.
- Experience building scalable systems, tooling, and automation in complex codebases.
- Experience driving large-scale automation and reducing recurring operational issues.
- Ability to translate technical signals into business and customer impact.
- Communication and stakeholder-management skills.
Responsibilities
- Define and drive the observability strategy and roadmap across teams.
- Design and improve scalable observability capabilities.
- Standardize monitoring, alerting, incident management, and postmortem practices.
- Evolve incident management, on-call practices, and post-incident learning.
- Lead cross-team reliability initiatives and resolve systemic risks.
- Use automation and AI-assisted workflows to accelerate root-cause analysis and reduce operational toil.
- Translate observability insights into strategic roadmap decisions.
- Detect, prevent, and mitigate large-scale issues using system and user signals.
- Optimize observability platforms for cost, scalability, and sustainability.
- Mentor engineers and improve organizational reliability and observability maturity.
Benefits
- Health plans
- Fertility and family-planning programs
- Mental health support
- Fitness benefits
- Paid time off
- Sick leave
- Annual bonus opportunities
- Long-term incentive opportunities
- 401(k) matching up to 5%
- Commuter benefits
- Pet insurance
- Medical insurance
- Vision insurance
- Dental insurance
- Life insurance
- Disability insurance
- Short-term and long-term incentive compensation
- Cash bonuses
- Stock program participation
- 14 paid company holidays
- Paid sick time
