Site Reliability Engineer
SpaceXAI builds frontier AI models and products including the Grok conversational assistant and a multimodal developer API.
Funding history
About SpaceXAI
SpaceXAI is a US-based AI company focused on accelerating scientific discovery and understanding the universe through reasoning, voice, image, video, and generative AI systems. Formerly xAI, it was acquired by SpaceX effective February 2, 2026, while continuing to operate the x.ai site and services.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design monitoring architecture, improve alert quality, and provide technical leadership during severe incidents. You will run postmortems, drive corrective actions, lead cross-functional reliability projects, maintain playbooks and dependency maps, define availability objectives, and join on-call incident response rotations.
Requirements
- Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience
- 5+ years of site reliability, systems engineering, plant operations, or large-scale production operations experience
- Large-scale incident command experience
- Fleet-scale or campus-scale monitoring and observability design experience
- Experience across at least two of compute, network, storage, power, and cooling or facilities telemetry
- Experience operating playbooks or runbooks with a 24/7 operations, control room, or NOC partner
- Python or Bash scripting proficiency
- Experience with a systems language such as C, C++, Java, Go, or Rust
- Problem-solving and cross-functional collaboration skills
Responsibilities
- Own monitoring architecture and alert signal quality
- Use NOC feedback to improve alert suppression and redesign
- Provide technical incident leadership and bridge coordination
- Run blameless postmortems and close corrective actions
- Lead reliability projects across compute, network, storage, and facilities
- Build and maintain playbooks, runbooks, and dependency maps
- Run game days
- Define error budgets and availability objectives
- Participate in on-call rotations and SEV incident response
