Director of Site Operations
SpaceXAI builds frontier AI models and products including the Grok conversational assistant and a multimodal developer API.
Funding history
About SpaceXAI
SpaceXAI is a US-based AI company focused on accelerating scientific discovery and understanding the universe through reasoning, voice, image, video, and generative AI systems. Formerly xAI, it was acquired by SpaceX effective February 2, 2026, while continuing to operate the x.ai site and services.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own node, rack, and cluster health across sites operating around the clock. You will lead site managers, shift supervisors, technicians, and site reliability engineers; direct incident response; coordinate remediation; manage vendors; and use operational data to improve uptime, repair times, and SLA performance.
Requirements
- Bachelor’s degree and 7+ years of large-scale operations experience with 5+ years leading technical people leaders, or 10+ years of large-scale operations experience with 5+ years leading technical people leaders
- Willingness to travel frequently to data center locations
- Ability to lift up to 50 lbs unassisted, stand for long periods, and occasionally use ladders
- Willingness to work extended hours and weekends as needed
Responsibilities
- Own node, rack, and cluster health across sites
- Lead site managers, shift supervisors, technicians, and site reliability engineers
- Drive recovery of failed nodes and racks
- Coordinate with facilities, network engineering, vendors, and tenant representatives
- Direct hardware rework, repairs, replacements, and capacity work
- Lead proactive monitoring, fault mitigation, root cause analysis, and reliability procedures
- Use operational data to improve uptime, repair time, and SLA performance
- Command response to cluster-impacting incidents
- Standardize best practices and scale operations across sites
