Site Reliability Engineer Data Center
SpaceXAI builds frontier AI models and products including the Grok conversational assistant and a multimodal developer API.
Funding history
About SpaceXAI
SpaceXAI is a US-based AI company focused on accelerating scientific discovery and understanding the universe through reasoning, voice, image, video, and generative AI systems. Formerly xAI, it was acquired by SpaceX effective February 2, 2026, while continuing to operate the x.ai site and services.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will evaluate firmware and hardware releases, investigate complex failures, manage RMA cases, and work with vendors on resolutions. You will build monitoring tools and automation, support real-time hardware troubleshooting, document reliability findings, and participate in hardware incident response and on-call rotations.
Requirements
- Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field, or equivalent experience
- 2+ years of hardware reliability engineering experience
- Expertise in firmware analysis, hardware specification review, and release validation
- Experience with RMA processes and vendor negotiations
- Ability to diagnose complex hardware failures
- Familiarity with data center hardware
- Python or Bash scripting proficiency
- Experience with a systems language such as C, C++, Java, or Rust
- Problem-solving and cross-functional collaboration skills
Responsibilities
- Analyze firmware packages and hardware specifications
- Run firmware security scanning and vulnerability analysis
- Identify safety issues before releases reach the data center
- Diagnose and validate complex hardware failures
- Manage RMA claims and vendor resolutions
- Troubleshoot and optimize hardware systems with operations technicians
- Develop monitoring tools, scripts, and processes
- Document failure modes, RCAs, reliability models, RMA outcomes, and hardware evaluations
- Participate in on-call rotations and hardware incident response
