Member of Technical Staff Datacenter Operations
4 days agoSeniorSalary: 150K - 300KSan Francisco, USARemoteFull TimeOperationsJobs by Prime Intellect
Prime IntellectVisit Prime Intellect website
AI infrastructure company providing an integrated stack for training, evaluating, deploying, and continuously improving agentic models.
Prime Intellect on X (Twitter)Prime Intellect on DiscordPrime Intellect on GitHubPrime Intellect on Documentation
Series ARecently funded34 current maintainers27 active leads7 new active leads9 lead step-downsTeam intelligence
Maintainer signals as of 9/23/2026
San Francisco, United States
Funding history
Projects
About Prime Intellect
Prime Intellect, Inc. operates AI infrastructure spanning RL environments, hosted training and evaluations, inference, secure sandboxes, and globally sourced GPU compute.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own the operational readiness of physical GPU infrastructure. You will coordinate deployments, maintain assets and documentation, lead hardware fault triage, establish maintenance procedures, and support incident response with datacenter partners.
Requirements
- 3+ years in datacenter operations, hardware infrastructure, or production systems operations
- Experience deploying and troubleshooting rack-mounted servers, networking equipment, and structured cabling
- Experience coordinating datacenter providers, remote hands, and hardware vendors
- Knowledge of Linux diagnostics, BMC consoles, and server hardware health tools
- Knowledge of GPU hardware, rack power, cooling, cabling, asset tracking, incident management, and scripting
Responsibilities
- Coordinate rack deployments, cabling, inventory, and acceptance testing
- Maintain asset records, rack layouts, power allocations, cabling documentation, and spare-parts inventories
- Lead hardware fault triage and coordinate remote hands, vendor escalations, replacements, and RMA workflows
- Establish maintenance plans and change procedures
- Track capacity readiness, failure trends, repair times, and operational risks
- Partner on power, cooling, environmental monitoring, and high-density deployment readiness
- Create runbooks and escalation procedures and support incident response
Benefits
- Equity incentives
