Member of Technical Staff Bare Metal and Fleet Provisioning
Prime IntellectVisit Prime Intellect website
AI infrastructure company providing an integrated stack for training, evaluating, deploying, and continuously improving agentic models.
Prime Intellect on X (Twitter)Prime Intellect on DiscordPrime Intellect on GitHubPrime Intellect on Documentation
Series ARecently funded34 current maintainers27 active leads7 new active leads9 lead step-downsTeam intelligence
Maintainer signals as of 9/23/2026
San Francisco, United States
Funding history
Projects
About Prime Intellect
Prime Intellect, Inc. operates AI infrastructure spanning RL environments, hosted training and evaluations, inference, secure sandboxes, and globally sourced GPU compute.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build the systems that turn bare-metal GPU servers into reliable production compute. You will automate provisioning, configuration, validation, monitoring, repair, and secure reuse throughout the machine lifecycle.
Requirements
- 3+ years of experience operating Linux servers or building bare-metal infrastructure automation in production
- Experience with PXE/i PXE, DHCP, image provisioning, and Redfish or IPMI
- Software engineering and debugging skills in Python, Go, or a comparable language and Bash
- Experience designing automation for partial failures, retries, and configuration drift
- Experience with Linux boot, systemd, kernel and driver troubleshooting, and OS image management
- Experience with Ansible and Terraform
- Knowledge of GPU server diagnostics, PCIe topology, BMC telemetry, networking, metrics, logs, and alerting
Responsibilities
- Build automated GPU server discovery, network boot, OS imaging, and configuration workflows
- Automate BIOS, BMC, NIC, GPU driver, and firmware configuration
- Develop hardware inventory and lifecycle services
- Create hardware acceptance tests and burn-in workflows
- Integrate provisioning and health checks with SLURM, Kubernetes, and compute allocation systems
- Build observability, quarantine, repair, and re-provisioning workflows
- Implement secure credential handling, tenant isolation, and data sanitization
Benefits
- Equity incentives
