Member of Technical Staff - GPU Infrastructure
1 month agoSeniorSalary: 150K - 300KSan Francisco, USAHybridFull TimeEngineeringJobs by Prime Intellect
Prime IntellectVisit Prime Intellect website
AI infrastructure company providing an integrated stack for training, evaluating, deploying, and continuously improving agentic models.
Prime Intellect on X (Twitter)Prime Intellect on DiscordPrime Intellect on GitHubPrime Intellect on Documentation
Series ARecently funded34 current maintainers27 active leads7 new active leads9 lead step-downsTeam intelligence
Maintainer signals as of 9/25/2026
San Francisco, United States
Funding history
Projects
About Prime Intellect
Prime Intellect, Inc. operates AI infrastructure spanning RL environments, hosted training and evaluations, inference, secure sandboxes, and globally sourced GPU compute.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Design, deploy, optimize, and support large-scale GPU infrastructure for customers, including GPU clusters, orchestration, high-performance networking, parallel filesystems, system performance, infrastructure troubleshooting, documentation, and operational support.
Requirements
- 3+ years of hands-on experience with GPU clusters and HPC environments
- Deep expertise with SLURM and Kubernetes in production GPU settings
- Experience with InfiniBand configuration and troubleshooting
- Strong understanding of NVIDIA GPU architecture, CUDA, and drivers
- Experience with Ansible and Terraform
- Proficiency in Python, Bash, and systems programming
- Customer-facing technical leadership experience
- Experience with NVIDIA drivers, Fabric Manager, and DCGM
- Experience configuring Docker, Containerd, and Enroot for GPUs
- Linux kernel tuning and performance optimization
- AI workload network topology design
- Knowledge of power and cooling requirements for high-density GPU deployments
Responsibilities
- Partner with clients to understand workload requirements and design GPU cluster architectures
- Create technical proposals and capacity plans for clusters ranging from 100 to 10,000+ GPUs
- Develop deployment strategies for LLM training, inference, and HPC workloads
- Present architectural recommendations to technical and executive stakeholders
- Deploy and configure SLURM and Kubernetes
- Implement InfiniBand, RoCE, and NVLink networking
- Optimize GPU utilization, memory management, and inter-node communication
- Configure Lustre, BeeGFS, and GPFS filesystems
- Tune kernel and CUDA configurations
- Resolve customer infrastructure issues
- Implement monitoring, alerting, and automated remediation
- Provide 24/7 on-call support for critical customer deployments
- Create runbooks and documentation
Benefits
- Equity incentives
