Search...

Member of Technical Staff - GPU Infrastructure

Prime Intellect logo
Prime Intellect

Prime Intellect builds the 'Open Superintelligence Stack' — an integrated compute, training, inference, and sandbox platform that lets companies train, deploy, and continuously improve their own AI models and agents. It serves AI startups, 'neolabs', and enterprises (over 6,000 customers, including Ramp and Zapier) that want to own their model optimization loop rather than rely solely on closed frontier labs.

Seed33 current maintainers31 active leads9 new active leads3 lead step-downsTeam intelligence

Maintainer signals as of 8/12/2026

San Francisco, USA
About Prime Intellect

Prime Intellect is a San Francisco-based AI infrastructure company building what it calls the Open Superintelligence Stack: a full-stack platform spanning GPU compute (on-demand and reserved clusters), large-scale reinforcement learning training ('Lab'), an Environments Hub with 2,500+ community RL environments, hosted evaluations, sandboxed code execution, and dedicated/serverless model inference with native LoRA support. The company maintains open-source libraries (verifiers and prime-rl) used to build and train RL environments, and publishes frontier open research such as the INTELLECT and SYNTHETIC model/dataset series. Prime Intellect works with AI startups, enterprises, and 'neolab' customers such as Ramp and Zapier, helping them turn production traces and evaluations into custom-trained, post-trained agent models that outperform closed frontier models on specific workflows at lower cost and latency. The company has raised over $150M in total funding, including a $130M Series A led by Radical Ventures with participation from NVIDIA Ventures, Intel Capital, and Dell Technologies Capital, and reports over $100M in annualized revenue.

View jobs by Prime Intellect

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design, deploy, optimize, and support large-scale GPU infrastructure for customers. You will architect GPU clusters, implement orchestration and high-performance networking, configure parallel filesystems, tune system performance, resolve infrastructure issues across the stack, and provide operational documentation and support.

Requirements

  • 3+ years of hands-on experience with GPU clusters and HPC environments
  • SLURM and Kubernetes in production GPU settings
  • InfiniBand configuration and troubleshooting
  • NVIDIA GPU architecture, CUDA ecosystem, and driver stack
  • Ansible and Terraform
  • Python, Bash, and systems programming
  • Customer-facing technical leadership
  • NVIDIA drivers, Fabric Manager, and DCGM
  • Docker, Containerd, and Enroot
  • Linux kernel tuning
  • AI workload network topology
  • Power and cooling requirements for high-density GPU deployments

Responsibilities

  • Partner with clients to understand workload requirements
  • Design GPU cluster architectures
  • Create technical proposals and capacity plans
  • Develop deployment strategies for LLM training, inference, and HPC workloads
  • Present architectural recommendations
  • Deploy SLURM and Kubernetes
  • Implement InfiniBand, RoCE, and NVLink networking
  • Optimize GPU utilization and memory management
  • Configure Lustre, BeeGFS, and GPFS filesystems
  • Tune kernel and CUDA configurations
  • Resolve customer infrastructure issues
  • Implement monitoring, alerting, and automated remediation
  • Provide on-call support
  • Create runbooks and documentation

Benefits

  • Equity incentives