Member of Technical Staff - Training Platform
Prime Intellect builds the 'Open Superintelligence Stack' — an integrated compute, training, inference, and sandbox platform that lets companies train, deploy, and continuously improve their own AI models and agents. It serves AI startups, 'neolabs', and enterprises (over 6,000 customers, including Ramp and Zapier) that want to own their model optimization loop rather than rely solely on closed frontier labs.
Maintainer signals as of 8/20/2026
Funding history
Investors
Projects
About Prime Intellect
Prime Intellect is a San Francisco-based AI infrastructure company building what it calls the Open Superintelligence Stack: a full-stack platform spanning GPU compute (on-demand and reserved clusters), large-scale reinforcement learning training ('Lab'), an Environments Hub with 2,500+ community RL environments, hosted evaluations, sandboxed code execution, and dedicated/serverless model inference with native LoRA support. The company maintains open-source libraries (verifiers and prime-rl) used to build and train RL environments, and publishes frontier open research such as the INTELLECT and SYNTHETIC model/dataset series. Prime Intellect works with AI startups, enterprises, and 'neolab' customers such as Ramp and Zapier, helping them turn production traces and evaluations into custom-trained, post-trained agent models that outperform closed frontier models on specific workflows at lower cost and latency. The company has raised over $150M in total funding, including a $130M Series A led by Radical Ventures with participation from NVIDIA Ventures, Intel Capital, and Dell Technologies Capital, and reports over $100M in annualized revenue.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build the hosted training platform and Kubernetes-based infrastructure for managed GPU training. You will develop orchestration, scheduling, autoscaling, control-plane agents, observability, backend APIs, monitoring tools, and product interfaces, while integrating new training capabilities.
Requirements
- Knowledge of open model families and fine-tuning techniques
- LoRA
- QLoRA
- RLHF
- RLAIF
- vLLM
- SGLang
- TensorRT-LLM
- GPU hardware
- Distributed training
- NCCL
- Kubernetes
- Helm
- CRDs
- KEDA
- GPU operator
- Terraform
- Ansible
- Prometheus
- Grafana
- Loki
- OpenTelemetry
- DCGM
- Linux
- Python
- FastAPI
- SQLAlchemy
- TypeScript
- React
- Next.js
- Tailwind
- REST
- tRPC
Responsibilities
- Design and operate Kubernetes-based training and inference orchestration
- Build and maintain Helm charts for reproducible training stacks
- Develop Python control-plane agents
- Implement scheduling and autoscaling for heterogeneous GPU hardware
- Operate GitOps workflows
- Build model caches, checkpoint pipelines, and shared storage
- Operate observability systems
- Build job submission and live run monitoring surfaces
- Develop FastAPI backend services and REST APIs
- Build real-time monitoring and debugging tools
- Ship product interfaces in Next.js, React, and TypeScript
- Interface with trainers, inference servers, and environment servers
- Productize new training capabilities
Benefits
- Significant equity
- Flexible work arrangement
- Full visa sponsorship
- Relocation support
- Professional development budget
- Regular team off-sites
- Conference attendance
