Member of Technical Staff - Training Platform
Prime Intellect provides an open superintelligence stack for training, evaluating, deploying, and continuously improving AI agents and models. Its platform combines RL environments, hosted training, inference, GPU compute, secure sandboxes, and open-source research tooling for researchers, startups, and enterprises.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Prime Intellect, Inc.
Prime Intellect operates an integrated AI infrastructure platform spanning Lab, hosted reinforcement-learning training, evaluations, environments, inference, secure sandboxes, and on-demand or reserved GPU compute. It also develops open-source tools including Verifiers, prime-rl, and Prime Agent, supporting workflows from environment creation and model evaluation through post-training and production deployment. The company serves researchers, startups, enterprises, and teams building agentic AI systems, with customer examples including Ramp and Zapier.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Prime Intellect is hiring a Member of Technical Staff to develop hosted training infrastructure and platform experiences spanning Kubernetes orchestration, GPU scheduling, autoscaling, observability, backend services, monitoring tools, and frontend product interfaces.
Requirements
- Strong knowledge of open model families and fine-tuning techniques including LoRA, QLoRA, full fine-tuning, RLHF, and RLAIF
- Familiarity with inference engines such as vLLM, SGLang, and TensorRT-LLM
- Understanding of GPU hardware tradeoffs and distributed training fundamentals
- Strong Kubernetes operations experience with Helm, CRDs, operators, KEDA, gang scheduling, and GPU operators
- Production cluster debugging experience
- Cloud platform experience, preferably GCP
- Infrastructure automation experience with Helm, Terraform, and Ansible
- Observability experience with Prometheus, Grafana, Loki, OpenTelemetry, and DCGM
- Linux networking, namespaces, and performance tuning fundamentals
- Strong Python backend development with FastAPI, async programming, and SQLAlchemy
- Experience building Python agents that interact with Kubernetes APIs
- Modern frontend development with TypeScript, React or Next.js, Tailwind, and shadcn
- REST and tRPC API design experience
- Experience building developer tools, dashboards, and live-monitoring UIs
Responsibilities
- Design and operate Kubernetes-based training and inference orchestration across multi-cluster, multi-cloud GPU fleets
- Build and maintain Helm charts for reproducible training stacks
- Develop Python control-plane agents that watch pods and synchronize cluster state
- Implement scheduling and autoscaling for heterogeneous GPU hardware
- Operate GitOps workflows and build model caches, checkpoint pipelines, and shared storage
- Operate observability systems and improve GPU cluster debugging
- Build job submission, live monitoring, logging, metrics, and model management surfaces
- Develop FastAPI backend services and REST APIs
- Build real-time monitoring and debugging tools
- Ship product interfaces with Next.js, React, and TypeScript
- Interface with trainers, inference servers, and environment servers
- Productize new training capabilities, model architectures, and reinforcement learning modes
Benefits
- Cash compensation of $150K–$300K
- Significant equity
- Flexible work arrangement with remote or San Francisco office options
- Full visa sponsorship
- Relocation support
- Professional development budget for courses and conferences
- Regular team off-sites
- Conference attendance
- Opportunity to shape decentralized AI development
