Search...

Member of Technical Staff - Inference

Prime Intellect logo
Prime Intellect

Prime Intellect builds the 'Open Superintelligence Stack' — an integrated compute, training, inference, and sandbox platform that lets companies train, deploy, and continuously improve their own AI models and agents. It serves AI startups, 'neolabs', and enterprises (over 6,000 customers, including Ramp and Zapier) that want to own their model optimization loop rather than rely solely on closed frontier labs.

Seed33 current maintainers31 active leads9 new active leads3 lead step-downsTeam intelligence

Maintainer signals as of 8/12/2026

San Francisco, USA
About Prime Intellect

Prime Intellect is a San Francisco-based AI infrastructure company building what it calls the Open Superintelligence Stack: a full-stack platform spanning GPU compute (on-demand and reserved clusters), large-scale reinforcement learning training ('Lab'), an Environments Hub with 2,500+ community RL environments, hosted evaluations, sandboxed code execution, and dedicated/serverless model inference with native LoRA support. The company maintains open-source libraries (verifiers and prime-rl) used to build and train RL environments, and publishes frontier open research such as the INTELLECT and SYNTHETIC model/dataset series. Prime Intellect works with AI startups, enterprises, and 'neolab' customers such as Ramp and Zapier, helping them turn production traces and evaluations into custom-trained, post-trained agent models that outperform closed frontier models on specific workflows at lower cost and latency. The company has raised over $150M in total funding, including a $130M Series A led by Radical Ventures with participation from NVIDIA Ventures, Intel Capital, and Dell Technologies Capital, and reports over $100M in annualized revenue.

View jobs by Prime Intellect

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build infrastructure for efficient multi-tenant LLM serving across cloud GPU fleets, design GPU-aware scheduling and failover, optimize inference frameworks and parallelism, integrate distributed inference into RL systems, and establish CI/CD, observability, documentation, and incident response practices.

Requirements

  • 3+ years building and operating large-scale ML or LLM services
  • Experience with vLLM, SGLang, or TensorRT-LLM
  • Familiarity with distributed and disaggregated serving infrastructure
  • Understanding of prefill, decode, KV-cache behavior, batching, sampling, speculative decoding, and parallelism
  • Experience debugging CUDA, NCCL, drivers, kernels, containers, service mesh, networking, and storage
  • Python
  • PyTorch
  • AWS or GCP
  • Kubernetes
  • CUDA
  • NCCL
  • InfiniBand

Responsibilities

  • Build a multi-tenant LLM serving platform across cloud GPU fleets
  • Design GPU-aware placement and scheduling algorithms
  • Implement multi-region and multi-zone failover and traffic shifting
  • Build autoscaling, routing, and load balancing
  • Optimize model distribution and cold-start times
  • Integrate and contribute to LLM inference frameworks
  • Optimize parallelism, caching, memory management, quantization, and speculative decoding
  • Profile kernels, memory bandwidth, and transport
  • Develop reproducible performance suites
  • Embed distributed inference into the RL stack
  • Establish CI/CD with artifact promotion and performance gates
  • Build observability and manage SLOs
  • Document architectures and playbooks
  • Mentor and collaborate cross-functionally

Benefits

  • Significant equity incentives
  • Flexible work arrangement
  • Full visa sponsorship
  • Relocation support
  • Professional development budget
  • Regular team off-sites
  • Conference attendance