Member of Technical Staff - Inference
Prime Intellect provides an open superintelligence stack for training, evaluating, deploying, and continuously improving AI agents and models. Its platform combines RL environments, hosted training, inference, GPU compute, secure sandboxes, and open-source research tooling for researchers, startups, and enterprises.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Prime Intellect, Inc.
Prime Intellect operates an integrated AI infrastructure platform spanning Lab, hosted reinforcement-learning training, evaluations, environments, inference, secure sandboxes, and on-demand or reserved GPU compute. It also develops open-source tools including Verifiers, prime-rl, and Prime Agent, supporting workflows from environment creation and model evaluation through post-training and production deployment. The company serves researchers, startups, enterprises, and teams building agentic AI systems, with customer examples including Ramp and Zapier.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Build infrastructure for efficient multi-tenant LLM serving across cloud GPU fleets, design GPU-aware scheduling and failover, optimize inference frameworks and parallelism, integrate distributed inference into RL systems, and establish CI/CD, observability, documentation, and incident response practices.
Requirements
- 3+ years building and operating large-scale ML or LLM services
- Experience with vLLM, SGLang, or TensorRT-LLM
- Familiarity with distributed and disaggregated serving infrastructure
- Understanding of prefill, decode, KV-cache behavior, batching, sampling, speculative decoding, and parallelism
- Experience debugging CUDA, NCCL, drivers, kernels, containers, service mesh, networking, and storage
- Python
- PyTorch
- AWS or GCP
- Kubernetes
- CUDA
- NCCL
- InfiniBand
- Nice-to-have experience with CUDA or Triton kernels, Nsight profiling, Rust, C++, Kafka or PubSub, Redis, gRPC or Protobuf, Prometheus or Grafana, OpenTelemetry, Terraform or Ansible, and open-source infrastructure contributions
Responsibilities
- Build a multi-tenant LLM serving platform across cloud GPU fleets
- Design GPU-aware placement and scheduling algorithms
- Implement multi-region and multi-zone failover and traffic shifting
- Build autoscaling, routing, and load balancing
- Optimize model distribution and cold-start times
- Integrate and contribute to LLM inference frameworks
- Optimize parallelism, caching, memory management, quantization, and speculative decoding
- Profile kernels, memory bandwidth, and transport
- Develop reproducible performance suites
- Embed distributed inference into the RL stack
- Establish CI/CD with artifact promotion and performance gates
- Build observability and manage SLOs
- Document architectures and playbooks
- Mentor and collaborate cross-functionally
Benefits
- Significant equity incentives
- Flexible work arrangement
- Full visa sponsorship
- Relocation support
- Professional development budget
- Regular team off-sites
- Conference attendance
