Member of Technical Staff - Compute Platform

Prime Intellect provides an open superintelligence stack for training, evaluating, deploying, and continuously improving AI agents and models. Its platform combines RL environments, hosted training, inference, GPU compute, secure sandboxes, and open-source research tooling for researchers, startups, and enterprises.

Maintainer signals as of 8/23/2026

San Francisco, USA
About Prime Intellect, Inc.

Prime Intellect operates an integrated AI infrastructure platform spanning Lab, hosted reinforcement-learning training, evaluations, environments, inference, secure sandboxes, and on-demand or reserved GPU compute. It also develops open-source tools including Verifiers, prime-rl, and Prime Agent, supporting workflows from environment creation and model evaluation through post-training and production deployment. The company serves researchers, startups, enterprises, and teams building agentic AI systems, with customer examples including Ramp and Zapier.

View jobs by Prime Intellect, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Build platform software and infrastructure for managing and monitoring AI workloads, including web interfaces, Python APIs, backend services, debugging tools, distributed training infrastructure, automation pipelines, cloud resources, container orchesation, and hardware scheduling systems.

Requirements

  • Strong Python backend development with FastAPI and async
  • Modern frontend development with TypeScript, React/Next.js, and Tailwind
  • Experience building developer tools and dashboards
  • RESTful API design and implementation
  • Systems programming experience with Rust
  • Infrastructure automation with Ansible and Terraform
  • Container orchestration with Kubernetes
  • Cloud platform expertise, preferably GCP
  • Observability tools such as Prometheus and Grafana
  • GPU computing or ML infrastructure experience

Responsibilities

  • Build web interfaces for AI workload management and monitoring
  • Develop REST APIs and backend services in Python
  • Create real-time monitoring and debugging tools
  • Implement user-facing features for resource management and job control
  • Design distributed training infrastructure in Rust
  • Build high-performance networking and coordination components
  • Create infrastructure automation pipelines with Ansible
  • Manage cloud resources and container orchestration
  • Implement scheduling systems for CPU, GPU, and TPU hardware
  • Integrate backend features into existing infrastructure