Member of Technical Staff - Compute Platform

AI infrastructure company providing an integrated stack for training, evaluating, deploying, and continuously improving agentic models.

Series ARecently funded34 current maintainers27 active leads7 new active leads9 lead step-downsTeam intelligence

Maintainer signals as of 9/25/2026

San Francisco, United States
About Prime Intellect

Prime Intellect, Inc. operates AI infrastructure spanning RL environments, hosted training and evaluations, inference, secure sandboxes, and globally sourced GPU compute.

View jobs by Prime Intellect

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Build platform software and infrastructure for managing and monitoring AI workloads, including web interfaces, Python APIs, backend services, debugging tools, distributed training infrastructure, automation pipelines, cloud resources, container orchesation, and hardware scheduling systems.

Requirements

  • Strong Python backend development with FastAPI and async
  • Modern frontend development with TypeScript, React/Next.js, and Tailwind
  • Experience building developer tools and dashboards
  • RESTful API design and implementation
  • Systems programming experience with Rust
  • Infrastructure automation with Ansible and Terraform
  • Container orchestration with Kubernetes
  • Cloud platform expertise, preferably GCP
  • Observability tools such as Prometheus and Grafana
  • GPU computing or ML infrastructure experience

Responsibilities

  • Build web interfaces for AI workload management and monitoring
  • Develop REST APIs and backend services in Python
  • Create real-time monitoring and debugging tools
  • Implement user-facing features for resource management and job control
  • Design distributed training infrastructure in Rust
  • Build high-performance networking and coordination components
  • Create infrastructure automation pipelines with Ansible
  • Manage cloud resources and container orchestration
  • Implement scheduling systems for CPU, GPU, and TPU hardware
  • Integrate backend features into existing infrastructure