Applied Research Evals and Data

Prime Intellect provides an open superintelligence stack for training, evaluating, deploying, and continuously improving AI agents and models. Its platform combines RL environments, hosted training, inference, GPU compute, secure sandboxes, and open-source research tooling for researchers, startups, and enterprises.

Maintainer signals as of 8/23/2026

San Francisco, USA
About Prime Intellect, Inc.

Prime Intellect operates an integrated AI infrastructure platform spanning Lab, hosted reinforcement-learning training, evaluations, environments, inference, secure sandboxes, and on-demand or reserved GPU compute. It also develops open-source tools including Verifiers, prime-rl, and Prime Agent, supporting workflows from environment creation and model evaluation through post-training and production deployment. The company serves researchers, startups, enterprises, and teams building agentic AI systems, with customer examples including Ramp and Zapier.

View jobs by Prime Intellect, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Work directly with customers to understand workflows, data sources, and bottlenecks; prototype AI agents, data pipelines, and evaluation harnesses; translate customer insights into research and product direction; design post-training and reinforcement-learning methods; build evaluation and verification systems; integrate applied data into model improvement; develop agent capabilities, distributed training and inference pipelines, and production observability.

Requirements

  • Machine learning engineering experience in post-training, reinforcement learning, or large-scale model alignment
  • Experience with applied data workflows and evaluation frameworks for large models or agents
  • Expertise in distributed training and inference frameworks such as vLLM, sglang, Ray, or Accelerate
  • Experience deploying containerized systems at scale using Docker, Kubernetes, and Terraform
  • Research contributions through publications, open-source contributions, or benchmarks
  • Knowledge of reasoning, measurement, and agentic AI systems

Responsibilities

  • Work with customers to understand workflows, data sources, and bottlenecks
  • Prototype agents, data pipelines, and evaluation harnesses for customer use cases
  • Translate customer insights and evaluation results into roadmap and research direction
  • Design and implement RL and post-training methods for domain-specific tasks
  • Build evaluation harnesses and verifiers for reasoning, robustness, and agentic behavior
  • Integrate applied data collection and analytics into post-training workflows
  • Prototype multi-agent and memory-augmented systems
  • Extend and integrate agent frameworks
  • Architect and maintain distributed training and inference pipelines
  • Develop observability and monitoring for production deployments

Benefits

  • Equity incentives
  • Flexible work, remote or San Francisco
  • Visa sponsorship
  • Relocation support
  • Professional development budget
  • Team off-sites
  • Conference attendance
Applied Research Evals and Data at Prime Intellect, Inc. | JobStash