Search...

Applied Research Evals and Data

Prime Intellect logo
Prime Intellect

Prime Intellect builds the 'Open Superintelligence Stack' — an integrated compute, training, inference, and sandbox platform that lets companies train, deploy, and continuously improve their own AI models and agents. It serves AI startups, 'neolabs', and enterprises (over 6,000 customers, including Ramp and Zapier) that want to own their model optimization loop rather than rely solely on closed frontier labs.

San Francisco, USA
About Prime Intellect

Prime Intellect is a San Francisco-based AI infrastructure company building what it calls the Open Superintelligence Stack: a full-stack platform spanning GPU compute (on-demand and reserved clusters), large-scale reinforcement learning training ('Lab'), an Environments Hub with 2,500+ community RL environments, hosted evaluations, sandboxed code execution, and dedicated/serverless model inference with native LoRA support. The company maintains open-source libraries (verifiers and prime-rl) used to build and train RL environments, and publishes frontier open research such as the INTELLECT and SYNTHETIC model/dataset series. Prime Intellect works with AI startups, enterprises, and 'neolab' customers such as Ramp and Zapier, helping them turn production traces and evaluations into custom-trained, post-trained agent models that outperform closed frontier models on specific workflows at lower cost and latency. The company has raised over $150M in total funding, including a $130M Series A led by Radical Ventures with participation from NVIDIA Ventures, Intel Capital, and Dell Technologies Capital, and reports over $100M in annualized revenue.

View jobs by Prime Intellect

Skills

About the Role

You will work directly with customers to understand their workflows, data sources, and bottlenecks. You will prototype AI agents, data pipelines, and evaluation harnesses, then translate customer insights and evaluation results into research and product direction. You will design post-training and reinforcement-learning methods, build evaluation and verification systems, and integrate applied data into model improvement. You will also develop agent capabilities, distributed training and inference pipelines, and production observability.

Requirements

  • Machine learning engineering experience in post-training, reinforcement learning, or large-scale model alignment
  • Experience with applied data workflows and evaluation frameworks for large models or agents
  • Expertise in distributed training and inference frameworks such as vLLM, sglang, Ray, or Accelerate
  • Experience deploying containerized systems at scale using Docker, Kubernetes, and Terraform
  • Research contributions through publications, open-source contributions, or benchmarks
  • Knowledge of reasoning, measurement, and agentic AI systems

Responsibilities

  • Work with customers to understand workflows, data sources, and bottlenecks
  • Prototype agents, data pipelines, and evaluation harnesses for customer use cases
  • Translate customer insights and evaluation results into roadmap and research direction
  • Design and implement RL and post-training methods for domain-specific tasks
  • Build evaluation harnesses and verifiers for reasoning, robustness, and agentic behavior
  • Integrate applied data collection and analytics into post-training workflows
  • Prototype multi-agent and memory-augmented systems
  • Extend and integrate agent frameworks
  • Architect and maintain distributed training and inference pipelines
  • Develop observability and monitoring for production deployments

Benefits

  • Equity incentives
  • Flexible work, remote or San Francisco
  • Visa sponsorship
  • Relocation support
  • Professional development budget
  • Team off-sites
  • Conference attendance