RL Systems Infra

Artificial-intelligence research and product company building customizable AI systems, including the Tinker training API and Inkling open-weight models.

Seed0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

Distributed
About Thinking Machines Lab

Thinking Machines Lab develops AI products that let researchers and developers fine-tune and use models, while also releasing open-weight multimodal models.

View jobs by Thinking Machines Lab

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design, build, and optimize infrastructure for large-scale reinforcement learning and post-training workloads. You will improve the reliability, scalability, and throughput of distributed RL training pipelines. You will develop monitoring and observability tools, collaborate with researchers to productionize algorithmic ideas, and build evaluation and benchmarking systems for helpfulness, safety, and factuality. You will share technical learnings through documentation, open-source libraries, or reports.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or a similar field
  • Strong engineering skills and ability to write performant, maintainable code and debug complex codebases
  • Understanding of deep learning frameworks such as PyTorch and JAX and their system architectures
  • Experience collaborating with cross-functional partners and subject matter experts
  • Initiative to work across stacks and teams to deliver work
  • Experience training or supporting large-scale language models
  • Experience with reinforcement learning workloads such as PPO, DPO, RLHF, or reward modeling
  • Background in high-performance or reliability engineering, distributed training frameworks, or cluster orchestration
  • Familiarity with monitoring and observability tools such as Prometheus, Grafana, and OpenTelemetry
  • Contributions to large-scale ML research or infrastructure, open-source frameworks, or performance optimization

Responsibilities

  • Design, build, and optimize infrastructure for large-scale reinforcement learning and post-training workloads
  • Improve the reliability, scalability, and throughput of RL training pipelines
  • Develop monitoring and observability tools for uptime, debuggability, and reproducibility
  • Translate algorithmic ideas into production-grade training pipelines
  • Build evaluation and benchmarking infrastructure for helpfulness, safety, and factuality
  • Publish technical learnings through documentation, open-source libraries, or technical reports

Benefits

  • Health benefits
  • Dental benefits
  • Vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support as needed
  • Visa sponsorship