Software Engineer Data Infrastructure

Artificial-intelligence research and product company building customizable AI systems, including the Tinker training API and Inkling open-weight models.

Seed0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

Distributed
About Thinking Machines Lab

Thinking Machines Lab develops AI products that let researchers and developers fine-tune and use models, while also releasing open-weight multimodal models.

View jobs by Thinking Machines Lab

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design, build, and operate fault-tolerant infrastructure for distributed compute, data orchestration, and multimodal storage. You will develop high-throughput ingestion and transformation systems, maintain data catalogs and quality controls, improve monitoring, and work with researchers to accelerate training cycles.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or a similar field
  • Proficiency in at least one backend language, such as Python or Rust
  • Fluency with distributed compute frameworks such as Apache Spark or Ray
  • Deep familiarity with cloud infrastructure, data lake architectures, and batch and streaming pipelines
  • Ability to own projects end-to-end across the stack
  • Hands-on experience with Kafka, dbt, Terraform, and Airflow
  • Experience building a web crawler
  • Experience with deduplication, data mining, and search
  • Knowledge of file formats and storage systems such as Parquet and Delta Lake
  • Experience with documentation, testing, and developer tooling

Responsibilities

  • Design, build, and operate scalable, fault-tolerant infrastructure for distributed compute, data orchestration, and multimodal storage
  • Develop high-throughput systems for data ingestion, processing, transformation, training data catalogs, deduplication, quality checks, and search
  • Build systems for traceability, reproducibility, and quality control throughout the data lifecycle
  • Implement and maintain monitoring and alerting for platform reliability and performance
  • Collaborate with research teams to improve data quality and accelerate training cycles

Benefits

  • Health benefits
  • Dental benefits
  • Vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support as needed