Software Engineer Inference

Artificial-intelligence research and product company building customizable AI systems, including the Tinker training API and Inkling open-weight models.

Seed0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

Distributed
About Thinking Machines Lab

Thinking Machines Lab develops AI products that let researchers and developers fine-tune and use models, while also releasing open-weight multimodal models.

View jobs by Thinking Machines Lab

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will operate and scale production inference systems serving live traffic. You will manage safe model rollouts, improve observability and capacity planning, lead incident response, productionize serving techniques, and design resilient failover and redundancy while balancing infrastructure cost and capacity.

Requirements

  • Experience operating large-scale, latency-sensitive production systems
  • Proficiency in Python and Go or another systems language
  • Experience with observability, monitoring, and incident response for production services
  • Strong understanding of distributed systems and failure modes at scale
  • Experience running production inference for large language models or large-scale ML systems
  • Experience with deployment and rollout systems such as canarying, blue/green deployments, or feature flags
  • Experience with capacity planning and cost optimization for GPU or TPU infrastructure
  • Familiarity with batching, caching, or quantization for inference

Responsibilities

  • Operate and scale production inference systems serving live traffic
  • Own safe, incremental rollouts of new models, model versions, and inference optimizations
  • Build and improve observability, alerting, and capacity planning
  • Productionize new serving techniques while maintaining reliability
  • Lead incident response, root cause analysis, and durable fixes for production issues
  • Design graceful degradation, failover, and redundancy
  • Manage capacity and cost tradeoffs for serving infrastructure

Benefits

  • Health benefits
  • Dental benefits
  • Vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support as needed