Network Engineer Supercomputing

Artificial-intelligence research and product company building customizable AI systems, including the Tinker training API and Inkling open-weight models.

Distributed
About Thinking Machines Lab

Thinking Machines Lab develops AI products that let researchers and developers fine-tune and use models, while also releasing open-weight multimodal models.

View jobs by Thinking Machines Lab

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will validate GPU network-fabric designs, debug production collective and interconnect failures, and own NVLink and NVSwitch health. You will build network instrumentation, dashboards, and alerts; investigate issues across cloud environments; and drive provider escalations through resolution.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or a similar field
  • Proficiency in at least one backend language
  • Experience operating large-scale clusters and container orchestration systems
  • Ability to operate across the stack and own projects end-to-end

Responsibilities

  • Validate GPU network-fabric design across deployments
  • Debug RDMA and RoCEv2 across NIC vendors
  • Diagnose NCCL collective failures, PFC and ECN tuning, and congestion-control behavior
  • Own NVLink and NVSwitch interconnect health
  • Build host-level network instrumentation, dashboards, and alerts
  • Triage fabric issues across NIC, driver, kernel, switch, and workload boundaries
  • Drive cloud-provider networking escalations through resolution

Benefits

  • Visa sponsorship
  • Health benefits
  • Dental benefits
  • Vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support