Network Engineer Supercomputing
Thinking Machines LabVisit Thinking Machines Lab website
Artificial-intelligence research and product company building customizable AI systems, including the Tinker training API and Inkling open-weight models.
Distributed
Funding history
About Thinking Machines Lab
Thinking Machines Lab develops AI products that let researchers and developers fine-tune and use models, while also releasing open-weight multimodal models.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will validate GPU network-fabric designs, debug production collective and interconnect failures, and own NVLink and NVSwitch health. You will build network instrumentation, dashboards, and alerts; investigate issues across cloud environments; and drive provider escalations through resolution.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, or a similar field
- Proficiency in at least one backend language
- Experience operating large-scale clusters and container orchestration systems
- Ability to operate across the stack and own projects end-to-end
Responsibilities
- Validate GPU network-fabric design across deployments
- Debug RDMA and RoCEv2 across NIC vendors
- Diagnose NCCL collective failures, PFC and ECN tuning, and congestion-control behavior
- Own NVLink and NVSwitch interconnect health
- Build host-level network instrumentation, dashboards, and alerts
- Triage fabric issues across NIC, driver, kernel, switch, and workload boundaries
- Drive cloud-provider networking escalations through resolution
Benefits
- Visa sponsorship
- Health benefits
- Dental benefits
- Vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
