Software Engineer ML Ops and Platform
Tools for Humanity is a technology company building products for humans in the age of AI.
Maintainer signals as of 9/2/2026
About Tools for Humanity
Tools for Humanity develops hardware, software, and services connected to World, described as an open-source real-human network. Its products include the Orb for privacy-focused proof-of-human verification, World App for World ID access, digital asset management, wallet functionality, transactions, and Mini Apps, and World Spaces for verification, community programming, and events. The company serves end users and organizations seeking to integrate World infrastructure into their projects.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own the machine learning lifecycle from data to device. You will design and operate production pipelines for training evaluation telemetry and deployment, build self-service platforms, enable staged rollouts and rollback, expose governed datasets and model artifacts through secure services, and implement monitoring lineage reproducibility and privacy controls.
Requirements
- 5+ years of experience building ML infrastructure data platforms or production ML systems at scale
- Experience delivering platforms and CI/CD pipelines used by ML or data teams
- Experience running large-scale training on multi-tenant GPU clusters
- Experience building versioned dataset and lineage systems with slice-level provenance and governed access
- Deep understanding of Docker Kubernetes or EKS and Infrastructure-as-Code tools such as Terraform CDK or CloudFormation
- Strong backend engineering skills in Python and/or Go
- Understanding of modern CI/CD model packaging and observability practices
- Experience operating production systems defining SLAs and handling rollout or incident workflows
- Comfort using modern agentic AI development
Responsibilities
- Design and operate infrastructure for training evaluation telemetry ingestion and deployment
- Maintain CI/CD workflows and automated pipelines
- Build edge-aware rollout services with staged deployment A/B experimentation and rollback
- Develop secure APIs and backend services for governed datasets and model artifacts
- Implement automated checks drift detection and real-time model monitoring
- Establish practices for data lineage reproducibility privacy security and secure edge delivery
- Collaborate with ML research product and firmware teams to improve delivery workflows
