Machine Learning Engineer ML Training Platform

Pluralis Research is a research lab developing decentralized AI through Protocol Learning, a communication-efficient approach to collaborative model training. It operates open training systems that allow distributed contributors to provide compute for collectively owned foundation models.

Distributed
About Pluralis Research

Pluralis Research develops Protocol Learning, which enables foundation models to be trained and served across globally distributed participants without requiring a single participant to hold the complete model. Its work covers low-bandwidth model parallelism, asynchronous distributed optimization, fault-tolerant training, privacy-preserving unextractable models, and collective ownership. The organization operates Agora and Node0 training systems, publishes research, provides participation documentation, and releases open-source software for distributed training and reinforcement-learning workflows.

View jobs by Pluralis Research

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will architect, build, and scale a platform for continuous experimentation and large-scale distributed training. You will provision and orchestrate multi-cloud compute, develop fault-tolerant ML infrastructure, and build systems that manage real-world network conditions and node churn.

Requirements

  • Production experience with infrastructure as code using Pulumi Terraform or CloudFormation
  • Experience managing multi-cloud deployments Docker Kubernetes EKS GPU workloads and heterogeneous clusters at scale
  • Knowledge of distributed training workflows including checkpointing data sharding model versioning and job orchestration
  • Knowledge of P2P networking NAT traversal traffic shaping and bandwidth constraints
  • Strong Python engineering skills including asyncio concurrency retry logic cloud SDKs and CLI tooling
  • Experience with observability SRE practice Prometheus Grafana performance profiling and incident response
  • Professional-level written and spoken English proficiency

Responsibilities

  • Design resource-management systems for multi-cloud compute provisioning and orchestration
  • Architect fault-tolerant distributed training and inference infrastructure
  • Build systems that simulate and handle bandwidth constraints latency packet loss and node churn

Benefits

  • Equity-heavy compensation package
  • Flexible remote-first work environment
  • Optional visa sponsorship and relocation support to Australia or the US