Distributed LLM Inference Engineer

Anyscale is an AI compute platform built by Ray’s creators for distributed data processing, model training, inference, and related production AI workloads.

San Francisco, United States
About Anyscale

Anyscale provides a managed, multi-cloud platform for building, running, and governing distributed AI workloads with Ray, including data processing, training, and model serving.

View jobs by Anyscale

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will deliver end-to-end batch and online inference solutions at scale. You will integrate Ray Data and LLM engines, optimize cost and performance, work with open-source communities, contribute improvements, and apply current research and engineering practices to large-scale ML inference.

Requirements

  • Experience running large-scale ML inference with high throughput and low latency
  • Knowledge of deep learning and deep learning frameworks such as PyTorch
  • Understanding of distributed systems and ML inference challenges

Responsibilities

  • Ship end-to-end batch and online inference solutions at high scale
  • Integrate Ray Data and LLM engines
  • Optimize large-scale ML inference for cost and performance
  • Integrate vLLM and contribute improvements to open-source software
  • Follow and implement state-of-the-art open-source and research practices

Benefits

  • Stock options
  • Healthcare plans with 99% employer-covered premiums for employees and dependents
  • 401k Retirement Plan
  • Education & Wellbeing Stipend
  • Paid Parental Leave
  • Fertility Benefits
  • Paid Time Off
  • Commute reimbursement
  • In-office meals
Distributed LLM Inference Engineer at Anyscale | JobStash