Distributed LLM Inference Engineer
AnyscaleVisit Anyscale website
Anyscale is an AI compute platform built by Ray’s creators for distributed data processing, model training, inference, and related production AI workloads.
San Francisco, United States
About Anyscale
Anyscale provides a managed, multi-cloud platform for building, running, and governing distributed AI workloads with Ray, including data processing, training, and model serving.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will deliver end-to-end batch and online inference solutions at scale. You will integrate Ray Data and LLM engines, optimize cost and performance, work with open-source communities, contribute improvements, and apply current research and engineering practices to large-scale ML inference.
Requirements
- Experience running large-scale ML inference with high throughput and low latency
- Knowledge of deep learning and deep learning frameworks such as PyTorch
- Understanding of distributed systems and ML inference challenges
Responsibilities
- Ship end-to-end batch and online inference solutions at high scale
- Integrate Ray Data and LLM engines
- Optimize large-scale ML inference for cost and performance
- Integrate vLLM and contribute improvements to open-source software
- Follow and implement state-of-the-art open-source and research practices
Benefits
- Stock options
- Healthcare plans with 99% employer-covered premiums for employees and dependents
- 401k Retirement Plan
- Education & Wellbeing Stipend
- Paid Parental Leave
- Fertility Benefits
- Paid Time Off
- Commute reimbursement
- In-office meals
