Software Engineer Ray Data
AnyscaleVisit Anyscale website
Anyscale is an AI compute platform built by Ray’s creators for distributed data processing, model training, inference, and related production AI workloads.
San Francisco, United States
About Anyscale
Anyscale provides a managed, multi-cloud platform for building, running, and governing distributed AI workloads with Ray, including data processing, training, and model serving.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will develop open-source Ray software, improve Ray Core and Datasets architecture, strengthen release testing, and communicate your work through talks, tutorials, and blog posts. You will work on large-scale data performance, ML and data-source integrations, streaming workloads, and hosted Ray data operations.
Requirements
- At least 5 years of relevant work experience
- Background in algorithms, data structures, and system design
- Experience building scalable, fault-tolerant distributed systems
- Experience with data processing or database internals, including Spark or Dask
Responsibilities
- Develop high-quality open-source software for Ray
- Identify, implement, and evaluate architectural improvements to Ray Core and Datasets
- Improve Ray testing and release processes
- Communicate work through talks, tutorials, and blog posts
- Optimize large-scale Ray Datasets performance
- Integrate ML training and data sources
- Develop stability and stress-testing infrastructure
- Lead streaming-workload integration initiatives
- Differentiate data operations in the hosted Ray service
