Data Developer
DRW is a Chicago-based diversified trading firm that uses technology, quantitative research, and risk management to provide liquidity across traditional and digital-asset markets.
About DRW
Founded by Don Wilson in 1992, DRW trades for its own account across global markets and operates strategies in cryptoassets, venture capital, real estate, carbon markets, and public-equity investments. Its crypto business, Cumberland, has provided institutional crypto-asset liquidity since 2014.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Design and build data pipelines for retrieval-augmented generation systems, ingest structured and unstructured data, optimize datasets for AI workloads, and deploy vector databases, embedding services, and processing pipelines. Monitor retrieval quality, improve data quality and latency, and collaborate with ML engineers on inference optimization.
Requirements
- Bachelor's or Master's degree in Computer Science, Data Engineering, or a related field
- 2-5 years building data systems and pipelines in production environments
- Experience with RAG architectures and vector databases
- Proficiency in Python and DAG-based orchestration platforms
- Experience with embedding models and semantic search systems
- Experience with Apache Spark, Ray, or Dask
- Understanding of LLM inference optimization and prompt engineering
- Familiarity with Docker, containerization, and orchestration platforms
- Knowledge of data modeling, ETL/ELT patterns, and data quality
Responsibilities
- Design and build data pipelines for RAG systems, including document ingestion, chunking, embedding generation, and vector storage.
- Build ingestion pipelines for structured and unstructured data sources into a centralized data lake.
- Develop data processing workflows to prepare and optimize datasets for fine-tuning and inference workloads.
- Build monitoring and evaluation frameworks for retrieval quality, latency, and system performance.
- Collaborate with ML engineers to optimize data formats and storage patterns for GPU-accelerated inference.
- Implement caching strategies and data versioning systems to support efficient model serving.
- Deploy and manage vector databases, embedding services, and data processing pipelines.
- Improve data quality, reduce latency, and enhance retrieval accuracy.
- Stay current with emerging data engineering and AI technologies.
- Contribute ideas for tools, process improvements, and technology adoption.
