Search...

Machine Learning Engineer

Virtu Financial logo
Virtu Financial

Stealth

Distributed
View jobs by Virtu Financial

Skills

About the Role

In this role, you will design and build the infrastructure that powers quantitative research at scale. You will create experiment tracking, job orchestration, and reproducibility systems so researchers can iterate quickly and recover from failures without losing work. You will build tools for every stage of the simulation lifecycle, including historical back-tests and production monitoring, and you will own visibility into GPU cluster utilization, tracking allocation and surfacing bottlenecks. You will diagnose and resolve performance issues across training pipelines — from data loading throughput and storage I/O to GPU utilization and inter-node communication in distributed training runs. You will build and maintain data pipelines that move financial data into training workflows with strong correctness and versioning guarantees, and develop feature storage and retrieval patterns for fast, reproducible access to training data at scale. You will work directly with researchers to reduce friction in their workflows, collaborate with infrastructure engineers on capacity planning and tooling decisions, and stay current with ML infrastructure developments to bring valuable ideas into the stack.

Requirements

  • 5+ years of experience in ML engineering, research infrastructure, or HPC environments
  • Strong Python engineering skills with clean, maintainable, well-tested code
  • Experience building or operating distributed training infrastructure with knowledge of collective communication libraries (NCCL, Horovod, or similar)
  • Practical experience with experiment tracking systems
  • Comfort working across the Linux systems stack including storage, networking, and job scheduling
  • Excellent communication skills and ability to work closely with researchers and engineers
  • Exposure to C++ in a performance-sensitive context is a plus
  • Experience with on-prem compute environments and job orchestration tools such as Slurm is desired
  • Familiarity with GPU profiling tools (NSight Systems, PyTorch Profiler) is desired
  • Experience with columnar data formats and tools such as Parquet, Arrow, and Polars is desired
  • Familiarity with workflow orchestration tools (Prefect, Dagster, or similar) is desired

Responsibilities

  • Design and build experiment tracking, job orchestration, and reproducibility infrastructure
  • Create tools for all stages of the simulation lifecycle including historical back-tests and production monitoring
  • Own visibility into GPU cluster utilization, track allocation, and surface bottlenecks
  • Diagnose and resolve performance issues across training pipelines
  • Build and maintain data pipelines that move financial data into training workflows with correctness and versioning guarantees
  • Develop feature storage and retrieval patterns for fast, reproducible access to training data at scale
  • Work directly with researchers to identify and reduce workflow friction
  • Collaborate with infrastructure engineers on capacity planning, cloud/on-prem tradeoffs, and tooling decisions
  • Stay current with ML infrastructure tooling and bring relevant ideas into the stack