Search...

Data & ML Engineer

Red Cell Partners logo
Red Cell Partners

Venture studio that launches, operates, and scales technology companies across healthcare, cyber, and national security.

McLean, Virginia, United States
About Red Cell Partners

Red Cell Partners is a McLean, Virginia-based venture studio founded in 2020. It incubates and scales mission-critical technology companies across three practice areas — healthcare, cyber, and national security — with 15 companies created, 12 publicly launched, and over $600M raised across the firm and its incubations.

View jobs by Red Cell Partners

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build the data and model layer for an AI-enabled decision-support system in an accredited environment. You will ingest and transform structured, semi-structured, and unstructured data, resolve records against a shared data model, develop calibrated relevance models, implement retrieval and grounded generation, and package, serve, version, and monitor models and pipelines. You will also establish provenance, auditability, validation, and secure release practices.

Requirements

  • 5+ years of experience in data engineering, data architecture, applied machine learning, ML engineering, or production analytics engineering
  • Strong Python and SQL skills
  • Experience delivering systems for sustained operational use
  • Experience working with large, imperfect operational data
  • Routine use of AI-assisted development
  • US citizenship
  • Active US Secret clearance
  • CAC eligibility
  • Willingness to travel up to 25%
  • Preferred experience with probabilistic matching, graph data modeling, PostgreSQL, pgvector, scikit-learn, XGBoost, PyTorch, retrieval-augmented generation, AWS Glue, Airflow, dbt, Spark, Kafka, or NiFi

Responsibilities

  • Design and maintain graphs of entities, records, and typed relationships
  • Implement probabilistic matching, candidate generation, scoring, clustering, and threshold policies
  • Build deduplication and known-record suppression
  • Establish provenance for source records
  • Develop relevance and priority models
  • Design calibration, threshold, and abstention policies
  • Perform feature engineering and error analysis
  • Implement embeddings, vector storage, retrieval, and grounded generation
  • Build secure ingestion, transformation, validation, and publishing pipelines
  • Implement quality checks, schema validation, lineage capture, audit logging, and source drift detection
  • Package, serve, version, and roll back models
  • Maintain audit trails and submit changes through gated release processes

Benefits

  • Fully remote work
  • Employer-paid medical, dental, and vision insurance for employees and their families
  • Unlimited paid time off
  • Flexible work schedule
  • 14 weeks of fully paid parental leave
  • 401(k) option
  • FSA
  • Equity incentives
  • Mental health benefits
  • GLP-1 solutions