Search...

Machine Learning Data Analyst

Incode Technologies logo
Incode Technologies

Incode is a leading provider of world-class identity solutions, reinventing online identity authentication and verification. Serving industries like finance, government, and retail, Incode's technology reduces fraud and enhances digital trust, aiming to create a single, secure identity for users everywhere.

San Francisco, USA
About Incode Technologies

Incode provides an AI-driven identity verification and biometric authentication platform, the Omni Platform, designed to eliminate fraud, ensure KYC/AML compliance, and streamline user experiences. The platform offers modular solutions for customer, business, and employee identity verification through SDKs, APIs, and a configurable dashboard. Built entirely in-house, Incode's technology leverages an advanced fraud network and features like liveness detection, document verification, and facial recognition, which can be matched against government biometric databases for high accuracy. The company serves a wide range of industries, including financial services, marketplaces, hospitality, gaming, and healthcare, aiming to power a world of trust with a single, secure identity for everyone.

View jobs by Incode Technologies

Skills

About the Role

You will design, build, and maintain automated data pipelines for collection, labeling, validation, and metric computation to support ML training and evaluation. You will establish and monitor data and labeling quality standards, perform accuracy audits and root-cause analysis, and implement automated model evaluation metrics and reporting. You will build scalable systems for performance tracking, dashboards, and monitoring, develop and operate workflow orchestration (Airflow, Prefect, or similar), write clean Python and performant SQL for large datasets (including Redshift and related AWS tooling), and collaborate with ML engineers, analysts, and product stakeholders to prioritize work and unblock execution.

Requirements

  • 3+ years of experience as a Data Analyst or in a similar data infrastructure role
  • Strong Python programming skills with focus on clean, maintainable code
  • Solid SQL expertise and experience with cloud or columnar databases such as AWS Redshift
  • Hands-on experience with workflow orchestration tools such as Airflow, Prefect, or Dagster
  • Proven experience in data quality management, data preparation, or ML data pipelines
  • Understanding of metric computation, data labeling, and automation in ML workflows
  • Strong collaboration and problem-solving skills
  • Background in mathematics, physics, or engineering

Responsibilities

  • Design, build, and maintain automated data pipelines for collection, labeling, validation, and metric computation that support ML training and evaluation
  • Establish and monitor data and labeling quality standards and drive consistency checks, accuracy audits, and root-cause analysis
  • Define, implement, and automate model evaluation metrics and reporting that reflect real-world product use cases and business goals
  • Build scalable systems for performance tracking, dashboards, and monitoring to enable fast, data-driven decisions
  • Develop and operate reliable workflow orchestration (Airflow, Prefect, or similar) to schedule, observe, and troubleshoot end-to-end pipelines
  • Write clean, maintainable Python code and performant SQL to process large datasets, leveraging AWS Redshift and related AWS tooling
  • Partner closely with ML engineers, analysts, and product stakeholders to prioritize work, unblock execution, and improve internal tooling for analysis and evaluation

Benefits

  • Flexible working hours and workplace
  • Open vacation policy