Founding Machine Learning Engineer

Arcade is an enterprise AI-agent actions runtime for governing, authorizing, executing, and auditing agent tool calls.

Series ARecently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

Burlingame, United States
About Arcade

Arcade AI, Inc. provides runtime infrastructure that lets AI agents securely access business systems under user-specific permissions, with policy enforcement, tool execution, and centralized auditability.

View jobs by Arcade

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own the end-to-end machine learning pipeline, from data ingestion and training through evaluation, release, and production monitoring. You will build and fine-tune models for tool selection, routing, retrieval, and agent memory; create evaluation systems; and deploy optimized models to customer VPCs and air-gapped environments. You will turn telemetry into training data, choose the ML stack, make build-versus-buy decisions, and shape the ML strategy and future hiring.

Requirements

  • 7+ years of software engineering experience
  • 4+ years training and shipping production ML systems
  • Experience training or fine-tuning models that reached production and improved customer-relevant metrics
  • Expertise in production agent systems, including harnesses, memory, skills, tool use, and sub-agents
  • Knowledge of fine-tuning and deriving training data from raw telemetry
  • Statistics fluency for evaluating whether metric changes are meaningful
  • Experience deploying models under latency, GPU, or infrastructure constraints
  • Python for training and ML work
  • TypeScript or Go for production model-serving services
  • Ability to make and defend technical stack decisions

Responsibilities

  • Own the end-to-end ML training pipeline from data through evaluation and release
  • Train and fine-tune models for tool selection, routing, retrieval, and agent memory
  • Expand model use cases for tool and agent-context search and recommendation
  • Build offline and online evaluation systems and compare models with Claude, GPT, and Gemini
  • Quantize, optimize, package, and serve models for customer VPCs and air-gapped environments
  • Turn production agent traces and tool-call telemetry into training data within enterprise data boundaries
  • Choose the ML stack, write the strategy, and inform ML hiring
  • Use expertise in agent systems to decide where models improve agents
  • Use AI tools to increase delivery speed

Benefits

  • Equity
  • Benefits