Senior or Staff ML Systems Engineer, LLMs
TRM Labs provides a blockchain intelligence platform to help organizations investigate, monitor, and detect crypto and digital asset fraud and financial crime. They serve government agencies, financial institutions, and crypto businesses worldwide.
Investors
Projects
About TRM Labs
TRM Labs provides a next-generation blockchain intelligence platform designed to investigate, monitor, and detect crypto and digital asset fraud and financial crime. The platform features extensive asset coverage, supporting over 200 million assets across more than 41 blockchains, including NFTs and DeFi protocols. It offers cross-chain analytics to trace the flow of funds seamlessly between different blockchains and utilizes over 150 risk categories, including FATF's money laundering predicate offenses, for customized risk scoring. TRM's data is built from a large, proprietary database of illicit activity combined with advanced data science. The company serves a global client base, including government agencies, financial institutions, and crypto businesses, helping them to safeguard the crypto financial system, maintain high standards for AML/CFT compliance, and build trust in digital assets.
Skills
About the Role
You will build and scale the technical infrastructure that powers large language models and agentic systems. You will create reusable CI/CD workflows for model training, evaluation, and deployment, automate model versioning and approval workflows, and implement compliance checks. You will design and operate modular AI infrastructure—vector databases, feature stores, model registries, and observability tooling—and embed models and agents into real-time applications. You will continuously evaluate and integrate state-of-the-art tools, monitor cost, latency, and performance, and run offline and online evaluation pipelines including regression tests and human-in-the-loop workflows. You will enable researchers by providing sandboxes, dashboards, and reproducible environments, and ensure data accuracy and reliability for model training and inference.
Requirements
- Write high quality maintainable software primarily in Python
- Strong background in scalable infrastructure including containerization and orchestration (Docker Kubernetes)
- Experience with infrastructure as code and deployment (Terraform CI/CD pipelines)
- Familiarity with monitoring and logging frameworks (Datadog Prometheus OpenTelemetry)
- Knowledge of MLOps best practices including model versioning rollback strategies automated evaluation and drift detection
- Experience with scalable model and agent serving infrastructure (vLLM Triton BentoML)
- Experience deploying and maintaining LLM and agentic workflows in production including monitoring cost latency and performance
- Ability to capture traces for analysis and optimize prompt response flows with real time data access
- Strong ownership pragmatism and ability to balance infrastructure elegance with iterative delivery
Responsibilities
- Build reusable CI/CD workflows for model training evaluation and deployment
- Automate model versioning approval workflows and compliance checks
- Design and maintain modular scalable AI infrastructure including vector databases feature stores model registries and observability tooling
- Embed AI models and agents into real time applications and workflows
- Evaluate and integrate state of the art AI tools and libraries
- Drive AI reliability governance and ensure compliance security and uptime
- Deploy infrastructure for offline and online evaluation including regression testing cost monitoring and human in the loop workflows
- Provide sandboxes dashboards and reproducible environments to accelerate research
- Ensure data accuracy consistency and reliability for model training and inferencing
Benefits
- Equity plan
- Remote work
