Machine Learning Engineer II - Inference Platform
Beatdapp provides a Trust & Safety OS that unifies identity, security, content integrity, fraud detection, and engagement integrity. Its primary clients include digital service providers, music distributors, labels, collective management organizations, gaming platforms, streaming services, and creator platforms.
Projects
About Beatdapp Software, Inc.
Beatdapp Software, Inc. develops and operates a unified Trust & Safety OS for digital platforms. Its capabilities include anomaly detection, customer due diligence, AI-generated music detection, account takeover detection, and manipulation-resistant recommendation systems. The platform analyzes large-scale behavioral and streaming data to detect fraud, protect royalties and revenue, verify users, secure accounts, and support trustworthy discovery. Beatdapp primarily serves the music industry, including DSPs, distributors, labels, and CMOs/PROs, while also targeting gaming, media and streaming, creator platforms, and marketplaces.
Skills
About the Role
You will sit at the intersection of ML engineering, platform and infrastructure work, and inference systems, bridging the gap between raw audio and the clean signals detection models depend on. You will partner with data scientists to bring models into production, bringing a production lens on latency, cost, and customer-facing edges into design conversations early. Your work will span GPU-bound inference containers, multi-cloud infrastructure, the API layer in front of the service, data and observability tooling, and the CI that ships it all. You will be trusted to make architectural judgment calls that contain drift and help the systems scale with minimal code, in an environment where roadmaps are measured in weeks rather than quarters and your scope will grow and shift alongside the team and systems.
Requirements
- Related STEM degree (BSc, MSc, or higher) and 3+ years of work experience in platform / infra / backend / ML / applied-ML / data engineering
- Ability to write clean, scalable, production-grade code in Python or performance-oriented languages (Go, Rust, C++)
- Architectural fluency across data stores, distributed systems, caching, and data transfer protocols
- Comfort building data processing pipelines and using SQL (Airflow, BigQuery, Postgres)
- Deep cloud infrastructure and networking experience across GCP and/or AWS
- Comfort with MLflow or similar tooling and model lifecycle processes
- Ability to write and modify Terraform modules and understand state and backends
- CI/CD discipline including cloud OIDC, image signing, and pinned versions
- Observability instincts across hardware, application, and model layers
- Comfort with inference performance tuning levers (micro-batching, concurrency, request queueing)
- Strong written communication for runbooks, design docs, PR descriptions, postmortems, and ticket hygiene
Responsibilities
- Build, tune, and ship inference containers including Dockerfiles and dependencies
- Manage image size, cold-starts, and GPU access patterns
- Operate multi-cloud orchestration (ECS, Cloud Run, GKE, EKS) for containers
- Maintain test coverage for the container surface and storage abstraction
- Tune concurrency, VRAM accounting, request timeouts, queueing, and rate limiting
- Handle multi-GPU distribution and right-sizing decisions
- Build and run scale and stress testing scenarios across mock deployments
- Characterize latency-vs-throughput curves and find breaking points
- Turn test results into autoscaling and instance-sizing decisions
- Operate the Terraform stack across multiple clouds (GCP, AWS)
- Manage networking, identity, GPU nodes, autoscaling, and per-tenant configurations
- Build and extend the customer-facing API layer including authentication, rate limiting, and data isolation
- Maintain and extend data orchestration pipelines for model evaluation and reporting
- Build and tune metrics, dashboards, logging, and alarms across inference service, instances, and models
