Data Platform Engineer
Beatdapp provides a Trust & Safety OS that unifies identity, security, content integrity, fraud detection, and engagement integrity. Its primary clients include digital service providers, music distributors, labels, collective management organizations, gaming platforms, streaming services, and creator platforms.
Projects
About Beatdapp Software, Inc.
Beatdapp Software, Inc. develops and operates a unified Trust & Safety OS for digital platforms. Its capabilities include anomaly detection, customer due diligence, AI-generated music detection, account takeover detection, and manipulation-resistant recommendation systems. The platform analyzes large-scale behavioral and streaming data to detect fraud, protect royalties and revenue, verify users, secure accounts, and support trustworthy discovery. Beatdapp primarily serves the music industry, including DSPs, distributors, labels, and CMOs/PROs, while also targeting gaming, media and streaming, creator platforms, and marketplaces.
Skills
About the Role
You'll be operating at the intersection of cloud infrastructure, distributed systems, and data engineering. You will be one of the architects of a system where trillions of data points are ingested, processed, and served as real-time recommendations. You will take full ownership of multi-cluster Kubernetes environments and backend service layers, ensuring API workloads scale seamlessly and data pipelines are robust and reliable. You will bridge the gap between raw streaming data and the clean, high-quality signals the models depend on, ensuring the systems that move, store, and serve data remain fast, secure, and resilient.
Requirements
- 3+ years of professional experience in Backend, DevOps, and/or Data Engineering, preferably supporting data-intensive or ML applications at scale
- Deep familiarity with Kubernetes, including compute instances, network configuration (VPCs/Subnets), and scaling API workloads
- Proficiency in writing clean, scalable backend services and data processing code, primarily in Python
- Comfortable writing production-grade code that handles large-scale data with stream processing, batch ETL, and API development in cloud-native environments
- Proven track record of building automated CI/CD pipelines, managing image registries (Docker/Podman), and handling complex code versioning
- Strong understanding of datastores (relational, non-relational, and columnar), distributed data systems, caching strategies, and data transfer protocols
- Experience with streaming platforms (e.g. Kafka, Pub/Sub) or query engines (e.g. BigQuery, Spark) is a strong asset
- Experience working with sensitive data, encryption, and secure cloud networking
Responsibilities
- Design, build, and maintain scalable data pipelines and processing workflows that move and transform high-volume streaming data
- Optimize batch and streaming workloads, manage data quality at ingestion, and ensure reliable delivery to downstream consumers including ML feature stores and serving layers
- Manage and optimize multi-cluster Kubernetes (K8s) environments
- Implement sophisticated autoscaling policies and node management strategies to support high-availability ML workloads
- Design and orchestrate live service deployments using strategies such as A/B testing and Canary releases
- Ensure the system supports seamless rollbacks and API versioning
- Design and maintain infrastructure using Infrastructure as Code principles to ensure environment consistency and rapid disaster recovery
- Own logging, tracing, and metrics components across backend services and data pipelines
- Work with Ops teams to define SLOs/error budgets, build dashboards, and maintain health monitoring systems
- Partner with security teams to enforce patch management, secrets handling, and data encryption protocols
- Automate routine operational tasks and environment provisioning
- Act as a primary stakeholder for system uptime, managing outages with clear communication
