Member of Technical Staff Research Infrastructure Engineer

Frontier AI research lab building the FLUX family of multimodal visual-intelligence models and delivering them through an API, playground, and open weights.

Freiburg im Breisgau, Germany
About Black Forest Labs

Black Forest Labs develops generative AI models for image, video, audio, and action prediction, alongside production API and enterprise deployment offerings.

View jobs by Black Forest Labs

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will maintain and scale research infrastructure for large-scale training workloads. You will optimize application and infrastructure performance, work with research teams on cost-efficient solutions, diagnose distributed-system bottlenecks, build telemetry and monitoring, and participate in on-call incident response.

Requirements

  • Experience building or operating large-scale training platforms
  • Experience with large-scale GPU compute clusters
  • Ability to debug performance and reliability issues across distributed fleets
  • Knowledge of Kubernetes, Infrastructure as Code, AWS, and GCP
  • Experience with SLURM
  • Knowledge of Python, Bash, Go, NVIDIA GPU drivers and operators, OpenTelemetry, and Prometheus

Responsibilities

  • Maintain and optimize research infrastructure
  • Scale infrastructure while maintaining reliability and performance
  • Design cost-efficient infrastructure solutions with research teams
  • Identify and resolve distributed-system performance bottlenecks
  • Build and evolve telemetry and monitoring systems
  • Participate in on-call rotations and incident response

Benefits

  • Equity
  • Reasonable travel costs covered
  • Relocation encouraged but not required