Member of Technical Staff Research Infrastructure Engineer
Black Forest LabsVisit Black Forest Labs website
Frontier AI research lab building the FLUX family of multimodal visual-intelligence models and delivering them through an API, playground, and open weights.
Freiburg im Breisgau, Germany
Funding history
About Black Forest Labs
Black Forest Labs develops generative AI models for image, video, audio, and action prediction, alongside production API and enterprise deployment offerings.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will maintain and scale research infrastructure for large-scale training workloads. You will optimize application and infrastructure performance, work with research teams on cost-efficient solutions, diagnose distributed-system bottlenecks, build telemetry and monitoring, and participate in on-call incident response.
Requirements
- Experience building or operating large-scale training platforms
- Experience with large-scale GPU compute clusters
- Ability to debug performance and reliability issues across distributed fleets
- Knowledge of Kubernetes, Infrastructure as Code, AWS, and GCP
- Experience with SLURM
- Knowledge of Python, Bash, Go, NVIDIA GPU drivers and operators, OpenTelemetry, and Prometheus
Responsibilities
- Maintain and optimize research infrastructure
- Scale infrastructure while maintaining reliability and performance
- Design cost-efficient infrastructure solutions with research teams
- Identify and resolve distributed-system performance bottlenecks
- Build and evolve telemetry and monitoring systems
- Participate in on-call rotations and incident response
Benefits
- Equity
- Reasonable travel costs covered
- Relocation encouraged but not required
