Member of Technical Staff - Multilingual Data
Reflection is an AI research lab building open frontier models and a full AI stack for developers, enterprises, and public-sector users.
Funding history
About Reflection
Reflection develops open-weight AI models, open-source software for customizing and running agents, AI-factory infrastructure, and related solutions. Its current research emphasizes large language models, reinforcement learning, and agentic reasoning.
Skills
About the Role
You will design and operate multilingual data pipelines for sourcing, cleaning, deduplication, language identification, and script normalization. You will establish corpus quality standards, run experiments on data efficiency, build language-specific evaluations and diagnostics, lead research projects, and collaborate to improve multilingual model capability.
Requirements
- Software engineering
- Distributed computing
- Data pipeline
- Language model
- Machine translation
- Speech processing
- Search
- Multilingualism
- Statistical analysis
- Experimentation
Responsibilities
- Design and operate multilingual data pipelines
- Source clean deduplicate and normalize multilingual data
- Define and enforce quality standards for multilingual corpora
- Run experiments on multilingual data efficiency
- Lead small research projects
- Build evaluation sets and diagnostics across languages registers and domains
- Collaborate to improve multilingual model capability
Benefits
- Stock options
- Medical insurance
- Dental insurance
- Vision insurance
- Life insurance
- Annual wellness allowance
- Daily office lunch and dinner
- 22 weeks of paid parental leave
- Unlimited paid time off in the U.S.
- 30 days of vacation in the U.K.
- Visa sponsorship
- Regular off-sites
- Happy hours
- Team celebrations
