Research Program Manager Research Infrastructure
Reflection is an AI research lab building open frontier models and a full AI stack for developers, enterprises, and public-sector users.
Funding history
About Reflection
Reflection develops open-weight AI models, open-source software for customizing and running agents, AI-factory infrastructure, and related solutions. Its current research emphasizes large language models, reinforcement learning, and agentic reasoning.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead cross-functional programs that scale training infrastructure and cluster reliability across pre-training, mid-training, and post-training. You will coordinate engineering leads and external partners, triage incidents, resolve escalations, and establish durable processes for handoffs, configuration management, and checkpoint workflows. You will identify bottlenecks, track training and infrastructure health, and translate technical complexity into clear updates and decision frameworks for leadership.
Requirements
- 7+ years of experience in technical program management, research operations, or infrastructure coordination
- Technical knowledge of distributed training frameworks, GPU cluster architecture, scheduler behavior, networking, and storage systems
- Experience operating in high-ambiguity, fast-moving environments
- Experience managing complex multi-team programs with competing priorities and hard deadlines
- Stakeholder management skills across technical individual contributors and senior leadership
- Ability to operate effectively during crises and prioritize under pressure
Responsibilities
- Own cross-functional programs spanning training infrastructure and cluster reliability across pre-training, mid-training, and post-training
- Drive end-to-end coordination of the training stack with engineering leads and external partners
- Triage incidents and escalations, coordinate responses, and drive resolution across teams
- Champion blameless post-mortems and turn incident learnings into system and process improvements
- Identify bottlenecks, define priorities, and align infrastructure investments with research velocity
- Maintain visibility into training health, cluster reliability, and infrastructure performance
- Create durable processes for cross-team handoffs, configuration management, and checkpoint workflows
- Translate technical complexity into status updates and decision frameworks for leadership
Benefits
- Stock options
- Comprehensive medical, dental, vision, and life insurance
- Annual wellness allowance
- Daily in-office lunch and dinner
- 22 weeks of paid parental leave
- Unlimited paid time off in the U.S.
- 30 vacation days in the U.K.
- Visa sponsorship support
- Regular off-sites, happy hours, and team celebrations
