Research Engineer Mid Training
CognitionVisit Cognition website
Cognition is an applied AI company that operates Devin, an autonomous software-engineering agent.
San Francisco, United States
Funding history
About Cognition
Cognition builds AI agents and models for software engineering. Its flagship product, Devin, can plan, write, test, and ship code in a customer’s existing codebase and tools.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own late-stage training decisions across data quality, data mixing, annealing schedules, context extension, capability improvement, and synthetic data. You will build evaluations, study scaling behavior, run training experiments, and turn research findings into measurable improvements for deployed agent capabilities.
Requirements
- Familiarity with the end-to-end LLM training pipeline
- Experience with continual pre-training, annealing, or late-stage data mixing
- Experience developing or evaluating synthetic data pipelines
- Proficiency in Python and PyTorch
- Knowledge of optimization, statistics, and machine learning theory
- Track record of original contributions through publications, open-source work, or internal results
Responsibilities
- Design and iterate on high-quality data mixtures for late-stage and annealing training runs
- Develop data sourcing, filtering, and weighting methods
- Drive coding, mathematics, and reasoning capability improvements through training interventions
- Develop and evaluate synthetic data pipelines
- Research and optimize learning-rate schedules, warmup strategies, and compute allocation
- Research and implement context-length extension methods
- Build evaluations that distinguish capability improvements from benchmark overfitting
- Measure how interventions scale with compute and data
