Member of Technical Staff, Training Infra
InceptionVisit Inception website
AI research and product company building diffusion-based language models for production applications.
Palo Alto, United States
Funding history
About Inception
Inception develops and deploys the Mercury family of diffusion LLMs, which generate and refine output in parallel to target lower-latency, lower-cost production AI workloads.
Skills
About the Role
You will design and optimize distributed training systems across large GPU fleets. You will improve training throughput and efficiency and develop reusable frameworks that strengthen reliability, reproducibility, and scalability for new model architectures.
Requirements
- PyTorch
- TensorFlow
- Engineering
- Python
- C++
- Rust
- Go
- Docker
- Kubernetes
- CI/CD
Responsibilities
- Design and optimize distributed training systems across GPUs and nodes
- Develop high-performance optimizations for training throughput and efficiency
- Develop reusable training frameworks and libraries
- Improve training reproducibility reliability and scalability
Benefits
- Equity
- Flexible vacation and paid time off
- Health insurance
- Dental insurance
- Vision insurance
- 401k match
- Catered meals
- Commuter subsidies
