Member of Technical Staff, LLM Evaluation Infra
InceptionVisit Inception website
AI research and product company building diffusion-based language models for production applications.
Palo Alto, United States
Funding history
About Inception
Inception develops and deploys the Mercury family of diffusion LLMs, which generate and refine output in parallel to target lower-latency, lower-cost production AI workloads.
Skills
About the Role
You will build automated evaluation pipelines for model training and deployment. You will develop benchmarks and metrics, analyze model outputs for failures and biases, and turn real-world use cases into meaningful evaluation criteria.
Requirements
- Machine learning evaluation
- Applied machine learning research
- Git
- Docker
- AWS
- GCP
- Azure
- Large language model fundamentals
- Python
- PyTorch
- Evaluation metrics
- Generative model benchmarks
- Communication
Responsibilities
- Build scalable automated evaluation pipelines for model training and deployment
- Design and maintain evaluation frameworks and benchmarks for language models
- Analyze model outputs to identify failure modes, biases, and performance gaps
- Translate real-world use cases into evaluation criteria
- Define metrics for model quality, safety, reliability, and regression detection
Benefits
- Equity
- Flexible vacation and paid time off
- Health insurance
- Dental insurance
- Vision insurance
- 401k match
- Catered meals
- Commuter subsidies
