Member of Technical Staff Applied Research
LlamaIndex provides AI document-processing and knowledge-agent infrastructure, centered on LlamaParse for parsing, extraction, indexing, and retrieval over unstructured enterprise documents.
About LlamaIndex
A San Francisco AI company offering a developer-first platform and commercial APIs for document workflows and knowledge agents. Its active products include LlamaParse, agentic OCR and structured extraction software, and the open-source local document parser LiteParse.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will develop and train vision-language models for document processing, build data pipelines for curation and synthetic data, and evaluate or fine-tune models against performance targets. You will improve production accuracy, latency, and cost, create benchmarks, and move successful prototypes into production.
Requirements
- Machine learning engineering
- Applied research
- Research engineering
- Model benchmarking
- Model training
- Python
- PyTorch
- Computer vision
- Vision-language model
- NLP
- Document AI
- OCR
- Information extraction
- Agentic AI
- Experimentation
- Production code
- Technical writing
- Synthetic data generation
- Post-training
- Fine-tuning
- vLLM
- Pydantic
- uv
- ruff
- mypy
- Claude Code
- Cursor
Responsibilities
- Develop and train vision-language models for document processing and understanding
- Build data pipelines for curation, synthetic data generation, labeling, and benchmark creation
- Evaluate base models and perform post-training or fine-tuning
- Improve model accuracy, latency, and cost-effectiveness
- Design and maintain benchmarks for extraction, layout, OCR, reasoning, and system reliability
- Work with real-world documents including PDFs, scanned documents, tables, charts, and forms
- Move successful research prototypes into production
- Translate customer requirements into benchmarks, experiments, and model improvements
- Research vision-language models, document AI, post-training, synthetic data, and agentic systems
- Use modern AI coding workflows and tools
