Machine Learning Researcher Multimodal LLMs
Bland AIVisit Bland AI website
Enterprise voice-AI platform for building, deploying, and monitoring AI phone agents.
San Francisco, United States
Funding history
Investors
About Bland AI
Bland provides in-house voice, language-model, speech-to-text, text-to-speech, telephony, testing, and observability infrastructure for inbound and outbound AI phone agents, primarily for regulated enterprise workflows.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will develop multimodal language models that combine speech, text, tools, and real-time reasoning. You will take models from research to production, integrate streaming audio and tool execution, design experiments, and improve user-facing conversational interactions.
Requirements
- Experience with LLMs, multimodal models, or speech-language systems
- Deep understanding of prompting, fine-tuning, and alignment
- Familiarity with neural audio codecs and multimodal LLM techniques
- Ability to design effective experiments
- Ability to translate modeling ideas into user-facing improvements
Responsibilities
- Develop a multimodal LLM stack combining speech, text, tools, and real-time reasoning
- Build conversational AI models from idea through production
- Integrate streaming audio, tool execution, and dynamic context
- Design experiments that answer research questions
- Translate modeling ideas into user-facing improvements
Benefits
- Equity
- Healthcare
- Dental insurance
- Vision insurance
- Office in Levi's Plaza
