Machine Learning Researcher Multimodal LLMs

Enterprise voice-AI platform for building, deploying, and monitoring AI phone agents.

San Francisco, United States
About Bland AI

Bland provides in-house voice, language-model, speech-to-text, text-to-speech, telephony, testing, and observability infrastructure for inbound and outbound AI phone agents, primarily for regulated enterprise workflows.

View jobs by Bland AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will develop multimodal language models that combine speech, text, tools, and real-time reasoning. You will take models from research to production, integrate streaming audio and tool execution, design experiments, and improve user-facing conversational interactions.

Requirements

  • Experience with LLMs, multimodal models, or speech-language systems
  • Deep understanding of prompting, fine-tuning, and alignment
  • Familiarity with neural audio codecs and multimodal LLM techniques
  • Ability to design effective experiments
  • Ability to translate modeling ideas into user-facing improvements

Responsibilities

  • Develop a multimodal LLM stack combining speech, text, tools, and real-time reasoning
  • Build conversational AI models from idea through production
  • Integrate streaming audio, tool execution, and dynamic context
  • Design experiments that answer research questions
  • Translate modeling ideas into user-facing improvements

Benefits

  • Equity
  • Healthcare
  • Dental insurance
  • Vision insurance
  • Office in Levi's Plaza
Machine Learning Researcher Multimodal LLMs at Bland AI | JobStash