Software Engineer Infrastructure
Tavus is an AI research lab and operating company focused on Human Computing, offering real-time human-like AI PALs, multimodal models, developer APIs, and a no-code PAL Maker.
About Tavus
Tavus builds foundational AI models for perception, conversational understanding, voice, and human rendering. Its current offerings include the Conversational Video Interface developer API, PAL Maker for no-code AI-human creation and deployment, and managed enterprise PAL solutions for use cases such as recruiting, healthcare, sales, education, and onboarding.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own and improve the systems behind a real-time conversational product. You will operate GPU inference infrastructure across providers and regions, expand clusters on EKS, improve routing and scheduling, raise uptime, support security and SOC2 work, and improve deployment pipelines. You will also investigate and fix issues across infrastructure, backend services, and product code.
Requirements
- Hands-on experience deploying and optimizing GPU inference workloads
- Experience building reliable systems on GPU cloud providers
- Deep Kubernetes and EKS knowledge, including routing and scheduling
- Experience designing workload placement across a fleet
- Experience writing services for workload placement
- Deep AWS experience
- Senior-level ownership experience, including setting technical direction and delivering ambiguous work
- Ability to explain complex ideas clearly to engineers and non-engineers
Responsibilities
- Own GPU inference deployments serving live conversations across multiple providers and regions
- Tune GPU generations and reduce cold-start and model load times
- Expand GPU capacity by onboarding providers and regions
- Stand up clusters on EKS
- Build routing, scheduling, and throughput for fast weight loading
- Improve infrastructure uptime
- Support security and SOC2 work
- Identify, fix, or escalate infrastructure and software problems
- Improve deployment pipelines
