Engineering Lead Inference Optimization
Venice AI provides private, uncensored AI services for text, image, video, audio, code, search, and agent development.
Maintainer signals as of 8/23/2026
Funding history
About Venice AI
Venice AI is an AI platform focused on privacy and uncensored access to leading third-party and open-source models. Its products include conversational AI, image and video generation, audio and music tools, code assistance, web search, and an OpenAI-compatible API for agents and multimodal applications. Privacy features include anonymization, zero data retention for self-hosted models, trusted execution environments, and end-to-end encryption. The platform offers free access alongside Pro, Pro+, and Max subscription tiers for creators, developers, and enterprises.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You own inference performance strategy, recruit engineers, optimize GPU infrastructure, and improve latency throughput and cost per token. You build reproducible benchmarks across inference engines, optimize load balancing, evaluate quantization and compilation techniques, develop performance kernels, and assess emerging inference hardware.
Requirements
- 8+ years in performance optimization or HPC
- 5+ years leading engineering teams
- Deep GPU architecture and parallel programming knowledge
- Proficiency in Python Rust or Go
- Hands-on production experience with an LLM inference engine such as vLLM or SGLang
- Experience with continuous batching PagedAttention KV cache management speculative decoding quantization CUDA graphs and torch.compile
- Experience with distributed inference strategies in multi-GPU and multi-node environments
- Fluency with GPU profiling tools
Responsibilities
- Own technical strategy for inference performance
- Recruit and lead the inference optimization function
- Optimize GPU infrastructure across multiple architectures
- Improve latency throughput and cost per token
- Build benchmarking harnesses across inference engines
- Optimize multivariate inference load balancing
- Evaluate kernels attention variants quantization schemes and compilation improvements
- Evaluate emerging inference hardware
