Member of Technical Staff AI Inference Engineer
PerplexityVisit Perplexity website
Perplexity is an AI-powered answer engine that provides real-time, cited answers and research capabilities.
San Francisco, United States
Funding history
About Perplexity
Private AI company founded in 2022. Its products combine conversational search and research with cited web sources; it also offers developer-facing Search, Agent, Router, and Embeddings APIs.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will support transformer-based retrieval, text-generation, and multimodal models in inference infrastructure. You will port CUDA kernels to CuTe DSL, develop the Rust serving runtime, optimize performance across the serving path, and build observability and remediation systems while responding to production incidents.
Requirements
- Experience with GPU programming and performance work
- Knowledge of CUDA, Triton, CUTLASS, or similar technologies
- Knowledge of LLM architectures
- Experience operating production distributed systems under load
- Experience with Rust, Python, CUDA, and CuTe DSL
- 3+ years of professional software engineering experience in ML inference or high-performance systems
- Familiarity with PyTorch, JAX, or TensorFlow
- Understanding of GPU architectures and inference optimization techniques
Responsibilities
- Support retrieval, text-generation, and multimodal models in inference infrastructure
- Port CUDA kernels to CuTe DSL
- Develop a Rust-based inference server
- Profile and fix performance bottlenecks
- Build dashboards, alerts, and automated remediation
- Respond to and learn from production incidents
Benefits
- Equity
