Member of Technical Staff AI Infrastructure Engineer
PerplexityVisit Perplexity website
Perplexity is an AI-powered answer engine that provides real-time, cited answers and research capabilities.
San Francisco, United States
Funding history
About Perplexity
Private AI company founded in 2022. Its products combine conversational search and research with cited web sources; it also offers developer-facing Search, Agent, Router, and Embeddings APIs.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design, deploy, and maintain Kubernetes clusters and Slurm environments for AI training and inference. You will develop APIs and orchestration systems, manage scheduling and workloads, improve system performance, build observability, respond to outages, and optimize cluster utilization and autoscaling.
Requirements
- Kubernetes administration expertise
- Slurm workload-management experience
- Experience managing distributed training systems at scale
- Knowledge of container orchestration and distributed systems architecture
- Knowledge of LLM architecture and training processes
- Experience managing GPU clusters
- Python and C++ programming experience
- PyTorch distributed-training experience
- Knowledge of networking, storage, and compute resource management
- Experience with APIs, debugging, and observability
Responsibilities
- Design, deploy, and maintain Kubernetes clusters for AI workloads
- Manage and optimize Slurm HPC environments
- Develop APIs and orchestration systems for training and inference
- Implement resource scheduling and job management
- Benchmark performance and resolve infrastructure bottlenecks
- Build monitoring, alerting, and observability solutions
- Respond to system outages and maintain availability
- Optimize cluster utilization and autoscaling
