Member of Technical Staff AI Infrastructure Engineer
PerplexityVisit Perplexity website
Perplexity is an AI-powered answer engine that provides real-time, cited answers and research capabilities.
San Francisco, United States
Funding history
About Perplexity
Private AI company founded in 2022. Its products combine conversational search and research with cited web sources; it also offers developer-facing Search, Agent, Router, and Embeddings APIs.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build, deploy, and optimize large-scale AI training and inference clusters. You will manage Kubernetes and Slurm environments, develop training and inference orchestration systems, improve performance and utilization, build observability solutions, and respond to outages.
Requirements
- Kubernetes administration, including custom resource definitions, operators, and cluster management
- Slurm workload management, job scheduling, resource allocation, and cluster optimization
- Distributed training systems
- Container orchestration and distributed systems architecture
- LLM architecture and training processes
- GPU cluster management and compute resource utilization
- Python
- C++
- PyTorch
- Networking, storage, and compute resource management for ML workloads
- API development and distributed systems management
- Debugging, monitoring, and observability tools
- Large-scale production Kubernetes deployment experience
- Slurm cluster administration and HPC workload management
- Experience supporting training jobs and high-availability inference services
- 3–5 years of relevant ML systems deployment experience
Responsibilities
- Design, deploy, and maintain scalable Kubernetes clusters for AI inference and training workloads
- Manage and optimize Slurm-based HPC environments for distributed LLM training
- Develop APIs and orchestration systems for training pipelines and inference services
- Implement resource scheduling and job management systems
- Benchmark performance, diagnose bottlenecks, and improve infrastructure
- Build monitoring, alerting, and observability solutions
- Respond to system outages and maintain uptime
- Optimize cluster utilization and implement autoscaling strategies
Benefits
- Equity
- Health insurance
- Dental insurance
- Vision insurance
- Retirement benefits
- Fitness account
- Commuter account
- Dependent care account
