Software Engineer
DeepInfraVisit DeepInfra website
AI inference cloud providing OpenAI-compatible APIs, private deployments, and GPU infrastructure for production AI workloads.
Palo Alto, United States
Funding history
About DeepInfra
DeepInfra operates an AI inference cloud for running LLMs, vision, embeddings, image/video generation, speech, and other machine-learning models at scale, including private GPU deployments and GPU rental.
Skills
About the Role
You will design, develop, and test inference solutions for AI models. You will implement, optimize, and evaluate models; operate production model-serving systems; monitor and debug services; improve performance; build features; and contribute to code reviews and system design.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field
- 3+ years of relevant experience
- Data structures, algorithms, and software design fundamentals
- Proficiency in Python
- Experience with AI and ML frameworks such as PyTorch or TensorFlow
- Experience building, shipping, and maintaining software systems
- Familiarity with AI models, Transformers, and Diffusers
- Experience with Git and collaborative development workflows
- Ability to debug, optimize, and improve systems
- Communication skills and ability to work independently
Responsibilities
- Design, develop, and test inference solutions for AI models
- Implement, optimize, and evaluate AI models using Python, C++, CUDA, and NCCL
- Own and operate production model-serving systems
- Monitor and debug production systems
- Build features and improve system performance
- Participate in code reviews and technical discussions
- Explore AI and ML techniques to improve model performance and efficiency
- Take ideas from concept to production
