Software Engineering Intern

AI inference cloud offering OpenAI-compatible APIs for machine-learning models, private deployments, and GPU infrastructure.

Palo Alto, United States
About DeepInfra

DeepInfra provides hosted AI inference for language, vision, embedding, image, video, and speech models, alongside private model deployments and dedicated GPU clusters.

View jobs by DeepInfra

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will work closely with experienced engineers to design, develop, test, and deploy inference solutions for AI models. You will implement and optimize models, monitor the live service, build features, fix bugs, review code, and participate in daily stand-ups and design discussions.

Requirements

  • Currently pursuing a Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related field
  • Fundamental knowledge of computer science, including data structures, algorithms, and software design patterns
  • Proficiency in Python and experience with AI/ML libraries and frameworks such as NumPy, pandas, SciPy, TensorFlow, and PyTorch
  • Familiarity with AI models, Transformers, and Diffusers
  • Experience with Git and agile development methodologies
  • Problem-solving ability, including debugging and code optimization
  • Communication and teamwork skills for cross-functional collaboration

Responsibilities

  • Collaborate with engineers to design, develop, and test inference solutions for AI models
  • Implement and optimize AI models using Python, C++, CUDA, and NCCL
  • Monitor and maintain the live service
  • Develop features, fix bugs, and conduct code reviews
  • Participate in daily stand-ups, code reviews, and design discussions
  • Stay up to date with AI and machine-learning trends and advancements
  • Try new approaches and ship software