Forward Deployed Engineer

AI inference cloud providing OpenAI-compatible APIs, private deployments, and GPU infrastructure for production AI workloads.

Palo Alto, United States
About DeepInfra

DeepInfra operates an AI inference cloud for running LLMs, vision, embeddings, image/video generation, speech, and other machine-learning models at scale, including private GPU deployments and GPU rental.

View jobs by DeepInfra

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own technical wins and proof-of-concept timelines from the first call through production. You will run benchmarks and evaluations, conduct provider bake-offs, tune model-to-hardware deployments, build cost models and migration plans, support security reviews, and create reusable benchmark reports and reference architectures.

Requirements

  • Customer-facing engineering experience with ownership of a technical outcome at an infrastructure or ML platform company
  • Python
  • Ability to communicate with engineering, finance, and procurement stakeholders

Responsibilities

  • Own technical wins and proof-of-concept timelines
  • Design and run benchmark harnesses and quality-parity evaluations
  • Run bake-offs against AI providers
  • Tune model-to-hardware deployments
  • Build cost-per-token models and migration plans
  • Handle enterprise security and compliance reviews
  • Drive post-signature usage reviews and expansion
  • Create reusable benchmark reports, reference architectures, and account-executive enablement materials