Senior Software Engineer - Model Performance

Inference is a distributed GPU network for running AI models efficiently across a global infrastructure.

Seed0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/2/2026

Distributed

Funding history

About Inference

Inference is a distributed GPU network for running AI models. Users can connect their devices to contribute computing power, access APIs to perform model inference, and monitor workloads through a dashboard. The platform offers an open infrastructure that helps scale machine learning workloads without relying on centralized servers.

View jobs by Inference

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will make the inference stack faster and more efficient by implementing optimization techniques, experimenting with novel approaches, profiling GPU workloads, and bringing performant model architectures into production.

Requirements

  • 2+ years of experience in ML systems, inference optimization, or GPU programming.
  • Strong proficiency in Python and familiarity with C++.
  • Hands-on experience with LLM inference frameworks such as vLLM, SGLang, TensorRT-LLM, or similar.
  • Deep understanding of GPU architecture and experience profiling GPU workloads.
  • Familiarity with quantization, speculative decoding, continuous batching, and KV cache management.
  • Experience with PyTorch and understanding of model execution on hardware.
  • Track record of measurably improving system performance.
  • Nice-to-have experience with CUDA programming, non-LLM model serving, distributed inference, open-source inference frameworks, Docker, and Kubernetes.

Responsibilities

  • Implement and productionize quantization, speculative decoding, KV cache optimization, continuous batching, and LoRA serving.
  • Debug and improve vLLM, SGLang, TensorRT-LLM, and underlying libraries.
  • Profile CUDA kernels and optimize GPU utilization across serving infrastructure.
  • Add support for new model architectures and ensure they meet production performance standards.
  • Experiment with novel inference techniques and productionize successful approaches.
  • Build tooling and benchmarks to track inference performance across the fleet.
  • Collaborate with applied ML engineers to ensure trained models can be served efficiently.

Benefits

  • Equity
  • Comprehensive benefits

Hiring Process

Applicants may send a resume and GitHub to amar@inference.net and/or apply through Ashby.