Senior Software Engineer - Model Performance
Inference is a distributed GPU network for running AI models efficiently across a global infrastructure.
Maintainer signals as of 9/2/2026
Funding history
About Inference
Inference is a distributed GPU network for running AI models. Users can connect their devices to contribute computing power, access APIs to perform model inference, and monitor workloads through a dashboard. The platform offers an open infrastructure that helps scale machine learning workloads without relying on centralized servers.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will make the inference stack faster and more efficient by implementing optimization techniques, experimenting with novel approaches, profiling GPU workloads, and bringing performant model architectures into production.
Requirements
- 2+ years of experience in ML systems, inference optimization, or GPU programming.
- Strong proficiency in Python and familiarity with C++.
- Hands-on experience with LLM inference frameworks such as vLLM, SGLang, TensorRT-LLM, or similar.
- Deep understanding of GPU architecture and experience profiling GPU workloads.
- Familiarity with quantization, speculative decoding, continuous batching, and KV cache management.
- Experience with PyTorch and understanding of model execution on hardware.
- Track record of measurably improving system performance.
- Nice-to-have experience with CUDA programming, non-LLM model serving, distributed inference, open-source inference frameworks, Docker, and Kubernetes.
Responsibilities
- Implement and productionize quantization, speculative decoding, KV cache optimization, continuous batching, and LoRA serving.
- Debug and improve vLLM, SGLang, TensorRT-LLM, and underlying libraries.
- Profile CUDA kernels and optimize GPU utilization across serving infrastructure.
- Add support for new model architectures and ensure they meet production performance standards.
- Experiment with novel inference techniques and productionize successful approaches.
- Build tooling and benchmarks to track inference performance across the fleet.
- Collaborate with applied ML engineers to ensure trained models can be served efficiently.
Benefits
- Equity
- Comprehensive benefits
Hiring Process
Applicants may send a resume and GitHub to amar@inference.net and/or apply through Ashby.
