Senior Staff LLM Inference Engineer
d-MatrixVisit d-Matrix website
AI-infrastructure company producing memory-centric hardware, networking, and software for low-latency generative-AI inference in data centers.
Santa Clara, United States
Funding history
About d-Matrix
d-Matrix develops digital in-memory-compute technology and a full-stack inference platform comprising Corsair accelerators, JetStream networking accelerators, and Aviator software.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will prototype LLM inference use cases and build proof-of-concept systems. You will optimize kernels, quantization, sparsity, and batching; develop runtimes, serving frameworks, and evaluation tools; and contribute to distributed inference, hardware workload insights, and customer-facing demonstrations.
Requirements
- Bachelor’s degree and 10+ years of relevant engineering experience, or equivalent demonstrated experience
- Strong proficiency in Python and C/C++
- Experience optimizing LLM inference, including attention kernels, KV cache, batching, and quantization
- Contributor-level experience with an inference framework such as vLLM, SGLang, TensorRT-LLM, or ONNX Runtime
- Familiarity with CUDA or Triton GPU kernel programming and performance profiling tools
Responsibilities
- Identify and prototype LLM inference use cases
- Build proof-of-concept systems
- Develop and tune kernels and operator-level optimizations
- Drive quantization, sparsity, and batching strategies
- Build and maintain inference runtimes, serving frameworks, and evaluation tooling
- Contribute to distributed inference systems
- Provide hardware architects with inference workload insights
- Translate proof-of-concepts into customer-facing demonstrations
- Contribute to technical publications, whitepapers, and open-source projects
Benefits
- Medical insurance
- Dental insurance
- Vision insurance
- 401k
- Equity
- Bonus
