AI Inference Engineer
Triune Infomatics Inc. is a privately held IT staffing and consulting company providing IT consulting, staffing, disability staffing, executive search, and AI-related consulting and talent services.
About Triune Infomatics Inc.
Founded in 2005, Triune Infomatics Inc. is a Pleasanton, California-based technology staffing and consulting company. Its current website offers client contact and job-application interactions, including active contract openings; its stated core business is staffing and consulting rather than a dedicated robotics product or service.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and optimize model-serving stacks and inference microservices, develop GPU kernels, tune LLM inference and KV-cache strategies, operate distributed multi-node GPU serving systems, build serving-platform components and OpenAI-compatible endpoints, and improve performance, reliability, observability, and model support.
Requirements
- Experience with production model-serving frameworks
- Proficiency in C++, Python, and Rust
- Experience writing GPU kernels using CUDA or ROCm
- Understanding of LLM inference internals, attention mechanisms, KV-cache management, continuous batching, and quantization
- Experience with distributed multi-node, multi-GPU serving environments
- Experience deploying and managing services on Kubernetes, OpenShift, or similar platforms
- Experience with performance profiling, benchmarking, and debugging latency or throughput issues
Responsibilities
- Build, operate, and optimize production model-serving stacks
- Develop and maintain high-throughput model-inference microservices
- Write and optimize custom GPU kernels
- Optimize LLM prefill, decoding, attention, and continuous batching
- Implement and tune quantization, speculative decoding, tensor parallelism, pipeline parallelism, and MoE serving
- Design and implement KV-cache optimization strategies
- Build and operate fault-tolerant serving systems on orchestration platforms
- Implement distributed computing across multi-node, multi-GPU clusters
- Contribute to distributed serving architecture components
- Build and maintain OpenAI-compatible endpoints
- Conduct profiling and benchmarking to resolve latency and throughput regressions
- Build telemetry-driven observability platforms
- Support a broad range of production model classes
