Software Engineer Inference
Luma AI develops multimodal creative AI agents, video and image generation models, and an API for generative media workflows.
Funding history
About Luma AI
Luma AI is an active AI company whose Luma Agents product plans, generates, iterates, and refines creative work across video, image, audio, and text. Its first-party materials describe proprietary Ray video and Uni image models, as well as developer API access.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will integrate model architectures into the inference engine and optimize deployments across clusters and hardware providers. You will build tools for inference jobs, automate and maintain reliable services, develop scheduling systems, and maintain CI/CD for model checkpoints and SDKs.
Requirements
- Python and system-architecture skills
- Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar tools
- Experience with queues, scheduling, traffic control, and fleet management at scale
- Experience with Linux, Docker, Kubernetes, orchestration, deployment, and scheduling
- Familiarity with Redis and S3-compatible storage
Responsibilities
- Integrate new model architectures into the inference engine
- Optimize model efficiency and deployments
- Build tooling to measure, profile, and track inference jobs and workflows
- Automate, test, and maintain inference services
- Manage and optimize inference workloads across clusters and hardware providers
- Build scheduling systems for GPU resources while meeting SLOs
- Maintain CI/CD for model checkpoints and SDKs
