Applied AI Research Engineer
Beam is an AI-native serverless cloud platform for GPU inference, task queues, and secure sandboxes.
Maintainer signals as of 9/23/2026
Funding history
About Beam
Smartshare, Inc. operates Beam, a developer platform for running AI and ML workloads—including inference endpoints, agents, task queues, and sandboxes—on CPUs and GPUs without managing underlying infrastructure.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead hands-on inference research to reduce token costs and latency for customer workloads. You will optimize inference systems, including speculative decoding, quantization, KV-cache usage, and memory management. You will work directly with customers, turn findings into platform improvements, and identify high-impact research opportunities.
Requirements
- Systems or research background in LLM inference
- Deep understanding of LLM serving from kernel to scheduler
- History of shipping products or research used in production-like scenarios
- Ability to collaborate closely with customers
- Knowledge of developer tools, cloud-native technologies, and open-source software
Responsibilities
- Optimize inference systems using speculative decoding, quantization, KV-cache, and memory management
- Work with customers to optimize production workloads
- Apply customer workload learnings to platform product improvements
- Identify high-upside research opportunities and guide the inference platform
Benefits
- Meaningful equity
- Health, dental, and vision benefits with 90% coverage for employees and 50% for dependents
- Opportunities to participate in cloud-native community events
- Fitness stipend
- Learning budget
