Job for Web3 Beginners

Software Engineer - Inference Platform

Baseten is an AI inference platform for deploying, optimizing, and scaling custom, open-source, and fine-tuned models in production.

San Francisco, United States
About Baseten

Baseten provides model runtimes, inference infrastructure, developer workflows, and deployment options including managed cloud, self-hosted, and hybrid environments.

View jobs by Baseten

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build and operate distributed runtime infrastructure for LLM inference. You will develop deployment orchestration, Model APIs, routing, autoscaling, observability, and platform controls while debugging production systems and owning projects from architecture through iteration.

Requirements

  • 3+ years building and operating distributed systems, backend infrastructure, or large-scale APIs
  • Experience owning low-latency backend services
  • Experience with rate limiting, authentication, quotas, metering, and migrations
  • Experience with profiling, tracing, capacity planning, and SLO management
  • Ability to debug across application, runtime, and infrastructure layers
  • Written communication and collaboration skills

Responsibilities

  • Build infrastructure and orchestration systems for distributed LLM inference
  • Design, build, and operate Model APIs
  • Implement API versioning, validation, metering, quotas, and authentication
  • Build observability and performance benchmarks
  • Debug and harden Kubernetes, networking, distributed runtime, and GPU systems
  • Own projects from architecture through deployment, monitoring, and iteration

Benefits

  • Equity
  • Medical, dental, and vision insurance for U.S. employees and dependents
  • Flexible PTO
  • Company-wide Winter Break
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Company-facilitated 401(k) for U.S. employees