Staff Software Engineer Inference API

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build and evolve production ML inference APIs for a heterogeneous serving system. You will integrate models and serving runtimes, coordinate GPU prefill with Cerebras decode, improve performance and reliability, develop validation and testing infrastructure, and create developer tools, SDKs, and documentation.

Requirements

  • 5+ years of software engineering experience with substantial individual-contributor ownership of production software or distributed systems
  • Strong Python and Go programming skills
  • Experience developing performance-sensitive or highly concurrent services in C++, Rust, or a similar systems language
  • Experience building stable APIs with validation, error handling, observability, compatibility, and versioning
  • Experience integrating software across service, framework, runtime, and infrastructure boundaries
  • Experience with OpenAI-compatible, gRPC, REST, or streaming inference APIs
  • Experience with Linux, containers, Kubernetes or comparable orchestration systems, CI/CD, and latency-sensitive production services
  • Ability to diagnose correctness, reliability, and performance issues in distributed serving systems
  • Bachelor's degree in a related discipline or equivalent practical experience

Responsibilities

  • Design, implement, and maintain production ML inference APIs
  • Create consistent request and response semantics across heterogeneous inference backends
  • Integrate foundation models, tokenizers, prompt formats, sampling methods, multimodal inputs, and model-specific features
  • Maintain API compatibility and establish versioning, deprecation, validation, and backward-compatibility practices
  • Extend and integrate inference services with vLLM, PyTorch, Hugging Face libraries, AMD ROCm, and Cerebras runtime components
  • Build control and data paths for GPU prefill and Cerebras decode coordination
  • Optimize streaming, latency, throughput, batching, serialization, tokenization, scheduling, and component communication
  • Build validation systems for functional and numerical correctness
  • Develop logging, tracing, metrics, dashboards, health checks, and diagnostic tooling
  • Create conformance tests, workload replay tools, validation suites, benchmarks, integration tests, and release gates
  • Build configuration, SDKs, documentation, examples, debugging tools, and self-service workflows
  • Translate model and customer requirements into scalable serving capabilities