Learning to Route Speculative Decoding Under Workload and Traffic Regimes

Workload-aware routing for speculative decoding in LLM serving, built around vLLM.

Placeholder page. Content is being written. Preview the section structure

Overview

Problem

Why it matters

Approach

System / architecture

Experiments

Results

What I learned

Links