Learning to Route Speculative Decoding Under Workload and Traffic Regimes
Workload-aware routing for speculative decoding in LLM serving, built around vLLM.
Placeholder page. Content is being written. Preview the section structure
Workload-aware routing for speculative decoding in LLM serving, built around vLLM.
Placeholder page. Content is being written. Preview the section structure