Speculative Decoding in vLLM on AMD GPUs
kept by night
Speculative decoding in vLLM enables vLLM to verify multiple draft tokens per target-model pass, boosting throughput on AMD GPUs depending on drafting method and acceptance rate.