Reading up on AMD Instinct MI355X
1 deep · digging since sep 12
- Speculative Decoding in vLLM on AMD GPUs
Speculative decoding in vLLM enables vLLM to verify multiple draft tokens per target-model pass, boosting throughput on AMD GPUs depending on drafting method and acceptance rate.