Reading up on ROCm
3 deep · digging since aug 20
- Speculative Decoding in vLLM on AMD GPUs
Speculative decoding in vLLM enables vLLM to verify multiple draft tokens per target-model pass, boosting throughput on AMD GPUs depending on drafting method and acceptance rate.
- AI At Home Part 2: Multi GPU Drifting
The author boosted Deepseek V4 Flash inference from ~10 to ~20 tokens/sec on four e-waste AMD V620 GPUs by using layer parallelism and a speculative draft model.