One line. Many voicesSeek and you shall find

jdagostino.github.io faviconAI At Home Part 2: Multi GPU Drifting

kept by

The author boosted Deepseek V4 Flash inference from ~10 to ~20 tokens/sec on four e-waste AMD V620 GPUs by using layer parallelism and a speculative draft model.

read later

For all the tabs you promised to read.
Save to read. Read to clear.

Close tabs. Keep links.