One line. Many voicesSeek and you shall find

www.seangoedecke.com faviconTwo different tricks for fast LLM inference

kept by

Anthropic's fast mode uses low-batch-size inference on the full model, while OpenAI's uses a smaller distilled model on Cerebras chips for much higher speed.

read later

For all the tabs you promised to read.
Save to read. Read to clear.

Close tabs. Keep links.