Inference cost at scale with napkin math
kept by eddie
Napkin math shows serving a 32B LLM on an NVIDIA B200 GPU costs ~$9.36 per user per month when 300 users share the GPU with typical idle duty cycles.
One line. Many voicesSeek and you shall find
kept by eddie
Napkin math shows serving a 32B LLM on an NVIDIA B200 GPU costs ~$9.36 per user per month when 300 users share the GPU with typical idle duty cycles.