Reading up on NVIDIA B200
1 deep · digging since jun 15
- Inference cost at scale with napkin math
Napkin math shows serving a 32B LLM on an NVIDIA B200 GPU costs ~$9.36 per user per month when 300 users share the GPU with typical idle duty cycles.
One topic. Every takeSeek and you shall find
1 deep · digging since jun 15
Napkin math shows serving a 32B LLM on an NVIDIA B200 GPU costs ~$9.36 per user per month when 300 users share the GPU with typical idle duty cycles.