Reading up on Batch API
1 deep · digging since feb 20
- Prompt Caching 201
Prompt caching reduces latency by up to 80% and costs by up to 90% by reusing computed key/value tensors for repeated prompt prefixes in OpenAI's API.
One topic. Every takeSeek and you shall find
1 deep · digging since feb 20
Prompt caching reduces latency by up to 80% and costs by up to 90% by reusing computed key/value tensors for repeated prompt prefixes in OpenAI's API.