KV cache · Key-Value cache
Stored attention tensors reused during decode; its size grows with context and concurrency, dominating inference memory.
Also written as: KV-cache · KV-caches
Stored attention tensors reused during decode; its size grows with context and concurrency, dominating inference memory.
Also written as: KV-cache · KV-caches