PuzzleKV proposes page-wise compression to reduce LLM KV-cache storage
A new arXiv paper proposes PuzzleKV, a training- and calibration-free method that independently compresses completed pages of large language model KV caches. It reports more than 96% of Full KV performance at about 60% of original storage, and more than 93% at 18.7% of original storage with quantization.