VisCache Reports Faster Vision-Language Model Inference With Selective Visual Cache Pruning
A new arXiv paper proposes VisCache, a no-training framework that reduces visual KV-cache storage for vision-language models while retaining 19% to 28% of the cache. The authors report speedups of up to 2.35 times, but the results have not been independently validated here.