GLANCE speeds up vision-language model decoding by drafting token blocks in one pass
A new arXiv paper introduces GLANCE, a speculative-decoding method that uses a vision-language model’s fused state to draft whole token blocks at once. The authors report lossless greedy outputs and up to 2.93× faster decoding in tested workloads.