GapSight teaches vision-language models when and where to look again
A new arXiv preprint introduces GapSight, a method that lets vision-language models selectively revisit image regions when a low-resolution global view loses important detail. The authors report higher average scores across six benchmarks, while noting that the results come from their own experiments and remain…