Study Finds Vision-Language Models Hit a Model-Class Ceiling in Few-Shot Adaptation
A new arXiv study reports that tuning the blend between text and image prototypes is not the main limit on few-shot vision-language model accuracy. In experiments spanning 4,800 evaluation cells, validation-free linear probes outperformed even a blend selected with test-set information.