Voltar às notícias
InovaçãoInstruções AI Understanding

Benchmark finds vision-language models struggle to identify ambiguous targets through dialogue

A new benchmark reports that large vision-language models perform well below human baselines when they must identify visual targets through incomplete descriptions and interaction.

Por 5 min read
Primary-source image accompanying Benchmark finds vision-language models struggle to identify ambiguous targets through dialogue
A versão curta

A new benchmark reports that large vision-language models perform well below human baselines when they must identify visual targets through incomplete descriptions and interaction.

O que aconteceu

A new arXiv preprint introduces a controlled framework for testing interactive visual grounding in large vision-language models. The study varies how much information about a visual target is provided initially and how much must be gathered through dialogue. Across four human-grounded visual contexts and four interaction protocols, the authors report that current models perform significantly below task-level human baselines.

The paper addresses visual grounding, which it describes as the task of connecting a referring expression to a visual target. The authors argue that common evaluations are usually one-shot: a model receives an informative description and maps it to the intended item. Their framework instead treats reference as an interactive process in which the initial information may be incomplete or ambiguous and the target may be established through dialogue.

The evaluation varies two central conditions: how much target information is supplied at the start and how much must be acquired through interaction. The source says the framework covers four human-grounded visual contexts and four interaction protocols. It does not identify those contexts or protocols in the arXiv page text provided here, so the scope and realism of the test settings cannot be independently assessed from this source alone.

Across the tested settings, the authors report that current large vision-language models perform significantly below task-level human baselines. The abstract does not provide the underlying scores, the number or identity of the evaluated models, the size of the human comparison group, or the statistical procedures used to establish that gap. The result should therefore be read as the paper's reported finding, not as a quantified industry-wide measurement.

The reported pattern is not uniformly negative. Interaction can help when follow-up questions refine or repair an initial target description. However, performance is lowest when no initial description is supplied and the model must acquire target information by asking questions. The authors describe this as a difficulty with proactive, question-driven grounding. Follow-up studies reportedly found similar patterns across human- and AI-generated descriptions, different reasoning efforts, repeated interactions, description providers, and visual contexts, but the source does not give the details needed to compare those conditions.

Leia a fonte primária: arxiv.org

Por que isso importa

The findings identify a practical weakness in multimodal AI systems that must resolve ambiguous references rather than simply match a complete description to an image. The models sometimes improve when follow-up questions clarify or repair an initial description, but perform worst when they must proactively ask questions to discover what target a person means. The reported overconfidence also raises concerns for systems that communicate uncertainty to users.

Many visual AI tasks are easier when the system receives a complete, precise description of what to find. The paper's contribution is to focus on the less orderly situation in which a person refers to something vaguely, assumes shared context, or provides information in stages. In that setting, success requires more than visual matching: the system must recognize ambiguity, decide what information is missing, ask a useful question, and combine the answer with what it sees.

The reported gap between models and human task-level baselines matters because a system can appear capable in one-shot demonstrations while remaining unreliable in ordinary conversations. A user may expect a model to understand phrases such as an unclear reference to an object in a scene, but the benchmark suggests that performance deteriorates when the model must actively establish which object is intended. The source does not claim that every deployed system has this weakness in the same measure, but it identifies a failure mode that standard one-shot tests can miss.

The paper also reports poor calibration: models often express more confidence than their empirical accuracy warrants. This creates a separate operational problem from simply getting an answer wrong. If a system sounds certain while selecting the wrong target, a user may have less reason to verify the result. The abstract does not state how confidence was elicited, how calibration was measured, or whether any model or protocol performed substantially better than the others.

The broader implication, based on the authors' framing, is that interactive visual grounding should be evaluated as a combination of visual matching, information seeking, and synthesis. That can inform testing for multimodal assistants and other systems that use dialogue to interpret images. Still, the source supplies no deployment evidence, user-impact measurements, safety incidents, or evidence that the benchmark predicts performance in a particular commercial or public-sector application.

O que assistir a seguir

The paper is an arXiv preprint and the source does not provide model names, sample sizes, evaluation metrics, task-level scores, or details about the four visual contexts and interaction protocols. Further scrutiny should assess whether the framework generalizes across models, languages, visual environments, and real-world use cases, and whether better question selection or uncertainty calibration improves performance.

The first question is reproducibility. The source identifies the paper as arXiv:2608.23978, submitted on Aug. 25, 2026, and labels it a version-one preprint. It says the work is associated with EMNLP 2026 main subjects, but does not establish acceptance or peer-reviewed publication. Readers should examine the full paper for the task construction, annotation process, model list, prompts, scoring rules, and statistical analysis.

A second issue is external validity. The abstract says the patterns persist across varied description sources, reasoning efforts, repeated interactions, description providers, and visual contexts, but it does not specify those variations in the supplied material. Follow-up evaluation across unfamiliar images, cluttered environments, different languages, accessibility scenarios, and live user dialogue would help determine whether the reported weakness is narrow or general.

A third area to monitor is intervention design. The paper indicates that follow-up questions can refine or repair an initial description, while proactive questioning remains difficult. That distinction suggests a useful test for future work: whether models can identify uncertainty early and ask targeted questions instead of guessing. The source does not report a successful method for improving that behavior, so any claim that a particular prompting, training, or interface technique solves the problem would require separate evidence.

Finally, calibration deserves continued attention. A model that asks clarifying questions but remains overconfident may still create avoidable errors. Future results should report both grounding accuracy and confidence quality, including performance under repeated interaction and across different visual conditions. Until those details are available, the paper supports a clear caution about current LVLM evaluation, but not a precise ranking of systems or a conclusion about readiness for a specific deployment.

Guias e questionários relacionados

Modelos de IA explicadosChatGPT e LLMAgentes de IATeste o que você sabe – experimente um teste gratuito de IAProcure um termo de IA em nosso glossário
Achou isso útil?