Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

2026-07-10Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionArtificial Intelligence
AI summary

The authors created a new dataset called Complex Social Behavior (CSB) with 100 images showing complicated human interactions to better test how computer models understand scenes. They compared older vision language models (VLMs) with newer multimodal large language models (MLLMs) on both CSB and simpler images from MS-COCO. They found that newer MLLMs describe complex scenes as accurately as humans and fix most mistakes that older models made, except sometimes they focus on different parts of the image than people do. The authors also identified which types of mistakes, like object detection and hallucination, affect description accuracy the most. Overall, this work gives a clearer picture of how these models have improved in understanding images over the past decade.

Vision Language ModelsMultimodal Large Language ModelsComplex Social Behavior DatasetMS-COCOScene DescriptionObject DetectionHallucination ErrorSpatial Dependence ErrorVisual ReasoningModel Evaluation
Authors
Shravan Murlidaran, Miguel P. Eckstein
Abstract
Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated human descriptions as a benchmark, and have not focused on understanding the model's error types. Here, we introduce the Complex Social Behavior (CSB) dataset, containing 100 images depicting complex social interactions/behaviors. We analyze the progression of scene descriptions over a decade (2017-2025) of VLMs (four pre-Multimodal Large Language Models, MLLMs, and five MLLMs). We evaluate the accuracy of the models and 20 human descriptions relative to a gold standard on the CSB dataset and on a sample from MS-COCO. We analyzed five visual-cognitive error types: object detection, recognition, hallucination, scene understanding, and spatial dependence. The CSB dataset showed a more pronounced improvement than MS-COCO in scene description accuracy, with pre-MLLMs achieving much lower accuracy than the bottom-ranked human descriptions and MLLMs attaining accuracies similar to the top-ranked human descriptions. We show that MLLMs have eliminated the gap in scene description accuracy between simpler MS-COCO scenes and scenes depicting complex behaviors (CSB). MLLMs have almost eliminated all error types in our tested datasets, except for occasionally relying on different image regions for scene descriptions than humans do (spatial dependence error). We also show that detection, recognition, and hallucination errors have the highest impact on scene description accuracy. Together, our findings provide a more thorough evaluation of how visual language models have advanced over the last decade.