ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
Abstract
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual entrainment in VLMs requires more than porting an existing text-only benchmark to the multimodal setting: it requires a taxonomically structured, dual-modality instrument whose conditions are constructed around the item at hand (the depicted image in the textual stream, the textual query in the visual stream). We argue that the move to VLMs is substantive rather than incremental. It makes entrainment a dual phenomenon, drivable independently by textual and by visual context, and it opens a veracity distinction (context that is false of the depicted scene yet possible in the world) that has no counterpart in the unimodal, world-knowledge-only formulation of prior work. To make this position concrete and actionable, we introduce ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes, i.e., the association of context with the item and its relationship to truth, and split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions). We do not claim to measure entrainment in any particular model; we provide the instrument, the taxonomy that motivates it, and the evaluation protocols it enables, so that the community can investigate the phenomenon rigorously. We will release the dataset and its documentation publicly.
Community
Contextual entrainment — a model's tendency to be pulled by auxiliary context regardless of whether that context is relevant, true, or meaningful — has been studied in unimodal language models but remains largely unexamined in vision-language models. We argue this multimodal setting is substantive rather than incremental: entrainment becomes a dual phenomenon drivable by both textual and visual context, and it opens a scene-relative veracity distinction with no counterpart in text-only work. We introduce ENTRAP-VL, a manually curated dataset of 1,500 items across eight categories, organized on two axes (association and veracity) and split into a textual-entrainment stream (eight conditions) and a visual-entrainment stream (three). We release the instrument, not measurements, so the community can investigate the phenomenon rigorously.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Do Vision-Language Models See or Guess? Measuring and Reducing Textual-Prior Reliance with a Phrasing-Controlled Benchmark (2026)
- Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering (2026)
- Sentence-Level Contextual Entrainment in Large Language Models (2026)
- How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA (2026)
- Diagnosing Visual Ignorance in Vision-Language Models (2026)
- Semantic Flip: Synthetic OOD Generation for Robust Refusal in Embodied Question Answering and Spatial Localization (2026)
- Pathways of Visual Information Flow in Vision-Language Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.20092 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper
