GeoPID: Decomposing and Steering Visual Information in Vision-Language Models
Abstract
While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context. In this work, we propose GEOPID, a training-free framework that analyzes multimodal information within VLMs from a geometric perspective. GEOPID decomposes information into Redundant, ModalityUnique, and Synergistic components through the geometric relationships between visual and textual representation subspaces. Through an extensive analysis across 22 VLMs and 14 benchmarks, we confirm that correct predictions exhibit stronger vision-unique components when questions strongly require visual grounding. Building on this geometric analysis, we introduce a targeted intervention technique that selectively amplifies visual representations along the visionunique subspace during inference. As a result, visual grounding capabilities were enhanced without any additional model parameter updates, achieving an average relative accuracy gain of 7.63%.
BibTeX
@article{kim2026geopid,
title={GeoPID: Decomposing and Steering Visual Information in Vision-Language Models},
author={Kim, Seulgi and Zhang, Zhixiong and Zhang, Xinwei and Ling, Jie and Shaw, Ronn},
journal={arXiv preprint arXiv:2610.08401},
year={2026}
}
Zhixiong Zhang