V‑ABS:面向动态视觉推理的行动‑观察者驱动束搜索

V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning

Zhiwei Ning, Xuanang Gao, Jie Yang, Wei Liu, et al.

ICML 2026

Abstract

Multimodal large language models (MLLMs)have achieved remarkable success in general per-ception, yet complex multi-step visual reasoningremains a persistent challenge. Although recentagentic approaches incorporate tool use, they of-ten neglect critical execution feedback. Conse-quently, they suffer from the imagination-action-observer (IAO) bias, a misalignment betweenprior imagination and observer feedback that un-dermines reasoning stability and optimality. Tobridge this gap, we introduce V-ABS, an action-observer driven beam search framework that en-ables deliberate reasoning through thinker-actor-observer iterations. We also propose an entropy-based adaptive weighting algorithm to mitigatethe IAO bias by dynamically balancing the con-fidence scores between the policy priors and theobservational feedback. Moreover, we constructa large-scale supervised fine-tuning (SFT) datasetcomprising over 80k samples to guide the modelto assign higher prior confidence to correct ac-tion paths. Extensive experiments across eightdiverse benchmarks show that V-ABS achievesstate-of-the-art performance, delivering an aver-age improvement of 19.7% on the Qwen3-VL-8Bbaseline and consistent gains across both open-source and proprietary models. Code is availableat https://github.com/pami-zwning/V-ABS.

Figure 1. Analysis of the IAO bias. The top panel reveals a significant discrepancy between the model’s prior scores Fpri and visual utility outcomes Fobs. The bottom panel demonstrates that V-ABS exhibits significant variance in confidence scores across different actions, which yields a substantial accuracy gain on the V* benchmark.

 

Figure 2. Overview of the V-ABS framework. (a) Action-observer driven algorithm: at each reasoning step, the thinker generates prior scores Fpri for each candidate action, the actor executes these actions via tool functions Ttool to update the visual state, and the observer evaluates the resulting states to obtain feedback scores Fobs, assigned by a heuristic score Fheur. The adaptive weighting mechanism dynamically balances these scores to ensure the optimal trajectory. (b) Visualization: an example of multi-step visual search where V-ABS actively updates the visual context through cropping operations. (c) Quantitative performance: V-ABS significantly outperforms the baseline models across diverse tasks.

 

https://doi.org/10.48550/arXiv.2605.10172

Copyright © 2025上海交通大学医疗机器人研究院 版权所有 沪交ICP备20190057   流量统计