Multimodal Large Language Models (MLLMs) have achieved impressive performance across a variety of vision–language tasks. However, their internal working mechanisms remain largely underexplored. In this work, we introduce FG-TRACER, a framework designed to analyze the information flow between visual and textual modalities in MLLMs in free-form generation. Notably, our numerically stabilized computational method enables the first systematic analysis of multimodal information flow in underexplored domains such as image captioning and chain-of-thought (CoT) reasoning. We apply FG-TRACER to three state-of-the-art MLLMs—LLaVA 1.5, LLaMA 3.2-Vision, and Qwen 2.5-VL—across three vision–language benchmarks—TextVQA, COCO 2014, and ChartQA—and we conduct a word-level analysis of multimodal integration. Our findings uncover distinct patterns of multimodal fusion across models and tasks, demonstrating that fusion dynamics are both model- and task-dependent. Overall, FG-TRACER offers a robust methodology for probing the internal mechanisms of MLLMs in free-form settings, providing new insights into their multimodal reasoning strategies. Our source code is publicly available at https://github.com/AImageLab-zip/FG-TRACER

FG-TRACER: Tracing Information Flow in Multimodal Large Language Models in Free-Form Generation / Saporita, A., Pipoli, V., Bolelli, F., Baraldi, L., Acquaviva, A., Ficarra, E.. - (2026). (2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026 Tucson, Arizona 6 - 10 Mar, 2026).

FG-TRACER: Tracing Information Flow in Multimodal Large Language Models in Free-Form Generation

Saporita, Alessia;Pipoli, Vittorio;Bolelli, Federico;Baraldi, Lorenzo;Ficarra, Elisa
2026

Abstract

Multimodal Large Language Models (MLLMs) have achieved impressive performance across a variety of vision–language tasks. However, their internal working mechanisms remain largely underexplored. In this work, we introduce FG-TRACER, a framework designed to analyze the information flow between visual and textual modalities in MLLMs in free-form generation. Notably, our numerically stabilized computational method enables the first systematic analysis of multimodal information flow in underexplored domains such as image captioning and chain-of-thought (CoT) reasoning. We apply FG-TRACER to three state-of-the-art MLLMs—LLaVA 1.5, LLaMA 3.2-Vision, and Qwen 2.5-VL—across three vision–language benchmarks—TextVQA, COCO 2014, and ChartQA—and we conduct a word-level analysis of multimodal integration. Our findings uncover distinct patterns of multimodal fusion across models and tasks, demonstrating that fusion dynamics are both model- and task-dependent. Overall, FG-TRACER offers a robust methodology for probing the internal mechanisms of MLLMs in free-form settings, providing new insights into their multimodal reasoning strategies. Our source code is publicly available at https://github.com/AImageLab-zip/FG-TRACER
2026
2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
Tucson, Arizona
6 - 10 Mar, 2026
Saporita, Alessia; Pipoli, Vittorio; Bolelli, Federico; Baraldi, Lorenzo; Acquaviva, Andrea; Ficarra, Elisa
FG-TRACER: Tracing Information Flow in Multimodal Large Language Models in Free-Form Generation / Saporita, A., Pipoli, V., Bolelli, F., Baraldi, L., Acquaviva, A., Ficarra, E.. - (2026). (2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026 Tucson, Arizona 6 - 10 Mar, 2026).
File in questo prodotto:
File Dimensione Formato  
2026wacv.pdf

Open access

Tipologia: AAM - Versione dell'autore revisionata e accettata per la pubblicazione
Dimensione 638.37 kB
Formato Adobe PDF
638.37 kB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

Licenza Creative Commons
I metadati presenti in IRIS UNIMORE sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono rilasciati con licenza Attribuzione 4.0 Internazionale (CC BY 4.0), salvo diversa indicazione.
In caso di violazione di copyright, contattare Supporto Iris

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11380/1390068
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact