TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding
TVI-CoT: 面向多模态理解的文本-视觉交错思维链推理
机构 * University of Science and Technology of China(中国科学技术大学)
专题命中 其他推理 :CoT(title,title_cn);reasoning(title,abstract);chain-of-thought(title,abstract)
AI总结 提出TVI-CoT框架,通过可学习控制令牌实现文本推理与视觉特征访问的动态交错,解决多模态LLM在推理过程中无法访问视觉特征的问题,在多个基准上取得最优结果。
Comments ICML2026