机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
ByteDance Inc.(字节跳动公司)
;
Zhejiang University of Technology(浙江工业大学)
;
Hong Kong Baptist University(香港浸会大学)
The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning
多模态焦点的潮起潮落:为基于视觉的语言模型推理调度视觉中继窗口
Wencheng Ye, Yi Bin, Yujuan Ding, Hongye Fang, Zheng Wang, Xing Xu, Jingkuan Song, Yun Zhang, Sirui Da, Heng Tao Shen
机构
*
School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)
;
School of Fashion and Textiles, The Hong Kong Polytechnic University(香港理工大学纺织及制衣学院)
;
School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)
Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning
Visual-Seeker:通过主动视觉推理实现视觉原生多模态智能搜索
Zhengbo Zhang, Changtao Miao, Jinbo Su, Zhaowen Zhou, Chunxia Zhang, Xukai Wang, Ruiqi Liu, Kaiyuan Zheng, Jiansheng Cai, Bo Zhang, Zhe Li, Shiming Xiang, Ying Yan
机构
*
School of Artificial Intelligence UCAS(中国科学院大学人工智能学院)
;
Institute of Automation CAS(中国科学院自动化研究所)
;
Ant Digital Technologies Ant Group(蚂蚁数字科技蚂蚁集团)
;
RUC(中国人民大学)
;
BIT(北京理工大学)
专题命中
视觉推理
:visual reasoning(title,abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.AI
Position: The Systemic Lack of Agency in Visual Reasoning
立场:视觉推理中系统性的能动性缺失
Yizhao Huang, Haoyang Chen, Shiqin Wang, Pohsun Huang, Jiayuan Li, Haoyuan Du, Yandong Shi, Zheng Wang, Zhixiang Wang
机构
*
National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science, Wuhan University(武汉大学计算机学院人工智能研究所多媒体软件工程研究中心)
;
Hubei Key Laboratory of Multimedia(湖北省多媒体与网络通信工程重点实验室)
;
Zhongguancun Academy, Beijing, China. 100094(中关村学院,北京,中国)
;
School of Automation, Beijing Institute of Technology(北京理工大学自动化学院)
;
Shanda AI Research Tokyo(上海大势人工智能研究东京)
Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction
在预测之前想象:用于视频事件预测的交错潜在视觉推理
Tianxiang Jiang, Linquan Wu, Sheng Xia, Songze Li, Ziang Yan, Haoyu Yang, Yu Qiao, Yi Wang
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
City University of Hong Kong(香港城市大学)
;
Nanjing University(南京大学)
;
Fudan University(复旦大学)
;
Zhejiang University(浙江大学)
;
University of Electronic Science and Technology of China(电子科技大学)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
VOLD:通过在线蒸馏将LLM推理能力转移到视觉语言模型
Walid Bousselham, Hilde Kuehne, Cordelia Schmid
机构
*
Tuebingen AI Center(图宾根人工智能中心)
;
University of Tuebingen(图宾根大学)
;
MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)
;
Inria, École Normale Supérieure, CNRS, PSL Research University(法国国家科学研究院、巴黎-萨克勒大学、École Normale Supérieure、PSL研究大学)
机构
*
Tandon School of Engineering, New York University(纽约大学工程学院)
;
Courant Institute of Mathematical Sciences, New York University(纽约大学数学科学学院)
;
Brookhaven National Laboratory(布鲁克海文国家实验室)
Beyond Text and Tables: Vision-Language Model Integration in ComProScanner for Extracting Materials Data from Scientific Figures with High Accuracy
超越文本与表格:ComProScanner中视觉-语言模型集成实现从科学图表中高精度提取材料数据
Aritra Roy, Enrico Grisan, Chiara Gattinoni, John Buckeridge
机构
*
Energy, Materials and Environment Research Centre, London South Bank University, London SE1 0AA, UK(能源、材料与环境研究中心,伦敦南银行大学)
;
School of Engineering and Design, London South Bank University, London SE1 0AA, UK(工程与设计学院,伦敦南银行大学)
;
Bioscience and Bioengineering Research Centre, London South Bank University, London SE1 0AA, UK(生物科学与生物工程研究中心,伦敦南银行大学)
;
Department of Physics, Kings College London, London WC2R 2LS, UK(物理系,伦敦国王学院)