Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning
Visual-Seeker:通过主动视觉推理实现视觉原生多模态智能搜索
Zhengbo Zhang, Changtao Miao, Jinbo Su, Zhaowen Zhou, Chunxia Zhang, Xukai Wang, Ruiqi Liu, Kaiyuan Zheng, Jiansheng Cai, Bo Zhang, Zhe Li, Shiming Xiang, Ying Yan
机构
*
School of Artificial Intelligence UCAS(中国科学院大学人工智能学院)
;
Institute of Automation CAS(中国科学院自动化研究所)
;
Ant Digital Technologies Ant Group(蚂蚁数字科技蚂蚁集团)
;
RUC(中国人民大学)
;
BIT(北京理工大学)
专题命中
视觉推理
:visual reasoning(title,abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.AI
Position: The Systemic Lack of Agency in Visual Reasoning
立场:视觉推理中系统性的能动性缺失
Yizhao Huang, Haoyang Chen, Shiqin Wang, Pohsun Huang, Jiayuan Li, Haoyuan Du, Yandong Shi, Zheng Wang, Zhixiang Wang
机构
*
National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science, Wuhan University(武汉大学计算机学院人工智能研究所多媒体软件工程研究中心)
;
Hubei Key Laboratory of Multimedia(湖北省多媒体与网络通信工程重点实验室)
;
Zhongguancun Academy, Beijing, China. 100094(中关村学院,北京,中国)
;
School of Automation, Beijing Institute of Technology(北京理工大学自动化学院)
;
Shanda AI Research Tokyo(上海大势人工智能研究东京)
Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction
在预测之前想象:用于视频事件预测的交错潜在视觉推理
Tianxiang Jiang, Linquan Wu, Sheng Xia, Songze Li, Ziang Yan, Haoyu Yang, Yu Qiao, Yi Wang
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
City University of Hong Kong(香港城市大学)
;
Nanjing University(南京大学)
;
Fudan University(复旦大学)
;
Zhejiang University(浙江大学)
;
University of Electronic Science and Technology of China(电子科技大学)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
VOLD:通过在线蒸馏将LLM推理能力转移到视觉语言模型
Walid Bousselham, Hilde Kuehne, Cordelia Schmid
机构
*
Tuebingen AI Center(图宾根人工智能中心)
;
University of Tuebingen(图宾根大学)
;
MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)
;
Inria, École Normale Supérieure, CNRS, PSL Research University(法国国家科学研究院、巴黎-萨克勒大学、École Normale Supérieure、PSL研究大学)
机构
*
Tandon School of Engineering, New York University(纽约大学工程学院)
;
Courant Institute of Mathematical Sciences, New York University(纽约大学数学科学学院)
;
Brookhaven National Laboratory(布鲁克海文国家实验室)
Beyond Text and Tables: Vision-Language Model Integration in ComProScanner for Extracting Materials Data from Scientific Figures with High Accuracy
超越文本与表格:ComProScanner中视觉-语言模型集成实现从科学图表中高精度提取材料数据
Aritra Roy, Enrico Grisan, Chiara Gattinoni, John Buckeridge
机构
*
Energy, Materials and Environment Research Centre, London South Bank University, London SE1 0AA, UK(能源、材料与环境研究中心,伦敦南银行大学)
;
School of Engineering and Design, London South Bank University, London SE1 0AA, UK(工程与设计学院,伦敦南银行大学)
;
Bioscience and Bioengineering Research Centre, London South Bank University, London SE1 0AA, UK(生物科学与生物工程研究中心,伦敦南银行大学)
;
Department of Physics, Kings College London, London WC2R 2LS, UK(物理系,伦敦国王学院)
Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
协同视见:基于多模态大语言模型的多机器人协作自体空间推理
Kunyu Peng, Zhikun Zhou, Kailun Yang, Di Wen, Ruiping Liu, Yufan Chen, Junwei Zheng, Hao Shi, Yi Zhou, M. Saquib Sarfraz, Danda Pani Paudel, Luc Van Gool
机构
*
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
Hunan University(湖南大学)
;
University of Oxford(牛津大学)
;
Zhejiang University(浙江大学)
;
ETH Zurich(苏黎世联邦理工学院)
;
Ant Group(蚂蚁集团)
专题命中
视觉推理
:multimodal large language model(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV
机构
*
Tencent Youtu Lab(腾讯云图实验室)
;
Sichuan University(四川大学)
;
University of the Chinese Academy of Sciences(中国科学院大学)
;
Fudan University(复旦大学)
;
Zhejiang University(浙江大学)
;
National University of Singapore(新加坡国立大学)
;
Nanyang Technological University(南洋理工大学)
专题命中
视觉推理
:visual reasoning(title);MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
Bringing Multimodal Large Language Models to Infrared-Visible Image Fusion Quality Assessment
将多模态大语言模型引入红外-可见图像融合质量评估
Yuchen Guo, Junli Gong, Yao Lu, Xintong Xu, Yiuming Cheung, Weifeng Su
机构
*
Northwestern University(西北大学)
;
Northeastern University(东北大学)
;
University of Washington(华盛顿大学)
;
Hong Kong Baptist University(香港 Baptist大学)
;
Beijing Normal - Hong Kong Baptist University(北京师范大学-香港 Baptist大学)
专题命中
视觉推理
:multimodal large language model(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
通过文本表示引导推理来解锁多模态大语言模型中的空间推理
Jiacheng Hua, Yishu Yin, Yuhang Wu, Tai Wang, Yifei Huang, Miao Liu
机构
*
College of AI, Tsinghua University, Beijing, China(清华大学人工智能学院,北京,中国)
;
Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室,上海,中国)
;
The University of Tokyo, Tokyo, Japan(东京大学,东京,日本)
专题命中
视觉推理
:multimodal large language model(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV