Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
计算机视觉中的推理:分类、模型、任务与方法论
Ayushman Sarkar, Zhenyu Yu, Mohd Yamani Idna Idris
机构
*
Department of Computer Science and Engineering, Birbhum Institute of Engineering and Technology(计算机科学与工程系,比罗尔理工学院)
;
College of Computer Science and Artificial Intelligence, Fudan University(计算机科学与人工智能学院,复旦大学)
;
Faculty of Computer Science and Information Technology, Universiti Malaya(计算机科学与信息技术学院,马来亚大学)
Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning
Visual-Seeker:通过主动视觉推理实现视觉原生多模态智能搜索
Zhengbo Zhang, Changtao Miao, Jinbo Su, Zhaowen Zhou, Chunxia Zhang, Xukai Wang, Ruiqi Liu, Kaiyuan Zheng, Jiansheng Cai, Bo Zhang, Zhe Li, Shiming Xiang, Ying Yan
机构
*
School of Artificial Intelligence UCAS(中国科学院大学人工智能学院)
;
Institute of Automation CAS(中国科学院自动化研究所)
;
Ant Digital Technologies Ant Group(蚂蚁数字科技蚂蚁集团)
;
RUC(中国人民大学)
;
BIT(北京理工大学)
专题命中
视觉推理
:visual reasoning(title,abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.AI
Position: The Systemic Lack of Agency in Visual Reasoning
立场:视觉推理中系统性的能动性缺失
Yizhao Huang, Haoyang Chen, Shiqin Wang, Pohsun Huang, Jiayuan Li, Haoyuan Du, Yandong Shi, Zheng Wang, Zhixiang Wang
机构
*
National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science, Wuhan University(武汉大学计算机学院人工智能研究所多媒体软件工程研究中心)
;
Hubei Key Laboratory of Multimedia(湖北省多媒体与网络通信工程重点实验室)
;
Zhongguancun Academy, Beijing, China. 100094(中关村学院,北京,中国)
;
School of Automation, Beijing Institute of Technology(北京理工大学自动化学院)
;
Shanda AI Research Tokyo(上海大势人工智能研究东京)
CycliST: A Video Language Model Benchmark for Reasoning on Cyclical State Transitions
CycliST:用于循环状态转换推理的视频语言模型基准
Simon Kohaut, Daniel Ochs, Shun Zhang, Benedict Flade, Julian Eggert, Kristian Kersting, Devendra Singh Dhami
机构
*
Artificial Intelligence and Machine Learning Lab, TU Darmstadt(人工智能与机器学习实验室,图腾斯达特技术大学)
;
Konrad Zuse School of Excellence in Learning and Intelligent Systems (ELIZA)(Konrad Zuse 学校(ELIZA))
;
Honda Research Institute Europe GmbH, Offenbach, Germany(本田欧洲研究院,奥芬巴赫,德国)
;
Uncertainty in Artificial Intelligence Group, TU Eindhoven(人工智能不确定性小组,埃因霍温技术大学)
;
Hessian Center for AI (hessian.AI)(黑森人工智能中心(hessian.AI))
;
Center for Cognitive Science(认知科学中心)
;
German Center for Artificial Intelligence (DFKI)(德国人工智能中心(DFKI))
UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics
UrbanWell: 面向时空城市福祉分析的多模态大语言模型基准测试
Yanxin Xi, Xiang Su, Jie Feng, Yu Liu, Sasu Tarkoma, Pan Hui
机构
*
University of Helsinki(赫尔辛基大学)
;
Zhongguancun Academy(中关村学院)
;
University of Oxford(牛津大学)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.AI
机构
*
Kean University(基恩大学)
;
Case Western Reserve University(凯斯西储大学)
;
University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
The Ohio State University(俄亥俄州立大学)
;
Tongji University(同济大学)
;
Duke Kunshan University(昆山杜克大学)
;
Royal Melbourne Institute of Technology(皇家墨尔本理工大学)
CommentsAccepted by SIGIR-ICTIR 2026, Oral Presentation
Journal refProceedings of the 2026 International ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR '26), July 25, 2026, Melbourne, VIC, Australia. ACM, New York, NY, USA, 12 pages
机构
*
Tsinghua University(清华大学)
;
Chongqing University(重庆大学)
;
Peking University(北京大学)
;
ZenoMind AI
;
Xi’an Jiaotong University(西安交通大学)
;
Beijing Institute of Technology(北京理工大学)
;
Southeast University(东南大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Joy Future Academy(京东探索研究院)
;
The University of Hong Kong(香港大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.AI
GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs
GroupToM-Bench: 多模态大语言模型中群体心智理论和非线性社会涌现的基准测试
Weidong Tang, Jierui Li, Yueling Hou, Zihan Mei, Can Zhang, Xinyan Wan, Zhiyuan Liang, Pengfei Zhou, Yang You, Wangbo Zhao
机构
*
Xidian University(西安电子科技大学)
;
National University of Singapore(新加坡国立大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
University of Science and Technology of China(中国科学技术大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV