SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
Mauro Orazio Drago, Luca Carlini, Pelinsu Celebi Balyemez, Dennis Pierantozzi, Chiara Lena, Cesare Hassan, Danail Stoyanov, Elena De Momi, Sophia Bano, Mobarak I. Hoque
机构
*
Dipartimento di Elettronica, Informazione e Bioingegneria (DEIB)(电子、信息与生物工程系)
;
Politecnico di Milano(米兰理工大学)
;
IRCCS Humanitas Research Hospital(IRCCS人类itas研究医院)
;
UCL Hawkes Institute and Department of Computer Science(UCL Hawkes研究所和计算机科学系)
;
University College London(伦敦大学学院)
;
University of Manchester(曼彻斯特大学)
EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting
Wei Zhang, Miaoxin Cai, Yaqian Ning, Tong Zhang, Yin Zhuang, Shijian Lu, He Chen, Jun Li, Xuerui Mao
机构
*
School of Interdisciplinary Science, Beijing Institute of Technology(交叉科学学院,北京理工大学)
;
College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)
;
National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing, Beijing Institute of Technology(空间智能信息处理国家重点实验室,北京理工大学)
;
School of Optics and Photonics, Beijing Institute of Technology(光学与 photonics 学院,北京理工大学)
;
State Key Laboratory of Explosion Science and Safety Protection, Beijing(爆炸科学与安全防护国家重点实验室,北京)
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(1 多媒体信息处理国家重点实验室,计算机学院,北京大学)
;
Beijing Academy of Artificial Intelligence(2 北京人工智能研究院)
;
Institute of Automation, Chinese Academy of Sciences(3 中国科学院自动化研究所)
;
Beihang University(4 北航大学)
机构
*
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)
;
Tsinghua University(清华大学)
;
Peking University(北京大学)
;
Beijing Institute of Technology(北京理工大学)
专题命中
视觉空间推理
:reasoning(abstract)
CommentsUpdate v3 of the NeurIPS 2025 Datasets and Benchmarks paper (v2), including additional evaluations of state-of-the-art multimodal large language models. Project page: https://anywhere-3d.github.io/
CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection
Fanxiao Li, Jiaying Wu, Canyuan He, Wei Zhou
机构
*
School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院)
;
National University of Singapore(新加坡国立大学)
;
Engineering Research Center of Cyberspace, Yunnan University(云南大学网络空间研究院)
ZING-3D: Zero-shot Incremental 3D Scene Graphs via Vision-Language Models
Pranav Saxena, Jimmy Chiun
机构
*
Birla Institute of Technology and Science Pilani, K.K Birla Goa Campus(比拉理工学院和科学学院,比拉戈阿校园)
;
Department of Mechanical Engineering, College of Design and Engineering, National University of Singapore(设计与工程学院机械工程系,新加坡国立大学)
Seeing through Uncertainty: Robust Task-Oriented Optimization in Visual Navigation
Yiyuan Pan, Yunzhe Xu, Zhe Liu, Hesheng Wang
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
School of Automation and Intelligent Sensing(自动化与智能感知学院)
;
Key Laboratory of System Control and Information Processing(系统控制与信息处理重点实验室)
;
Ministry of Education of China(中华人民共和国教育部)
;
National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家级重点实验室)
;
Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院)
Immersive Explainability: Visualizing Robot Navigation Decisions through XAI Semantic Scene Projections in Virtual Reality
Jorge de Heuvel, Sebastian Müller, Marlene Wessels, Aftab Akhtar, Christian Bauckhage, Maren Bennewitz
机构
*
University of Bonn(波恩大学)
;
Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔人工智能与机器学习研究所)
;
Center for Robotics(机器人中心)
;
University of Mainz(美因茨大学)
;
Fraunhofer Institute for Intelligent Analysis and Information Systems IAIS(弗劳恩霍夫智能分析与信息系统研究所)
机构
*
Simon Fraser University(西蒙弗雷泽大学)
;
University of Padova(帕多瓦大学)
;
Fondazione Bruno Kessler (FBK)(布鲁诺·克塞勒基金会)
;
Alberta Machine Intelligence Institute (Amii)(阿尔伯塔人工智能研究所)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
Jingli Lin, Chenming Zhu, Runsen Xu, Xiaohan Mao, Xihui Liu, Tai Wang, Jiangmiao Pang
机构
*
Shanghai AI Laboratory(上海人工智能实验室)
;
Shanghai Jiao Tong University(上海交通大学)
;
The University of Hong Kong(香港大学)
;
The Chinese University of Hong Kong(香港中文大学)
专题命中
视觉空间推理
:reasoning(abstract)
Comments30 pages, a benchmark designed to evaluate Online Spatio-Temporal understanding from the perspective of an agent actively exploring a scene. Project Page: https://rbler1234.github.io/OSTBench.github.io/