CommentsReframed CVT-Bench around spatial-state integrity, added real-world benchmark derived from 3dSGG, context controls and other minor changes. 21 pages, 10 figures, 15 tables. Project page: this https URL (https://shanmukha-here.github.io/CVT-Bench)
Topology-Aware Reasoning over Incomplete Knowledge Graph with Graph-Based Soft Prompting
基于拓扑的不完整知识图谱推理与图基软提示
Shuai Wang, Xixi Wang, Yinan Yu
机构
*
Chalmers University of Technology and University of Gothenburg, Sweden(楚姆勒大学技术学院和哥德堡大学,瑞典)
;
Technical University of Denmark, Kgs. Lyngby, Denmark(丹麦技术大学,基斯勒吕辛,丹麦)
机构
*
National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室)
;
School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)
;
AMAP, Alibaba Group(阿里集团AMAP)
;
School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机与人工智能学院)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.LG
Ruiqi Wu, Yuang Yao, Tengfei Ma, Chenran Zhang, Na Su, Tao Zhou, Geng Chen, Wen Fan, Yi Zhou
机构
*
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Department of Ophthalmology, The First Affiliated Hospital of Nanjing Medical University(南京医科大学第一附属医院眼科学系)
;
School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
;
School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.AI
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
OmniAD:基于多模态推理的工业异常检测与理解
Shifang Zhao, Yiheng Lin, Lu Han, Yao Zhao, Yunchao Wei
机构
*
Institute of Information Science, Beijing Jiaotong University(北京交通大学信息科学学院)
;
Visual Intelligence + X International Joint Laboratory of the Ministry of Education(教育部视觉智能+X国际合作实验室)
;
Key Laboratory of Noise and Vibration Research, Institute of Acoustics, Chinese Academy of Sciences(中国科学院声学研究所噪声与振动重点实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
Vision-Language-Policy Model for Dynamic Robot Task Planning
用于动态机器人任务规划的视觉-语言-策略模型
Jin Wang, Kim Tien Ly, Jacques Cloete, Jin Jin, Nikos Tsagarakis, Ioannis Havoutis
机构
*
Dynamic Robot Systems Group, Oxford Robotics Institute, University of Oxford(牛津大学机器人研究所动态机器人系统组)
;
Humanoids and Human-Centered Mechatronics (HHCM), Istituto Italiano di Tecnologia(意大利技术研究所人形与以人为中心的机电系统)
TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning
TSRouter:用于时间序列推理的动态模态-模型选择
Fangxu Yu, Tao Feng, Dehai Min, Lu Cheng, Ge Liu, Tianyi Zhou
机构
*
University of Maryland, College Park(马里兰大学帕克分校)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
;
MBZUAI(Mohamed Bin Zayed University of Artificial Intelligence)
DECIS: Dual-Evidence Corrective Verification for Interpretable Strabismus Diagnostic Decision-Making
MAGIS:基于证据的多智能体推理用于可解释的斜视临床决策
Xikai Tang, Yifan Wang, Jiafan Zhuang, Li Luo, Jinming Guo, Xiaoling Xie, Jiacheng Liu, Peiwei Wei, Lihao Zhong, Xiaoli Kang, Jie Cen, Guangqiang Yin, Kunliang Qiu, Ce Zheng, Zhun Fan
机构
*
School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)
;
Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(电子科技大学深圳高等研究院)
;
Joint Shantou International Eye Center of Shantou University and The Chinese University of Hong Kong(汕头大学·香港中文大学联合汕头国际眼科中心)
;
School of Artificial Intelligence, Guangzhou City Polytechnic(广州城市职业学院人工智能学院)
;
Medical College, Shantou University(汕头大学医学院)
;
College of Engineering, Shantou University(汕头大学工学院)
;
Department of Ophthalmology, Xinhua Hospital Affiliated to Shanghai Jiaotong University School of Medicine(上海交通大学医学院附属新华医院眼科)
;
Shenzhen Loop Area Institute(深圳河套学院)
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
MMR-V:未言明的是什么?视频中多模态深度推理的基准测试
Kejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li, Shangqing Tu, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院,北京,中国)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Tsinghua University(清华大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing
VCG-Bench:迈向统一的视觉导向基准,用于结构化生成与编辑
Xiaoyan Su, Peijie Dong, Zhenheng Tang, Song Tang, Yuyao Zhai, Kaitao Lin, Liang Chen, Gai Yuhang, Yuyu Luo, Qiang Wang, Xiaowen Chu
机构
*
The Hong Kong University of Science and Technology (GuangZhou)(香港科学与技术大学(广州))
;
Huawei Technologies Co., Ltd(华为技术有限公司)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
South China University of Technology(华南理工大学)