机构
*
Institute of Digital Twin, Eastern Institute of Technology(数字孪生研究院,东技术学院)
;
Ocean University of China(中国海洋大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Munich Center for Machine Learning, LMU Munich(慕尼黑机器学习中心,慕尼黑大学)
;
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室)
专题命中
视觉推理
:vision-language model(title);vision language model(abstract);分类 cs.CV
StruVis: Enhancing Reasoning-based Text-to-Image Generation via Thinking with Structured Vision
StruVis: 通过结构化视觉进行推理的文本到图像生成增强
Yuanhuiyi Lyu, Kaiyu Lei, Ziqiao Weng, Xu Zheng, Lutao Jiang, Teng Li, Yangfu Li, Ziyuan Huang, Linfeng Zhang, Xuming Hu
机构
*
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Ant Group(蚂蚁集团)
;
Shanghai Jiao Tong University(上海交通大学)
;
Hong Kong University of Science and Technology(香港科技大学)
;
East China Normal University(华东师范大学)
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
VisioMath: 评估LMMs中基于图的数学推理能力的基准测试
Can Li, Ying Liu, Ting Zhang, Mei Wang, Hua Huang
机构
*
School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)
;
Beijing Key Laboratory of Artificial Intelligence for Education(北京人工智能教育重点实验室)
;
Engineering Research Center of Intelligent Technology and Educational Application, Ministry of Education(教育部智能技术与教育应用工程研究中心)
OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis
OralGPT-Plus:通过强化学习学习使用视觉工具进行全景X射线分析
Yuxuan Fan, Jing Hao, Hong Chen, Jiahao Bao, Yihua Shao, Yuci Liang, Kuo Feng Hung, Hao Tang
机构
*
The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学)
;
Faculty of Dentistry, The University of Hong Kong(香港大学牙医学院)
;
School of Computer Science, Peking University(北京大学计算机学院)
;
Shanghai Jiao Tong University(上海交通大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
MultiHaystack: Benchmarking Multimodal Retrieval and Reasoning over 40K Images, Videos, and Documents
MultiHaystack:用于40,000张图像、视频和文档的多模态检索和推理的基准测试
Dannong Xu, Zhongyu Yang, Jun Chen, Yingfang Yuan, Ming Hu, Lei Sun, Luc Van Gool, Danda Pani Paudel, Chun-Mei Feng
机构
*
INSAIT
;
Lanzhou University(兰州大学)
;
King Abdullah University of Science and Technology(国王阿卜杜勒-阿齐兹科学与技术大学)
;
Heriot-Watt University(赫瑞-沃德大学)
;
Monash University(莫纳什大学)
;
University College Dublin(都柏林大学学院)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
DEX-AR: 一种用于自回归视觉-语言模型的动态可解释性方法
Walid Bousselham, Angie Boggust, Hendrik Strobelt, Hilde Kuehne
机构
*
Tuebingen AI Center University of Tuebingen(图宾根人工智能中心图宾根大学)
;
MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
;
MIT-IBM Watson AI Lab(麻省理工-IBM沃森人工智能实验室)
;
IBM Research(IBM研究院)
Cut to the Chase: Training-free Multimodal Summarization via Chain-of-Events
直击核心:一种无需训练的多模态摘要方法 via 事件链
Xiaoxing You, Qiang Huang, Lingyu Li, Xiaojun Chang, Jun Yu
机构
*
School of Computer Science, Hangzhou Dianzi University(杭州电子科技大学计算机科学学院)
;
School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)智能科学与工程学院)
;
School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院)