Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum
Wiki-R1: 通过数据和采样课程激励基于知识的多模态推理用于VQA
Shan Ning, Longtian Qiu, Xuming He
机构
*
ShanghaiTech University(上海科技大学)
;
Shanghai Engineering Research Center of Intelligent Vision and Imaging(上海智能视觉与成像工程研究中心)
;
Lingang Laboratory(临港实验室)
专题命中
视觉问答
:visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV
MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding
MedStreamBench: 面向流式与主动式医疗视频理解的时间感知基准
Yuan Wang, Shujian Gao, Songtao Jiang, Zhengyu Hu, Zuozhu Liu
机构
*
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
Department of Electrical and Computer Engineering, Carnegie Mellon University(卡内基梅隆大学电气与计算机工程系)
机构
*
Nanyang Technological University(南洋理工大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)
专题命中
视觉问答
:multimodal large language model(abstract);分类 cs.AI
Teaching Vision-Language-Action Models What to See and Where to Look
教视觉-语言-动作模型看什么和看哪里
Yuguang Yang, Canyu Chen, Zhewen Tan, Yizhi Wang, Zichao Feng, Chunyang Liu, Kehua Sheng, Juan Zhang, Linlin Yang, Baochang Zhang, Yan Wang, Bo Zhang, Xianbin Cao
机构
*
School of Electronic Information Engineering, Beihang University(北京航空航天大学电子信息工程学院)
;
National College for Excellent Engineers, Beihang University(北京航空航天大学卓越工程师学院)
;
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
;
DiDi(滴滴出行)
;
State Key Laboratory of Media Convergence and Communication, Communication University of China(中国传媒大学媒体融合与传播国家重点实验室)
;
School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
;
School of Cyber Science and Technology, Beihang University(北京航空航天大学网络安全科学与技术学院)
;
School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)
MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support
MMIR-TCM:面向中医临床决策支持的记忆集成多模态推理与检索
Lihui Luo, Joongwon Chae, Ziyan Chen, Yang Liu, Siyi Cheng, Weihan Gao, Zelin Zeng, Xiaoming Yin, Samaneh Beheshti Kashi, Dongmei Yu, Lian Zhang, Jing Sui, Zeming Liang, Jiansong Ji, Peter E. Lobie, Peiwu Qin
机构
*
Institute of Biopharmaceutics and Health Engineering, Tsinghua Shenzhen International Graduate School(清华深圳国际研究生院生物医药与健康工程研究院)
;
Chinese Medicine Guangdong Laboratory(广东省中医药实验室)
;
Beijing Normal University(北京师范大学)
;
First Hospital of Hebei Medical University(河北医科大学第一医院)
;
Lishui Hospital of Zhejiang University(浙江大学丽水医院)
;
XiaoMing TCM Hospital(小明中医医院)
专题命中
视觉推理
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.AI
Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning
在思考之前,先学会决策:面向高效视觉推理的主动路由
Yinan Zhou, Haokun Lin, Yichen Wu, Yuxin Chen, Teng Wang, Caifeng Shan, Zhenan Sun, Chen Ma, Li Zhu, Ying Shan
机构
*
Xi’an Jiaotong University(西安交通大学)
;
ARC Lab, Tencent IEG(腾讯IEG ARC实验室)
;
City University of Hong Kong(香港城市大学)
;
Institute of Automation, CAS(中国科学院自动化研究所)
;
Harvard University(哈佛大学)
;
Nanjing University(南京大学)
机构
*
School of Information Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)信息科学与技术学院)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)计算机科学与技术学院)
;
Pengcheng Laboratory(鹏城实验室)
;
School of Computer Science and Technology, Shandong Jianzhu University(山东建筑大学计算机科学与技术学院)
;
Zhongguancun Academy(中关村学院)
;
City University of Hong Kong(香港城市大学)
EduArt: An educational-level benchmark for evaluating art history knowledge in large language models
EduArt:评估大型语言模型艺术史知识的教育级基准
Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza
机构
*
University of Bologna(博洛尼亚大学)
;
Villa i Tatti – The Harvard University Center for Italian Renaissance Studies(哈佛大学意大利文艺复兴研究中心(I Tatti))
;
University of Copenhagen(哥本哈根大学)
Simulating Validity: Modal Decoupling in MLLM Generated Feedback on Science Drawings
模拟有效性:在MLLM生成的科学图表反馈中的模态解耦
Arne Bewersdorff, Nejla Yuruk, Xiaoming Zhai
机构
*
University of Georgia, AI4STEM Education Center(佐治亚大学AI4STEM教育中心)
;
Gazi University, Department of Mathematics and Science Education(加齐大学数学与科学教育系)
专题命中
视觉定位与Grounding
:MLLM(title,title_cn);grounding(summary_cn,abstract);multimodal large language model(abstract);分类 cs.AI