机构
*
National Engineering Research Center of Visual Technology, National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(视觉技术国家工程研究中心、多媒体信息处理国家重点实验室、计算机学院、北京大学)
;
Meituan Inc.(美团公司)
;
National Biomedical Imaging Center, Peking University(生物医学成像国家中心、北京大学)
专题命中
视觉推理
:MLLM(title,title_cn);multimodal large language model(abstract);分类 cs.CV
Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models
RLVR能否扩展推理边界?探究视觉-语言模型中的能力扩展
Minghe Shen, Zhuo Zhi, Chonghan Liu, Shuo Xing, Zhengzhong Tu, Che Liu
机构
*
University College London(伦敦大学学院)
;
Samsung R&D Institute UK(三星英国研发院)
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
Texas A&M University(德克萨斯农工大学)
;
Imperial College London(伦敦帝国学院)
机构
*
School of Data Science and Engineering, East China Normal University, Shanghai 200062, China(数据科学与工程学院,华东师范大学,上海200062,中国)
;
Medical Artificial Intelligence Laboratory, Westlake University, Hangzhou 310024, China(医学人工智能实验室,西湖大学,杭州310024,中国)
专题命中
视觉推理
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
CamReasoner:通过结构化空间推理强化相机运动理解
Hang Wu, Yujun Cai, Zehao Li, Haonan Ge, Bowen Sun, Junsong Yuan, Yiwei Wang
机构
*
University of California, Merced(加州大学梅尔德分校)
;
University of Queensland(昆士兰大学)
;
Ant Group(蚂蚁集团)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
University at Buffalo, State University of New York(纽约州立大学布法罗分校)
CommentsWe evaluate whether VLMs can comprehend multi-scale visual stock price data like human analysts with a proposed benchmark, identifying current VLMs' weak predictive power, significant biases, and limited sensitivity to forecast horizons and prompts