Rethinking Video-Language Model from the Language Input Perspective
从语言输入角度重新思考视频-语言模型
Xiang Fang, Wanlong Fang, Changshuo Wang, Xiaoye Qu, Daizong Liu
机构
*
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院)
;
Nanyang Technological University, Singapore(新加坡南洋理工大学)
;
University College London(伦敦大学学院)
;
Huazhong University of Science and Technology(华中科技大学)
;
Wuhan University(武汉大学)
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
如何以及想象什么?统一多模态模型中的视觉思维用于跨视角空间推理
Qian Yang, Ankur Sikarwar, Huy Le, Le Zhang, Zhuan Shi, Perouz Taslakian, Aishwarya Agrawal
机构
*
Mila - Québec AI Institute(蒙特利尔AI研究所)
;
Université de Montréal(蒙特利尔大学)
;
McGill University(麦吉尔大学)
;
ServiceNow AI Research(ServiceNow人工智能研究)
;
Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)
GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision
GeoSolver: 利用细粒度过程监督扩展遥感中的测试时推理
Lang Sun, Ronghao Fu, Zhuoran Duan, Haoran Liu, Xueyan Liu, Bo Yang
机构
*
College of Computer Science and Technology(计算机科学与技术学院)
;
Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education Jilin University(教育部符号计算与知识工程重点实验室)
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
专题命中
视觉推理
:visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV
机构
*
Tsinghua University, SIGS(清华大学 SIGS)
;
Meituan(美团)
;
The Chinese University of Hong Kong(香港中文大学)
;
National University of Singapore(新加坡国立大学)
;
LMMs-Lab(LMMs实验室)
;
University of California, Los Angeles(加州大学洛杉矶分校)
机构
*
Computer Vision Institute, School of Computer Science and Software Engineering, Shenzhen University(计算机视觉研究院,计算机科学与软件工程学院,深圳大学)
;
Guangdong Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室,深圳大学)
;
School of Computer Science, University of Nottingham Ningbo China(Nottingham Ningbo 中国计算机科学学院)
;
Department of Electrical and Computer Engineering, National University of Singapore(电子与计算机工程系,新加坡国立大学)
;
Department of Radiation Oncology, Stanford University(放射肿瘤科,斯坦福大学)
;
Sun Yat-sen University(中山大学)
;
School of Computer Science, University of Nottingham(计算机科学学院,Nottingham大学)
机构
*
School of Computer Science, National Engineering Research Center for Multimedia Software and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University, China(计算机学院,国家多媒体软件工程技术研究中心和湖北多媒体与网络通信工程重点实验室,武汉大学,中国)
;
Alibaba Group, Hangzhou, China(阿里巴巴集团,杭州,中国)
;
Independent Researcher(独立研究者)
;
Department of Machine Learning, Mohamed bin Zayed University of Artificial Intelligence, United Arab Emirates(机器学习系,Mohamed bin Zayed人工智能大学,阿拉伯联合酋长国)
专题命中
视觉推理
:grounding(abstract);multimodal large language model(abstract);分类 cs.CV
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
OPPO AI Center(OPPO AI中心)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract_cn);分类 cs.CV
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Tsinghua University(清华大学)
;
National University of Singapore(新加坡国立大学)
;
Zhongguancun Academy(中关村学院)
专题命中
视觉推理
:visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV
CheXTemporal: A Dataset for Temporally-Grounded Reasoning in Chest Radiography
CheXTemporal:用于胸部X光影像中时间感知推理的数据集
Eva Prakash, Yunhe Gao, Chong Wang, Justin Xu, Neal Prakash, Arne Michalson, Seena Dehkharghani, Eun Kyoung Hong, Julie Bauml, Roger Boodoo, Jean-Benoit Delbrouck, Sophie Ostmeier, Curtis Langlotz
机构
*
Stanford University(斯坦福大学)
;
University of Oxford(牛津大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
HOPPR
;
University Hospital Zurich(苏黎世大学医院)
Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning
Pest-Thinker: 通过强化学习学习像昆虫学家一样思考和推理
Xueheng Li, Yu Wang, Tao Hu, Ji Huang, Ke Cao, Qize Yang, Rui Li, Jie Zhang, Chengjun Xie
机构
*
Institute of Intelligent Machines, Hefei Institute of Physical Science, Chinese Academy of Sciences(智能机器研究所、合肥物理科学研究所、中国科学院)
;
University of Science and Technology of China(中国科学技术大学)
;
Zhongke Hefei Institute of Technology Innovation Engineering(中科创新合肥科技研究院)
;
Hefei University of Technology(合肥工业大学)
专题命中
视觉推理
:visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV
AstroAlertBench: Evaluating the Accuracy, Reasoning, and Honesty of Multimodal LLMs in Astronomical Classification
AstroAlertBench: 评估多模态大语言模型在天文学分类中的准确性、推理和诚实性
Claire Chen, Jiabao Sean Xiao, Shuze Daniel Liu, Facundo Perez Paolino, Luke Handley, Theophile Jegou du Laz, Ricky Nilsson, Alice Zou, Matthew Graham, Ashish Mahabal
机构
*
California Institute of Technology(加州理工学院)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Purdue University(普渡大学)
专题命中
视觉推理
:grounding(abstract);multimodal large language model(abstract);分类 cs.AI