CommentsThis manuscript has been withdrawn by the authors. It reproduced the methodology of Gardinazzi et al., arXiv:2410.11042, without citation, and utilized code and data from the associated repository (github.com/RitAreaSciencePark/ZigZagLLMs) without disclosure or violate the MIT License. A revised future version with full attribution may be prepared. For any feedback, please contact Pengcheng Zheng
机构
*
Kyoto University(京都大学)
;
NII LLMC(日本国立信息与通信技术研究所语言模型中心)
;
RIKEN AIP(日本理化学研究所先进理工研究所)
;
Case Western Reserve University(凯斯西储大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
The University of Osaka(大阪大学)
;
University of Tokyo(东京大学)
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.CV
机构
*
Department of Automation, Tsinghua University, Beijing, China(清华大学自动化系)
;
Department of Electronic Engineering, Tsinghua University, Beijing, China(清华大学电子工程系)
;
Zhongguancun Academy, Beijing, China(中关村学院)
;
China Agricultural University, Beijing, China(中国农业大学)
;
Peking University, Beijing, China(北京大学)
;
Beijing Institute of Technology, Beijing, China(北京理工大学)
;
Institute of Automation, Chinese Academy of Sciences, Beijing, China(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学人工智能学院)
;
WeChat Vision, Tencent Inc, Beijing, China(微信视觉,腾讯公司)
Self-supervised Hierarchical Visual Reasoning with World Model
基于世界模型的自监督分层视觉推理
Yuanfei Xu, Lin Liu, Wengang Zhou, Mingxiao Feng, Houqiang Li
机构
*
Department of Electronic Engineering and Information Science, University of Science and Technology of China(电子工程与信息科学系,中国科学技术大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(人工智能研究院,合肥综合性国家科学中心)
机构
*
Department of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院)
;
Department of Computer Science, National Taiwan University(国立台湾大学计算机科学系)
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.CV
ChatSR: Multimodal Large Language Models for Scientific Formula Discovery
ChatSR:用于科学公式发现的多模态大语言模型
Yanjie Li, Lina Yu, Weijun Li, Min Wu, Liping Zhang, Jingyi Liu, Yusong Deng, Mingzhu Wan, Xin Ning
机构
*
AnnLab, Institute of Semiconductors, Chinese Academy of Sciences, Beijing, China(安 lab,半导体研究所,中国科学院,北京,中国)
;
School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing, China(电子、电气与通信工程学院,中国科学院大学,北京,中国)
;
Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing 101408, China(先进交叉科学学院,中国科学院大学,北京101408,中国)
;
College of Materials Science and Opto-Electronic Technology, University of Chinese Academy of Sciences, Beijing, 100049, China(材料科学与光电技术学院,中国科学院大学,北京100049,中国)
;
School of Integrated Circuits, University of Chinese Academy of Sciences, Beijing 100049, China(集成电路学院,中国科学院大学,北京100049,中国)
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.AI
机构
*
School of Mathematics, Tianjin University, Tianjin, China(天津大学数学学院)
;
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China(中国科学院信息工程研究所)
;
Shanghai Advanced Institute of Finance (SAIFS), East China Normal University, Shanghai, China(上海先进金融研究所(SAIFS),东华大学)
;
ENN Group, Digital Technology Research Institute, China(ENN集团,数字技术研究院)
;
Nanyang Technological University, Singapore(南洋理工大学)
;
National University of Singapore, Singapore(新加坡国立大学)
机构
*
School of Computer Science and Engineering, School of Intelligence Science and Engineering and Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Southeast University, China(东南大学计算机科学与工程学院、智能科学与工程学院及新一代人工智能技术及其跨学科应用关键实验室,中国)
;
Wangxuan Institute of Computer Technology, Peking University, China(北京大学王轩计算机技术研究所,中国)
;
University of Copenhagen, Denmark(丹麦哥本哈根大学)
GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language Models
GeoArena:在大视觉-语言模型中评估开放世界地理推理
Pengyue Jia, Yingyi Zhang, Xiangyu Zhao, Sharon Li
机构
*
Department of Data Science, City University of Hong Kong(香港城市大学数据科学系)
;
Department of Computer Sciences, University of Wisconsin-Madison(威斯康星大学麦迪逊分校计算机科学系)
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
AVRT:通过单模态教师模型实现音频-视觉推理迁移
Edson Araujo, Saurabhchand Bhati, M. Jehanzeb Mirza, Brian Kingsbury, Samuel Thomas, Rogerio Feris, James R. Glass, Hilde Kuehne
机构
*
University of Tübingen, Germany(图宾根大学)
;
MIT, Cambridge MA, USA(麻省理工学院)
;
IBM Research, USA(IBM研究院)
;
MIT-IBM Watson AI Lab, USA(麻省理工-IBM沃森人工智能实验室)
;
Tuebingen AI Center, Germany(图宾根人工智能中心)
HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning
HyperGVL:超图理解与推理中大视觉-语言模型的基准测试与改进
Yanbin Wei, Chun Kang, Siwei Li, Haoxuan Che, Yang Chen, Hua Liu, Jian Liu, Zhuang Liu, Can Ouyang, Fei Xing, Lei Sha, Rui Liu, Yu Zhang, James Kwok
机构
*
Southern University of Science and Technology(南方科技大学)
;
Hong Kong University of Science and Technology(香港科技大学)
;
Huawei Research(华为研究)
;
Beihang University(北航)