机构
*
Department of Control Science and Engineering, College of Electronic and Information Engineering, and Shanghai Institute of Intelligent Science and Technology, Tongji University(同济大学电子与信息工程学院控制科学与工程系,以及上海智能科学与技术研究院)
;
Department of Electronic Engineering, Faculty of Engineering, The Chinese University of Hong Kong(香港中文大学工程学院电子工程系)
;
Shanghai Operation Robot Co., Ltd.(上海微创医疗机器人(集团)股份有限公司)
RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning
RSICCLLM:用于遥感图像变化描述的多模态大语言模型
Yelin Wang, Zijia Song, Shuo Ye, Chuanguang Yang, Miaoyu Wang, Yong Xu, Zhulin An, Yongjun Xu, Zitong Yu
机构
*
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China(人工智能安全国家重点实验室,计算技术研究所,中国科学院,北京,中国)
;
Great Bay University, Dongguan, China(东莞Great Bay大学,中国)
;
Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学深圳学院,中国)
;
Dongguan Key Laboratory for Intelligence and Information Technology, Dongguan, China(东莞智能与信息科技重点实验室,中国)
机构
*
College of Intelligence and Computing, Tianjin University(天津大学智能与计算学部)
;
Faculty of Computer Science and Artificial Intelligence, Shenzhen University of Advanced Technology(深圳理工大学计算机科学与人工智能学院)
;
School of Artificial Intelligence, Nanchang University(南昌大学人工智能学院)
Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment
回答前看清楚:通过显著性驱动的感知重新对齐减轻LVLMs中的幻觉
Pengxu Chen, Yao Zhu, Guangming Zhu, Jun Sheng, Jincai Huang, Xiangyang Ji, Liang Zhang
机构
*
Xidian University(西安电子科技大学)
;
Tsinghua University(清华大学)
;
Shanghai Road Transport Development Center(上海市道路运输发展中心)
;
Hunan Institute of Advanced Technology(湖南先进技术研究院)
SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models
SAB-LVLM: 面向大型视觉-语言模型的重要性感知二值化
Qi Lyu, Jiahua Dong, Baichen Liu, Xudong Wang, Mingfei Han, Yulun Zhang, Fahad Shahbaz Khan, Salman Khan, Lianqing Liu, Zhi Han
机构
*
State Key Laboratory of Robotics and Intelligent Systems(机器人学国家重点实验室)
;
Shenyang Institute of Automation, Chinese Academy of Sciences(中国科学院沈阳自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Shanghai Jiao Tong University(上海交通大学)
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
ReasonCLIP-58M: CLIP的视觉基础常识推理监督
Sicheng Zhang, Muzammal Naseer, Binzhu Xie, Naufal Suryanto, Shi Qiu, Jamal Bentahar, Naveed Akhtar, Mubarak Shah
机构
*
Khalifa University(卡利法大学)
;
University of Western Australia(西澳大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
University of Melbourne(墨尔本大学)
;
University of Central Florida(佛罗里达中央大学)