机构
*
School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院)
;
School of Computer Science and Technology, Tiangong University(天津工大学计算机科学与技术学院)
;
Key Research Center for Surface Monitoring and Analysis of Relics, State Administration of Cultural Heritage(文物表面监测与分析关键研究中心,国家文物局)
机构
*
State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology, Anhui University, Hefei, China(光电信息采集与防护技术国家重点实验室,安徽大学,合肥,中国)
;
School of Artificial Intelligence, Anhui University, Hefei, China(人工智能学院,安徽大学,合肥,中国)
;
College of Intelligence Science and Technology, National University of Defense Technology, Changsha, China(智能科学与技术学院,国防科技大学,长沙,中国)
机构
*
Institute of Trustworthy Embodied AI, Fudan University(可信具身AI研究院,复旦大学)
;
Shanghai Key Laboratory of Multimodal Embodied AI(上海多模态具身AI重点实验室)
;
City University of Hong Kong(香港城市大学)
SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
SPEX:一种用于光谱遥感图像土地覆盖提取的视觉-语言模型
Dongchen Si, Di Wang, Erzhong Gao, Xiaolei Qin, Liu Zhao, Jing Zhang, Minqiang Xu, Jianbo Zhan, Jianshe Wang, Lin Liu, Bo Du, Liangpei Zhang
机构
*
College of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院)
;
iFlytek Co., Ltd.(iFlytek公司)
;
National Engineering Research Center of Speech and Language Information Processing(语音与语言信息处理国家工程研究中心)
;
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
Zhongguancun Academy(中关村学院)
;
National Engineering Research Center for Multimedia Software(多媒体软件国家工程研究中心)
;
Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University(湖北省多媒体与网络通信工程重点实验室,武汉大学)
;
State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(测绘遥感信息工程国家重点实验室,武汉大学)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
驯服持续音频-视觉分割中的模态纠缠
Yuyang Hong, Qi Yang, Tao Zhang, Zili Wang, Zhaojin Fu, Kun Ding, Bin Fan, Shiming Xiang
机构
*
School of Artificial Intelligence, UCAS(人工智能学院,UCAS)
;
MAIS, Institute of Automation(自动化研究所MAIS)
;
School of Intelligent Science and Technology, University of Science and Technolog Beijing(智能科学与技术学院,北京理工大学)
Listening with the Eyes: Benchmarking Egocentric Co-Speech Grounding across Space and Time
用眼睛倾听:跨时空的自体视觉共指基准测试
Weijie Zhou, Xuantang Xiong, Zhenlin Hu, Xiaomeng Zhu, Chaoyang Zhao, Honghui Dong, Zhengyou Zhang, Ming Tang, Jinqiao Wang
机构
*
Beijing Jiaotong University(北京交通大学)
;
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences (CASIA)(基础模型研究中心、自动化研究所、中国科学院(CASIA))
;
Tencent Robotics X(腾讯机器人X)
;
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology (HKUST)(计算机科学与工程系、香港科学与技术大学(HKUST))
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳学院)