Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
检索头能看见图像吗?长上下文视觉语言模型中的多模态检索头
Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao, Yiming Du, Haobo Li, Xiyu Ren, Ginny Wong, Simon See, Lishu Luo, Haodong Duan, Pasquale Minervini, Yangqiu Song
机构
*
HKUST(香港科技大学)
;
University of Edinburgh(爱丁堡大学)
;
CUHK(香港中文大学)
;
NVAITC, NVIDIA, Santa Clara, USA(NVIDIA Santa Clara 分公司)
;
Tsinghua University(清华大学)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Sun Yat-sen University(中山大学)
;
Southern University of Science and Technology(南方科技大学)
;
National University of Singapore(新加坡国立大学)
Industrial3D: A Water-Treatment TLS Point Cloud Dataset and Cross-Paradigm Benchmark for MEP Scene Understanding
Industrial3D: 一个用于工业基础设施的地面激光雷达点云数据集和跨范式基准
Chao Yin, Hongzhe Yue, Qing Han, Difeng Hu, Zhenyu Liang, Fangzhou Lin, Bing Sun, Boyu Wang, Mingkai Li, Wei Yao, Jack C. P. Cheng
机构
*
Department of Civil and Environmental Engineering, The Hong Kong University of Science and Technology(香港科技大学土木与环境工程系)
;
Guangzhou Institute of Geography, Guangdong Academy of Sciences(广东省科学院广州地理研究所)
;
School of Civil Engineering, Southeast University(东南大学土木工程学院)
;
College of Geography and Tourism, Hengyang Normal University(衡阳师范学院地理与旅游学院)
;
Department of Architecture and Civil Engineering, City University of Hong Kong(香港城市大学建筑与土木工程系)
;
S.M.A.R.T. Construction Research Group, New York University Abu Dhabi(纽约大学阿布扎比分校S.M.A.R.T.建筑研究组)
;
Department of the Built Environment, National University of Singapore(新加坡国立大学建筑环境系)
;
Spatial Intelligence and Urban Computing, Institute of Urban Environment, Chinese Academy of Sciences(中国科学院城市环境研究所空间智能与城市计算)
;
State Key Lab of Ecological Security of Regions and Cities, Institute of Urban Environment, Chinese Academy of Sciences(中国科学院城市环境研究所区域与城市生态安全国家重点实验室)
机构
*
School of Artificial Intelligence and Data Science, University of Science and Technology of China(中国科学技术大学人工智能与数据科学学院)
;
Artificial Intelligence Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能学域)
;
RMIT University(皇家墨尔本理工大学)
;
State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)
;
iFLYTEK Research, iFLYTEK(科大讯飞研究院)
机构
*
Hong Kong University of Science and Technology(香港科学与技术大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Harbin Institute of Technology Shenzhen China(哈尔滨工业大学深圳中国)
AULLM++: Structured-Token-Conditioned Large Language Models for Micro-Expression Action Unit Detection
AULLM++:用于微表情动作单元检测的结构化令牌条件大语言模型
Zhishu Liu, Kaishen Yuan, Bo Zhao, Hui Ma, Zitong Yu
机构
*
Great Bay University(广东东莞大学)
;
Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
LLM-based Human Simulations Have Not Yet Been Reliable
基于大语言模型的人类模拟尚不可靠
Qian Wang, Jiaying Wu, Zichen Jiang, Zhenheng Tang, Bingqiao Luo, Nuo Chen, Wei Chen, Huacan Wang, Bingsheng He
机构
*
National University of Singapore(新加坡国立大学)
;
The Hong Kong University of Science and Technology(香港理工大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
Chuanhao Yan, Fengdi Che, Xuhan Huang, Xu Xu, Xin Li, Yizhi Li, Xingwei Qu, Jingzhe Shi, Chenghua Lin, Yaodong Yang, Binhang Yuan, Hang Zhao, Yu Qiao, Bowen Zhou, Jie Fu
机构
*
Shanghai AI Lab(上海人工智能实验室)
;
University of Alberta(阿尔伯塔大学)
;
Tsinghua University(清华大学)
;
Chinese University of Hong Kong, Shenzhen(香港大学(深圳))
;
Hong Kong University of Science and Technology(香港科技大学)
;
Nanyang Technological University(南洋理工大学)
;
University of Manchester(曼彻斯特大学)
;
Peking University(北京大学)
机构
*
The Hong Kong University of Science and Technology(香港科技大学)
;
ShanghaiTech University(上海科技大学)
;
Zhejiang University(浙江大学)
;
Nanjing University(南京大学)
;
Huazhong University of Science and Technology(华中科技大学)