CLIP4VI-ReID: Learning Modality-shared Representations via CLIP Semantic Bridge for Visible-Infrared Person Re-identification
CLIP4VI-ReID:通过CLIP语义桥学习模态共享表征用于可见光-红外行人重识别
Xiaomei Yang, Xizhan Gao, Sijie Niu, Fa Zhu, Guang Feng, Xiaofeng Qu, David Camacho
机构
*
Shandong Key Laboratory of Ubiquitous Intelligent Computing, School of Information Science and Engineering, University of Jinan(山东 ubiquitous 智能计算重点实验室,信息科学与工程学院,济南大学)
;
College of Information Science and Technology & College of Artificial Intelligence, Nanjing Forestry University(信息科学与技术学院及人工智能学院,南京林业大学)
;
Computer Systems Engineering Department, Universidad Politécnica de Madrid(计算机系统工程系,马德里理工大学)
Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?
我们真的需要参数超过10亿的多模态情感语言模型吗?
Kaiwen Zheng, Junchen Fu, Wenhao Deng, Hu Han, Joemon M. Jose, Xuri Ge
机构
*
University of Glasgow(格拉斯哥大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
School of Artificial Intelligence, Shandong University(山东大学人工智能学院)
机构
*
National and Local Joint Engineering Laboratory of Computer Aided Design, School of Software Engineering, Dalian University(大连大学软件工程学院计算机辅助设计国家地方联合工程实验室)
;
Department of Radiology, Xinhua Hospital Affiliated to Dalian University(大连大学附属新华医院放射科)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
University of California, San Francisco(加州大学旧金山分校)
;
Yale University(耶鲁大学)
;
The Hong Kong Polytechnic University(香港理工大学)
SVL: Empowering Spiking Neural Networks for Efficient 3D Open-World Understanding
SVL:基于脉冲的视觉-语言预训练用于高效的3D开放世界理解
Xuerui Qiu, Peixi Wu, Yaozhi Wen, Shaowei Gu, Yuqi Pan, Xinhao Luo, Bo XU, Guoqi Li
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Future Technology, University of Chinese Academy of Sciences(中国科学院大学未来技术学院)
;
Zhongguancun Academy(中关村学院)
;
University of Science and Technology of China(中国科学技术大学)
;
Peking University(北京大学)
机构
*
Hong Kong University of Science and Technology(香港科学与技术大学)
;
Zhejiang University(浙江大学)
;
National University of Singapore(新加坡国立大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Peking University(北京大学)
;
Independent Researcher(独立研究者)
机构
*
MemTensor (Shanghai) Technology Co., Ltd.(墨芯(上海)科技有限公司)
;
Renmin University of China(中国人民大学)
;
National University of Singapore(新加坡国立大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Tongji University(同济大学)
专题命中
多模态Agent
:multimodal(abstract);multimodal foundation model(abstract);分类 cs.CL
Do VLMs Align Better with Humans than LLMs during Natural Reading?
VLMs 在自然阅读中可能不会全局性地增强与人类的对齐性优于 LLMs
Jinzhou Wu, Zhengwu Ma, Jixing Li, Baoping Tang, Zitong Lu
机构
*
Department of Mechanical and Vehicle Engineering, Chongqing University(重庆大学机械与车辆工程学院)
;
Department of Linguistics and Translation, City University of Hong Kong(香港城市大学语言学与翻译系)
;
McGovern Institute for Brain Research, Massachusetts Institute of Technology(麻省理工学院麦戈文脑科学研究所)