AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu, Nicolas Chapados, Yoshua Bengio, Enamul Hoque, Christopher Pal, Issam H. Laradji, David Vazquez, Perouz Taslakian, Spandana Gella, Sai Rajeswar
机构
*
ServiceNow
;
York University(约克大学)
;
Mila – Quebec AI Institute(魁北克人工智能研究院)
;
École de Technologie Supérieure(魁北克高等技术学院)
;
Université de Montréal(蒙特利尔大学)
;
McGill University(麦吉尔大学)
;
University of Waterloo(滑铁卢大学)
;
CIFAR AI Chair(CIFAR人工智能 chair)
;
Polytechnique Montréal(蒙特利尔理工学院)
;
University of British Columbia(不列颠哥伦比亚大学)
MedAlign: A Synergistic Framework of Multimodal Preference Optimization and Federated Meta-Cognitive Reasoning
Siyong Chen, Jinbo Wen, Jiawen Kang, Tenghui Huang, Xumin Huang, Yuanjia Su, Hudan Pan, Zishao Zhong, Dusit Niyato, Shengli Xie, Dong In Kim
机构
*
School of Automation, Guangdong University of Technology(广东工业大学自动化学院)
;
College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院)
;
State Key Laboratory of Traditional Chinese Medicine Syndrome, The Second Affiliated Hospital of Guangzhou University of Chinese Medicine, Guangdong Provincial Hospital of Chinese Medicine, Guangdong Provincial Academy of Chinese Medical Sciences(广东省中医药科学院中医证候重点实验室,广州中医药大学第二附属医院,广东省中医院,广东省中医药科学院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院)
;
Department of Electrical and Computer Engineering, Sungkyunkwan University(成均馆大学电子与计算机工程系)
机构
*
School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
;
School of Computing and Information Technology, Great Bay University(大亚湾大学计算机与信息科技学院)
;
College of Computer Science, Nankai University(南开大学计算机学院)
;
School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院)
;
Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室)
;
National Engineering Laboratory of Big Data System Computing Technology, Shenzhen University(大数据系统计算技术国家工程实验室)
专题命中
图文多模态
:multimodal(title,abstract);分类 cs.CV
CommentsAccepted by IEEE Transactions on Instrumentation and Measurement (TIM)
Scope: Selective Cross-modal Orchestration of Visual Perception Experts
Tianyu Zhang, Suyuchen Wang, Chao Wang, Juan Rodriguez, Ahmed Masry, Xiangru Jian, Yoshua Bengio, Perouz Taslakian
机构
*
ServiceNow
;
Université de Montréal(蒙特利尔大学)
;
École de Technologie Supérieure(高级技术学院)
;
University of Waterloo(滑铁卢大学)
;
McGill University(麦吉尔大学)
;
York University(约克大学)
;
CIFAR AI Chair(CIFAR人工智能主席)
;
Mila
;
Law Zero
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院)
;
Xiamen University(厦门大学)
;
The Hong Kong University of Science and Technology(香港理工大学)
;
Nanyang Technological University(南洋理工大学)
EchoingECG: An Electrocardiogram Cross-Modal Model for Echocardiogram Tasks
Yuan Gao, Sangwook Kim, Chris McIntosh
机构
*
Peter Munk Cardiac Centre, University Health Network (UHN)(彼得·默克心脏中心,大学健康网络)
;
Department of Medical Biophysics, UofT(医学生物物理学系)
;
Ted Rogers Centre for Heart Research, UHN(泰德·罗杰斯心脏病研究中心,大学健康网络)
;
Department of Computer Science, University of Toronto (UofT)(计算机科学系,多伦多大学)
;
Toronto General Hospital Research Institute, UHN(多伦多总医院研究 institute)
;
Department of Medical Imaging, UofT(医学影像学系)
;
Vector Institute, Toronto(向量研究所)
Benchmarking Vision-Language and Multimodal Large Language Models in Zero-shot and Few-shot Scenarios: A study on Christian Iconography
Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza
机构
*
Gianmarco Spinaci Department of Classical Philology and Italian Studies, University of Bologna, Italy Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Gianmarco Spinaci 文艺复兴研究系,博洛尼亚大学,意大利 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利)
;
Lukas Klic Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Lukas Klic 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利)
;
Giovanni Colavizza Department of Classical Philology and Italian Studies, University of Bologna, Italy Department of Communication, University of Copenhagen, Denmark(Giovanni Colavizza 文艺复兴研究系,博洛尼亚大学,意大利 传播系,哥本哈根大学,丹麦)
Seeing 3D Through 2D Lenses: 3D Few-Shot Class-Incremental Learning via Cross-Modal Geometric Rectification
Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He
机构
*
South China University of Technology(南方科技大学)
;
State Key Laboratory of Subtropical Building Science(亚热带建筑科学国家重点实验室)
;
Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information(广东省计算智能与网络信息重点实验室)
;
Ministry of Education Key Laboratory of Big Data and Intelligent Robot(教育部大数据与智能机器人重点实验室)
;
Singapore Management University(新加坡国立大学)
Image-Text-Image Knowledge Transfer for Lifelong Person Re-Identification with Hybrid Clothing States
Qizao Wang, Xuelin Qian, Bin Li, Yanwei Fu, Xiangyang Xue
机构
*
School of Automation, Northwestern Polytechnical University(自动化学院,西北工业大学)
;
School of Computer Science, Shanghai Key Lab of Intelligent Information Processing, Fudan University(计算机学院,上海智能信息处理重点实验室,复旦大学)
;
Shenzhen Research Institute of Northwestern Polytechnical University(西北工业大学深圳研究院)
LightEMMA: Lightweight End-to-End Multimodal Model for Autonomous Driving
Zhijie Qiao, Haowei Li, Zhong Cao, Henry X. Liu
机构
*
Department of Civil and Environmental Engineering, University of Michigan(土木与环境工程系,密歇根大学)
;
University of Michigan Transportation Research Institute(密歇根大学交通研究所)
BadPromptFL: A Novel Backdoor Threat to Prompt-based Federated Learning in Multimodal Models
Maozhen Zhang, Mengnan Zhao, Wei Wang, Bo Wang
机构
*
School of Information and Communication Engineering, Dalian University of Technology(信息与通信工程学院,大连理工大学)
;
School of Computer Science and Technology, Anhui University(计算机科学与技术学院,安徽大学)
;
New Laboratory of Pattern Recognition (NLPR) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS) Institute of Automation, Chinese Academy of Sciences (CASIA)(模式识别新实验室(NLPR)多模态人工智能系统国家重点实验室(MAIS)自动化研究所,中国科学院(CASIA))
OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection
Chen Hu, Shan Luo, Letizia Gionfrida
机构
*
Department of Informatics, King's College London(伦敦国王学院信息学院)
;
Department of Engineering, King's College London(伦敦国王学院工程学院)
;
John A. Paulson School of Engineering and Applied Sciences, Harvard University(哈佛大学约翰·A·保罗森工程与应用科学学院)
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
Jintao Cheng, Weibin Li, Jiehao Luo, Xiaoyu Tang, Zhijian He, Jin Wu, Yao Zou, Wei Zhang
机构
*
Hong Kong University of Science(香港科学与技术大学)
;
South China Normal University, Shanwei, Guangdong, China(华南师范大学,汕尾,广东,中国)
;
Shenzhen Technology University, Shenzhen, Guangdong, China(深圳科技大学,深圳,广东,中国)
;
University of Science(科学大学)
机构
*
University of Trento(特伦托大学)
;
University of Bergamo(贝拉姆奥大学)
;
Indian Institute of Technology Bombay(印度班加罗尔理工学院)
;
LNMIIT Jaipur(斋普尔LNMIIT)
;
The University of Queensland(昆士兰大学)
;
Fondazione Bruno Kessler(布鲁诺·凯塞勒基金会)