arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45986 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4644 篇

2508.20013 2025-11-11 cs.LG cs.AI cs.IR 79%

Cross-Platform E-Commerce Product Categorization and Recategorization: A Multimodal Hierarchical Classification Approach

Lotte Gross, Rebecca Walter, Nicole Zoppi, Adrien Justus, Alessandro Gambetti, Qiwei Han, Maximilian Kaiser

机构 * Nova School of Business and Economics(诺瓦商业与经济学院) Nova School of Business(诺瓦商业学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accetped at IEEE BigData 2025, 10 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23769 2025-11-07 cs.CV 79%

TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models

Yao Xiao, Qiqian Fu, Heyi Tao, Yuqun Wu, Zhen Zhu, Derek Hoiem

机构 * Siebel School of Computing and Data Science(塞比尔计算与数据科学学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Published in TMLR, with a J2C Certification

Journal ref Transactions on Machine Learning Research, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00916 2025-11-04 cs.CV 79%

Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs

Yan Shu, Chi Liu, Robin Chen, Derek Li, Bryan Dai

机构 * Ubiquant(乌比量化)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00821 2025-11-04 cs.CV 79%

OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models

Ruoxiang Huang, Xindian Ma, Rundong Kong, Zhen Yuan, Peng Zhang

机构 * Tianjin University(天津大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02095 2025-11-04 cs.CV cs.LG 79%

Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences

Hyojin Bahng, Caroline Chan, Fredo Durand, Phillip Isola

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 图文多模态 :image-text(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01341 2025-11-04 cs.CL 79%

AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding

Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu, Nicolas Chapados, Yoshua Bengio, Enamul Hoque, Christopher Pal, Issam H. Laradji, David Vazquez, Perouz Taslakian, Spandana Gella, Sai Rajeswar

机构 * ServiceNow York University(约克大学) Mila – Quebec AI Institute(魁北克人工智能研究院) École de Technologie Supérieure(魁北克高等技术学院) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) University of Waterloo(滑铁卢大学) CIFAR AI Chair(CIFAR人工智能 chair) Polytechnique Montréal(蒙特利尔理工学院) University of British Columbia(不列颠哥伦比亚大学)

专题命中 图文多模态 :multimodal(title);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24331 2025-10-29 cs.LG cs.CV 79%

What do vision-language models see in the context? Investigating multimodal in-context learning

Gabriel O. dos Santos, Esther Colombini, Sandra Avila

机构 * Instituto de Computação, Universidade Estadual de Campinas (UNICAMP)(计算机学院,Campinas州立大学(UNICAMP))

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21093 2025-10-27 cs.AI 79%

MedAlign: A Synergistic Framework of Multimodal Preference Optimization and Federated Meta-Cognitive Reasoning

Siyong Chen, Jinbo Wen, Jiawen Kang, Tenghui Huang, Xumin Huang, Yuanjia Su, Hudan Pan, Zishao Zhong, Dusit Niyato, Shengli Xie, Dong In Kim

机构 * School of Automation, Guangdong University of Technology(广东工业大学自动化学院) College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) State Key Laboratory of Traditional Chinese Medicine Syndrome, The Second Affiliated Hospital of Guangzhou University of Chinese Medicine, Guangdong Provincial Hospital of Chinese Medicine, Guangdong Provincial Academy of Chinese Medical Sciences(广东省中医药科学院中医证候重点实验室,广州中医药大学第二附属医院,广东省中医院,广东省中医药科学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) Department of Electrical and Computer Engineering, Sungkyunkwan University(成均馆大学电子与计算机工程系)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16198 2025-10-21 cs.CL 79%

EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

Mohamed Gamil, Abdelrahman Elsayed, Abdelrahman Lila, Ahmed Gad, Hesham Abdelgawad, Mohamed Aref, Ahmed Fares

机构 * Department of Electrical Engineering, Faculty of Engineering at Shoubra, Benha University, Cairo 11629, Egypt(电气工程系,谢布拉工程学院,本海大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16036 2025-10-21 cs.CV 79%

IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection

Zewen Li, Zitong Yu, Qilang Ye, Weicheng Xie, Wei Zhuo, Linlin Shen

机构 * School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) School of Computing and Information Technology, Great Bay University(大亚湾大学计算机与信息科技学院) College of Computer Science, Nankai University(南开大学计算机学院) School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室) National Engineering Laboratory of Big Data System Computing Technology, Shenzhen University(大数据系统计算技术国家工程实验室)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Instrumentation and Measurement (TIM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12974 2025-10-20 cs.CV 79%

Scope: Selective Cross-modal Orchestration of Visual Perception Experts

Tianyu Zhang, Suyuchen Wang, Chao Wang, Juan Rodriguez, Ahmed Masry, Xiangru Jian, Yoshua Bengio, Perouz Taslakian

机构 * ServiceNow Université de Montréal(蒙特利尔大学) École de Technologie Supérieure(高级技术学院) University of Waterloo(滑铁卢大学) McGill University(麦吉尔大学) York University(约克大学) CIFAR AI Chair(CIFAR人工智能主席) Mila Law Zero

专题命中 图文多模态 :cross-modal(title);image-text(abstract);分类 cs.CV

Comments 14 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25373 2025-10-17 cs.AI 79%

From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models

Chenyue Zhou, Mingxuan Wang, Yanbiao Ma, Chenxu Wu, Wanyi Chen, Zhe Qian, Xinyu Liu, Yiwei Zhang, Junhao Wang, Hengbo Xu, Fei Luo, Xiaohua Chen, Xiaoshuai Hao, Hehan Li, Andi Zhang, Wenxuan Wang, Kaiyan Zhang, Guoli Jia, Lingling Li, Zhiwu Lu, Yang Lu, Yike Guo

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) Xiamen University(厦门大学) The Hong Kong University of Science and Technology(香港理工大学) Nanyang Technological University(南洋理工大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10342 2025-10-14 cs.CV 79%

Ordinal Scale Traffic Congestion Classification with Multi-Modal Vision-Language and Motion Analysis

Yu-Hsuan Lin

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments 7 pages, 4 figures. Preprint submitted to arXiv in October 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10117 2025-10-14 cs.AI 79%

DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay

Yunxiang Mo, Tianshi Zheng, Qing Zong, Jiayu Liu, Baixuan Xu, Yauwai Yim, Chunkit Chan, Jiaxin Bai, Yangqiu Song

机构 * Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China(计算机科学与工程系,香港科技大学,香港特别行政区,中国)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments EMNLP 2025 Wordplay (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10104 2025-10-14 cs.CV 79%

Answer-Consistent Chain-of-thought Reinforcement Learning For Multi-modal Large Langauge Models

Minbin Huang, Runhui Huang, Chuanyang Zheng, Jingyao Li, Guoxuan Chen, Han Shi, Hong Cheng

机构 * The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09358 2025-10-13 cs.CV 79%

Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models

Qihang Ma, Shengyu Li, Jie Tang, Dingkang Yang, Shaodong Chen, Yingyi Zhang, Chao Feng, Jiao Ran

机构 * ByteDance Douyin Content Group(字节跳动抖音内容团队)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments EMNLP2025. Code is avaible at https://github.com/bytedance/DynamicCoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06092 2025-10-10 cs.CV 79%

Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation

Yachun Mi, Yu Li, Yanting Li, Chen Hui, Tong Zhang, Zhixuan Li, Chenyue Song, Wei Yang Bryan Lim, Shaohui Liu

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18743 2025-10-07 cs.CV 79%

SAR-TEXT: A Large-Scale SAR Image-Text Dataset Built with SAR-Narrator and A Progressive Learning Strategy for Downstream Tasks

Yiguo He, Xinjun Cheng, Junjie Zhu, Chunping Qiu, Jun Wang, Xichuan Zhang, Qiangjuan Huang, Ke Yang

机构 * Intelligent Game and Decision Lab(智能游戏与决策实验室)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments IEEE Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25791 2025-10-01 cs.CV 79%

EchoingECG: An Electrocardiogram Cross-Modal Model for Echocardiogram Tasks

Yuan Gao, Sangwook Kim, Chris McIntosh

机构 * Peter Munk Cardiac Centre, University Health Network (UHN)(彼得·默克心脏中心,大学健康网络) Department of Medical Biophysics, UofT(医学生物物理学系) Ted Rogers Centre for Heart Research, UHN(泰德·罗杰斯心脏病研究中心,大学健康网络) Department of Computer Science, University of Toronto (UofT)(计算机科学系,多伦多大学) Toronto General Hospital Research Institute, UHN(多伦多总医院研究 institute) Department of Medical Imaging, UofT(医学影像学系) Vector Institute, Toronto(向量研究所)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments MICCAI 2025

Journal ref Medical Image Computing and Computer Assisted Intervention - MICCAI 2025. MICCAI 2025. Lecture Notes in Computer Science, vol 15964. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18839 2025-09-24 cs.CV 79%

Benchmarking Vision-Language and Multimodal Large Language Models in Zero-shot and Few-shot Scenarios: A study on Christian Iconography

Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza

机构 * Gianmarco Spinaci Department of Classical Philology and Italian Studies, University of Bologna, Italy Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Gianmarco Spinaci 文艺复兴研究系,博洛尼亚大学,意大利 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利) Lukas Klic Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Lukas Klic 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利) Giovanni Colavizza Department of Classical Philology and Italian Studies, University of Bologna, Italy Department of Communication, University of Copenhagen, Denmark(Giovanni Colavizza 文艺复兴研究系,博洛尼亚大学,意大利 传播系,哥本哈根大学,丹麦)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14958 2025-09-23 cs.CV 79%

Seeing 3D Through 2D Lenses: 3D Few-Shot Class-Incremental Learning via Cross-Modal Geometric Rectification

Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He

机构 * South China University of Technology(南方科技大学) State Key Laboratory of Subtropical Building Science(亚热带建筑科学国家重点实验室) Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information(广东省计算智能与网络信息重点实验室) Ministry of Education Key Laboratory of Big Data and Intelligent Robot(教育部大数据与智能机器人重点实验室) Singapore Management University(新加坡国立大学)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16600 2025-09-19 cs.CV 79%

Image-Text-Image Knowledge Transfer for Lifelong Person Re-Identification with Hybrid Clothing States

Qizao Wang, Xuelin Qian, Bin Li, Yanwei Fu, Xiangyang Xue

机构 * School of Automation, Northwestern Polytechnical University(自动化学院,西北工业大学) School of Computer Science, Shanghai Key Lab of Intelligent Information Processing, Fudan University(计算机学院,上海智能信息处理重点实验室,复旦大学) Shenzhen Research Institute of Northwestern Polytechnical University(西北工业大学深圳研究院)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted by TIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08679 2025-09-17 cs.CV 79%

ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way

Rajarshi Roy, Devleena Das, Ankesh Banerjee, Arjya Bhattacharjee, Kousik Dasgupta, Subarna Tripathi

机构 * Kalyani Government Engineering College(卡利尼政府工程学院) Intel Labs(英特尔实验室)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00284 2025-09-16 cs.RO cs.AI 79%

LightEMMA: Lightweight End-to-End Multimodal Model for Autonomous Driving

Zhijie Qiao, Haowei Li, Zhong Cao, Henry X. Liu

机构 * Department of Civil and Environmental Engineering, University of Michigan(土木与环境工程系,密歇根大学) University of Michigan Transportation Research Institute(密歇根大学交通研究所)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08618 2025-09-11 cs.CV 79%

CLAPS: A CLIP-Unified Auto-Prompt Segmentation for Multi-Modal Retinal Imaging

Zhihao Zhao, Yinzheng Zhao, Junjie Yang, Xiangtong Yao, Quanmin Liang, Shahrooz Faghihroohi, Kai Huang, Nassir Navab, M. Ali Nasseri

机构 * Technical University of Munich(慕尼黑技术大学) Sun Yat-Sen University(中山大学)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments BIBM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08040 2025-09-09 cs.LG cs.AI 79%

BadPromptFL: A Novel Backdoor Threat to Prompt-based Federated Learning in Multimodal Models

Maozhen Zhang, Mengnan Zhao, Wei Wang, Bo Wang

机构 * School of Information and Communication Engineering, Dalian University of Technology(信息与通信工程学院,大连理工大学) School of Computer Science and Technology, Anhui University(计算机科学与技术学院,安徽大学) New Laboratory of Pattern Recognition (NLPR) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS) Institute of Automation, Chinese Academy of Sciences (CASIA)(模式识别新实验室(NLPR)多模态人工智能系统国家重点实验室(MAIS)自动化研究所,中国科学院(CASIA))

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04324 2025-09-05 cs.RO cs.CV 79%

OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection

Chen Hu, Shan Luo, Letizia Gionfrida

机构 * Department of Informatics, King's College London(伦敦国王学院信息学院) Department of Engineering, King's College London(伦敦国王学院工程学院) John A. Paulson School of Engineering and Applied Sciences, Harvard University(哈佛大学约翰·A·保罗森工程与应用科学学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19608 2025-09-04 cs.AI cs.LG 79%

ChordPrompt: Orchestrating Cross-Modal Prompt Synergy for Multi-Domain Incremental Learning in CLIP

Zhiyuan Wang, Bokui Chen

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University, China(清华大学深圳国际研究生院,清华大学,中国)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.AI

Comments Accepted by the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML-PKDD 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02129 2025-09-03 cs.LG cs.CV 79%

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time

Jintao Cheng, Weibin Li, Jiehao Luo, Xiaoyu Tang, Zhijian He, Jin Wu, Yao Zou, Wei Zhang

机构 * Hong Kong University of Science(香港科学与技术大学) South China Normal University, Shanwei, Guangdong, China(华南师范大学,汕尾,广东,中国) Shenzhen Technology University, Shenzhen, Guangdong, China(深圳科技大学,深圳,广东,中国) University of Science(科学大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20860 2025-09-03 cs.CV 79%

FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models

Mainak Singha, Subhankar Roy, Sarthak Mehrotra, Ankit Jha, Moloud Abdar, Biplab Banerjee, Elisa Ricci

机构 * University of Trento(特伦托大学) University of Bergamo(贝拉姆奥大学) Indian Institute of Technology Bombay(印度班加罗尔理工学院) LNMIIT Jaipur(斋普尔LNMIIT) The University of Queensland(昆士兰大学) Fondazione Bruno Kessler(布鲁诺·凯塞勒基金会)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏