arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3437 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3437 篇

2511.11216 2025-11-17 cs.CV 83%

Positional Bias in Multimodal Embedding Models: Do They Favor the Beginning, the Middle, or the End?

Kebin Wu, Fatima Albreiki

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments accepted to AAAI 2026 main track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10997 2025-11-17 cs.CV cs.LG 83%

PROMISE: Prompt-Attentive Hierarchical Contrastive Learning for Robust Cross-Modal Representation with Missing Modalities

Jiajun Chen, Sai Cheng, Yutao Yuan, Yirui Zhang, Haitao Yuan, Peng Peng, Yi Zhong

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted by AAAI'2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10641 2025-11-07 cs.LG cs.AI 83%

Test-Time Warmup for Multimodal Large Language Models

Nikita Rajaneesh, Thomas Zollo, Richard Zemel

机构 * Columbia University(哥伦比亚大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17784 2025-10-31 cs.LG cs.AI 83%

Revealing Multimodal Causality with Large Language Models

Jin Li, Shoujin Wang, Qi Zhang, Feng Liu, Tongliang Liu, Longbing Cao, Shui Yu, Fang Chen

机构 * University of Technology Sydney(技术大学悉尼大学) Tongji University(同济大学) University of Melbourne(墨尔本大学) University of Sydney(悉尼大学) Macquarie University(麦考瑞大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06888 2025-10-09 cs.IR cs.AI 83%

M3Retrieve: Benchmarking Multimodal Retrieval for Medicine

Arkadeep Acharya, Akash Ghosh, Pradeepika Verma, Kitsuchart Pasupa, Sriparna Saha, Priti Singh

机构 * Indian Institute of Technology Patna(印度理工学院帕纳瓦分校) King Mongkut’s Institute of Technology Ladkrabang(拉差丹awan技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments EMNLP Mains 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03458 2025-10-07 cs.CL 83%

Omni-Embed-Nemotron: A Unified Multimodal Retrieval Model for Text, Image, Audio, and Video

Mengyao Xu, Wenfei Zhou, Yauhen Babakhin, Gabriel Moreira, Ronay Ak, Radek Osmulski, Bo Liu, Even Oldridge, Benedikt Schifferer

机构 * NVIDIA

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25177 2025-09-30 cs.CV 83%

Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding

Bingkui Tong, Jiaer Xia, Kaiyang Zhou

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) Hong Kong Baptist University(香港 Baptist大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19994 2025-09-30 cs.CV 83%

Improving Generalizability and Undetectability for Targeted Adversarial Attacks on Multimodal Pre-trained Models

Zhifang Zhang, Jiahan Zhang, Shengjie Zhou, Qi Wei, Shuo He, Feng Liu, Lei Feng

机构 * Southeast University(东南大学) Johns Hopkins University(约翰霍普金斯大学) Chongqing University(重庆大学) Nanyang Technological University(南洋理工大学) University of Melbourne(墨尔本大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18450 2025-09-30 cs.CL 83%

BRIT: Bidirectional Retrieval over Unified Image-Text Graph

Ainulla Khan, Yamada Moyuru, Srinidhi Akella

机构 * Fujitsu Research India(富士通印度研究)

专题命中 跨模态检索 :image-text(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted in EMNLP-2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15472 2025-09-29 cs.CV 83%

Efficient Multimodal Dataset Distillation via Generative Models

Zhenghao Zhao, Haoxuan Wang, Junyi Wu, Yuzhang Shang, Gaowen Liu, Yan Yan

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Central Florida(中央佛罗里达大学) Cisco Research(思科研究)

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21556 2025-09-29 cs.CL 83%

VAT-KG: Knowledge-Intensive Multimodal Knowledge Graph Dataset for Retrieval-Augmented Generation

Hyeongcheol Park, Jiyoung Seo, MinHyuk Jang, Hogun Park, Ha Dam Baek, Gyusam Chang, Hyeonsoo Im, Sangpil Kim

机构 * Korea University(韩国大学) Sungkyunkwan University(全北大学) Hanwha Systems(韩华系统)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Project Page: https://vatkg.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16597 2025-09-23 cs.CL 83%

MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models

Luyan Zhang

机构 * Northeastern University(东北大学)

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CL

Comments 13 pages, 6 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17265 2025-09-23 cs.LG cs.AI 83%

SUA: Stealthy Multimodal Large Language Model Unlearning Attack

Xianren Zhang, Hui Liu, Delvin Ce Zhang, Xianfeng Tang, Qi He, Dongwon Lee, Suhang Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Amazon(亚马逊) University of Sheffield(谢菲尔德大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments EMNLP25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02962 2025-09-18 cs.CV 83%

Resilient Multimodal Industrial Surface Defect Detection with Uncertain Sensors Availability

Shuai Jiang, Yunfeng Ma, Jingyu Zhou, Yuan Bian, Yaonan Wang, Min Liu

机构 * School of Artificial Intelligence and Robotics(人工智能与机器人学院) National Engineering Research Center for Robot Visual Perception and Control Technology(机器人视觉感知与控制技术国家工程研究中心) Hunan University(湖南大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to IEEE/ASME Transactions on Mechatronics

Journal ref IEEE/ASME Transactions on Mechatronics, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01275 2025-09-17 cs.AI 83%

Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D

Artemis Panagopoulou, Le Xue, Honglu Zhou, silvio savarese, Ran Xu, Caiming Xiong, Chris Callison-Burch, Mark Yatskar, Juan Carlos Niebles

机构 * Salesforce AI Reseach(Salesforce人工智能研究院) University of Pennsylvania(宾夕法尼亚大学)

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11054 2025-09-16 cs.IT cs.CV math.IT 83%

Rate-Distortion Limits for Multimodal Retrieval: Theory, Optimal Codes, and Finite-Sample Guarantees

Thomas Y. Chen

机构 * Department of Computer Science, Columbia University(计算机科学系,哥伦比亚大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments ICCV MRR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10467 2025-09-16 cs.IR cs.AI cs.CL cs.CV cs.MM 83%

DSRAG: A Domain-Specific Retrieval Framework Based on Document-derived Multimodal Knowledge Graph

Mengzheng Yang, Yanfei Ren, David Osei Opoku, Ruochang Li, Peng Ren, Chunxiao Xing

机构 * School of Software, Henan University, Kaifeng 475004, China(河南大学软件学院) BNRist, DCST, RIIT, Tsinghua University, Beijing 100084, China(清华大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 12 pages, 5 figures. Accepted to the 22nd International Conference on Web Information Systems and Applications (WISA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08897 2025-09-12 cs.CV cs.AI cs.CL cs.MM 83%

Recurrence Meets Transformers for Universal Multimodal Retrieval

Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

机构 * Department of Education and Humanities, University of Modena and Reggio Emilia(教育与人文学院, Modena and Reggio Emilia大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01360 2025-09-03 cs.CV cs.LG 83%

M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision

Che Liu, Zheng Jiang, Chengyu Fang, Heng Guo, Yan-Jie Zhou, Jiaqi Qu, Le Lu, Minfeng Xu

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Imperial College London(帝国理工学院) Tsinghua University(清华大学) Hupan Lab(华潘实验室)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16701 2025-09-03 cs.IR cs.CL 83%

AlzheimerRAG: Multimodal Retrieval Augmented Generation for Clinical Use Cases using PubMed articles

Aritra Kumar Lahiri, Qinmin Vivian Hu

机构 * Department of Computer Science, Toronto Metropolitan University(计算机科学系,多伦多 Metropolitan 大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Journal ref Machine Learning and Knowledge Extraction. 2025; 7(3):89

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20188 2025-08-29 cs.CV cs.LG 83%

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study

Max Torop, Masih Eskandar, Nicholas Kurtansky, Jinyang Liu, Jochen Weber, Octavia Camps, Veronica Rotemberg, Jennifer Dy, Kivanc Kose

机构 * Northeastern University(东北大学) Memorial Sloan Kettering Cancer Center(纪念斯隆凯特琳癌症中心)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20057 2025-08-28 cs.MM 83%

ProMSC-MIS: Prompt-based Multimodal Semantic Communication for Multi-Spectral Image Segmentation

Haoshuo Zhang, Yufei Bo, Meixia Tao

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

Comments arXiv admin note: text overlap with arXiv:2508.17920

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02826 2025-08-27 cs.CV 83%

Confidence-driven Gradient Modulation for Multimodal Human Activity Recognition: A Dynamic Contrastive Dual-Path Learning Approach

Panpan Ji, Junni Song, Yifan Lu, Hang Xiao, Hanyu Liu, Chao Li

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13901 2025-08-20 cs.RO cs.CV 83%

Multimodal Data Storage and Retrieval for Embodied AI: A Survey

Yihao Lu, Hao Tang

机构 * School of Economics and Management, South China Normal University(经济管理学院,华南师范大学) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01226 2025-08-05 cs.IR cs.MM 83%

CM$^3$: Calibrating Multimodal Recommendation

Xin Zhou, Yongjie Wang, Zhiqi Shen

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.MM

Comments Working Paper: https://github.com/enoche/CM3

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23331 2025-08-01 cs.CV 83%

Contrastive Learning-Driven Traffic Sign Perception: Multi-Modal Fusion of Text and Vision

Qiang Lu, Waikit Xiu, Xiying Li, Shenyu Hu, Shengbo Sun

机构 * School of Intelligent Systems Engineering, Sun Yat-sen University(中山大学智能系统工程学院) Guangdong Provincial Key Laboratory of Intelligent Transportation System(广东省智能交通系统重点实验室)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 11pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21917 2025-07-30 cs.CV 83%

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

Nicola Fanelli, Gennaro Vessio, Giovanna Castellano

机构 * Department of Computer Science University of Bari Aldo Moro(计算机科学系巴里大学Aldo Moro)

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20326 2025-07-29 cs.LG cs.AI 83%

MIPS: a Multimodal Infinite Polymer Sequence Pre-training Framework for Polymer Property Prediction

Jiaxi Wang, Yaosen Min, Xun Zhu, Miao Li, Ji Wu

机构 * Department of Electronic Engineering, Tsinghua University Beijing China Zhongguancun Institute of Artificial Intelligence Beijing China Department of Electronic Engineering \& College of AI, Tsinghua University Beijing National Research Center for Information Science Department of Electronic Engineering, Tsinghua University Zhongguancun Institute of Artificial Intelligence

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 14 pages, 8 figures, accepted by ACM Multimedia 2025 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01334 2025-07-22 cs.IR cs.CV 83%

Composed Multi-modal Retrieval: A Survey of Approaches and Applications

Kun Zhang, Jingyu Li, Zhe Li, Jingjing Zhang, Fan Li, Yandong Liu, Rui Yan, Zihang Jiang, Nan Chen, Lei Zhang, Yongdong Zhang, Zhendong Mao, S. Kevin Zhou

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22589 2025-07-15 cs.CV 83%

LIGHT: Multi-Modal Text Linking on Historical Maps

Yijun Lin, Rhett Olson, Junhan Wu, Yao-Yi Chiang, Jerod Weinman

机构 * University of Minnesota(明尼苏达大学) Grinnell College(格林内尔学院)

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted at ICDAR2025

详情

展开后加载摘要…

URL PDF HTML 收藏