arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3454 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3454 篇

2605.27449 2026-05-28 cs.IR cs.AI 77%

Checking Fact with Better Retrieval: Dynamic Contrastive Learning for Evidence Retrieval

用更好的检索核查事实:用于证据检索的动态对比学习

Zhongtian Hua, Yi Luo, Meijia Yu, Yingjie Han

机构 * Zhengzhou University(郑州大学) Henan University of Science and Technology(河南科技大学)

专题命中 跨模态检索 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.AI

AI总结 提出动态自适应对比学习方法DACLR,通过事件级特征提取、两阶段检索和动态对比损失优化,提升多模态证据检索的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14928 2026-05-15 cs.CL 77%

Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA

过程链:用于过程问答的层次视觉语言推理

Guanhua Chen, Yutong Yao, Shenghe Sun, Ci-Jun Gao, Shudong Liu, Lidia S. Chao, Feng Wan, Derek F. Wong

机构 * NLP CT Lab, Department of Computer and Information Science, University of Macau(计算机与信息科学系,澳门大学CT实验室) Department of Electrical and Computer Engineering, University of Macau(电气与计算机工程系,澳门大学) Centre for Cognitive and Brain Sciences, University of Macau(认知与脑科学中心,澳门大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CL

AI总结 本文提出ProcedureVQA基准,针对视觉过程问答任务,通过过程链框架解决现有VLMs在跨模态检索和步骤分解上的不足,实验显示其在六个VLMs上提升达13%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13511 2026-05-04 cs.CV cs.IR 77%

Adapting MLLMs for Nuanced Video Retrieval

为精细化视频检索适应大语言模型

Piyush Bagad, Andrew Zisserman

机构 * Visual Geometry Group, University of Oxford(牛津大学视觉几何组)

专题命中 跨模态检索 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 本文提出一种统一嵌入模型,通过对比损失微调文本训练的多模态大语言模型,以处理视频检索中的时间、否定和多模态细微差别,实现最先进的检索性能。

Comments 38 Pages. Project page at http://bpiyush.github.io/tara-website

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12193 2026-04-20 cs.CV 77%

VeRVE: Versatile Retrieval for Videos via Unified Embeddings

VeRVE:通过统一嵌入实现视频的多功能检索

Shaunak Halbe, Bhagyashree Puranik, Jayakrishnan Unnikrishnan, Kushan Thakkar, Vimal Bhat, Toufiq Parag

机构 * Georgia Institute of Technology(佐治亚理工学院) Amazon(亚马逊) Keystone AI

专题命中 跨模态检索 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 本文提出VeRVE框架,结合语义和时刻级检索能力,支持复杂多模态查询,通过对比对齐视觉和文本嵌入实现高效检索,优于其他多模态大语言模型方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15946 2026-04-17 cs.LG cs.AI cs.CR 77%

Fall into a Pit, Gain in a Wit: Cognitive-Guided Harmful Meme Detection via Misjudgment Risk Pattern Retrieval

跌入陷阱,获得智慧:通过误判风险模式检索的认知引导有害迷因检测

Wenshuo Wang, Ziyou Jiang, Junjie Wang, Mingyang Li, Jie Huang, Yuekai Huang, Zhiyuan Chang, Feiyan Duan, Qing Wang

机构 * State Key Laboratory of Complex System Modeling and Simulation Technology(复杂系统建模与仿真技术国家重点实验室) Science and Technology on Integrated Information System Laboratory(集成信息系统技术研究所) Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 跨模态检索 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.AI

AI总结 本文提出PatMD方法,通过学习并主动缓解潜在误判风险,识别有害迷因的深层误判风险模式,提升多模态大语言模型的检测能力,实验显示在5项有害检测任务中,F1-score和准确率均显著提升。

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08003 2025-10-10 cs.CV 77%

CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning

Weihuang Lin, Yiwei Ma, Jiayi Ji, Xiaoshuai Sun, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(教育部多媒体可信感知与高效计算重点实验室,厦门大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12819 2025-07-18 cs.CV 77%

MCoT-RE: Multi-Faceted Chain-of-Thought and Re-Ranking for Training-Free Zero-Shot Composed Image Retrieval

Jeong-Woo Park, Seong-Whan Lee

机构 * Department of Artificial Intelligence, Korea University(人工智能系,韩国大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

Comments 6 pages, 4 figures, 2025 IEEE International Conference on Systems, Man, and Cybernetics

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12378 2025-07-17 cs.IR cs.CL 77%

Developing Visual Augmented Q&A System using Scalable Vision Embedding Retrieval & Late Interaction Re-ranker

Rachna Saxena, Abhijeet Kumar, Suresh Shanmugam

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);MLLM(abstract);分类 cs.CL

Comments Presented at NLP@IR workshop at SIGIR conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08958 2023-02-20 cs.CV 77%

Towards Unifying Medical Vision-and-Language Pre-training via Soft Prompts

Zhihong Chen, Shizhe Diao, Benyou Wang, Guanbin Li, Xiang Wan

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.07364 2018-07-20 cs.CV 77%

Revisiting Cross Modal Retrieval

Shah Nawaz, Muhammad Kamran Janjua, Alessandro Calefati, Ignazio Gallo

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments 14 pages. Under review at ECCVW (MULA 2018)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14264 2026-07-17 cs.CV cs.CL 新提交 76%

MonteRET: AI Agent Enhancing Multimodal LLMs with Multi-granularity Knowledge Retrieval for Chest CT Report Generation

MonteRET:通过多粒度知识检索增强多模态大语言模型以生成胸部CT报告的人工智能代理

Yi Lin, Yihao Ding, Elana Benishay, Elefterios Trikantzopoulos, David Nauheim, Hanley Ong, Jiang Bian, Hua Xu, Yuzhe Yang, George Shih, Yifan Peng

机构 * Weill Cornell Medicine(威尔康乃尔医学院) University of Western Australia(西澳大利亚大学) Indiana University(印第安纳大学) Regenstrief Institute(瑞根斯特里夫研究所) Yale University(耶鲁大学) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.CL

AI总结 研究针对自动生成胸部CT报告的挑战,提出MonteRET框架,它整合多种特征,通过知识检索和报告重写代理完善报告。经训练和评估,该框架在多方面提升了报告质量,相比其他方法有显著优势,获放射科住院医师青睐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24530 2026-05-26 cs.CL cs.CV 76%

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval

Unveil: 统一视觉-文本集成与蒸馏的多模态文档检索

Hao Sun, Yingyan Hou, Jiayan Guo, Bo Wang, Chunyu Yang, Jinsong Ni, Yan Zhang

机构 * State Key Laboratory of General Artificial Intelligence, Peking University(北京理工大学通用人工智能国家重点实验室) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空航天信息研究所) Key Laboratory of Target Cognition and Application Technology(目标认知与应用技术重点实验室) Beijing Institute of Technology(北京理工大学) Ucap Cloud(Ucap云)

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV、cs.CL

AI总结 提出Unveil框架,通过视觉-文本嵌入和知识蒸馏实现鲁棒的文档检索,兼顾布局与语义信息。

Comments ACL 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23709 2026-01-23 cs.CV cs.AI cs.LG 76%

Skin Lesion Phenotyping via Nested Multi-modal Contrastive Learning

通过嵌套多模态对比学习进行皮肤病变表型分析

Dionysis Christopoulos, Sotiris Spanos, Eirini Baltzi, Valsamis Ntouskos, Konstantinos Karantzalos

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV、cs.AI

AI总结 SLIMP通过嵌套多模态对比学习,结合图像与元数据提升皮肤病变分类性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14154 2026-01-21 cs.CV cs.AI 76%

LLM Augmented Intervenable Multimodal Adaptor for Post-operative Complication Prediction in Lung Cancer Surgery

基于大语言模型的可干预多模态适配器用于肺癌手术后并发症预测

Shubham Pandey, Bhavin Jawade, Srirangaraj Setlur, Venu Govindaraju, Kenneth Seastedt

机构 * University at Buffalo(布法罗大学) Roswell Park Comprehensive Cancer Center(罗斯威尔帕克综合癌症中心)

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.AI

AI总结 MIRACLE通过整合术前临床和放射学数据,利用超球面嵌入空间和干预式深度学习模块,实现肺癌手术后并发症风险的预测与可解释性管理。

Comments Accepted to P2P-CV @ WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19864 2025-12-29 cs.CL cs.AI 76%

HARMON-E: Hierarchical Agentic Reasoning for Multimodal Oncology Notes to Extract Structured Data

HARMON-E:多模态肿瘤病历的分层代理推理以提取结构化数据

Shashi Kant Gupta, Arijeet Pramanik, Jerrin John Thomas, Regina Schwind, Lauren Wiener, Avi Raju, Jeremy Kornbluth, Yanshan Wang, Zhaohui Su, Hrituraj Singh

机构 * Triomics University of Pittsburgh(匹兹堡大学) Ontada

专题命中 跨模态检索 :multimodal(title);分类 cs.CL、cs.AI

AI总结 HARMON-E通过分层代理推理框架,实现对多模态肿瘤病历的结构化数据提取,达到高精度和大规模应用。

Comments 39 Pages, Supplementary Included

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10266 2025-10-30 cs.CV cs.AI 76%

SignMouth: Leveraging Mouthing Cues for Sign Language Translation by Multimodal Contrastive Fusion

Wenfang Wu, Tingting Yuan, Yupeng Li, Daling Wang, Xiaoming Fu

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13888 2025-09-18 cs.CL cs.AI cs.IR 76%

Combating Biomedical Misinformation through Multi-modal Claim Detection and Evidence-based Verification

Mariano Barone, Antonio Romano, Giuseppe Riccio, Marco Postiglione, Vincenzo Moscato

机构 * University of Naples Federico II(那不勒斯费迪里奇二世大学) Northwestern University(西北大学)

专题命中 跨模态检索 :multi-modal(title);分类 cs.CL、cs.AI

Journal ref SIGIR '25: Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22938 2025-08-01 cs.CL cs.AI 76%

A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents

Sumit Soman, H. G. Ranjani, Sujoy Roychowdhury, Venkata Dharma Surya Narayana Sastry, Akshat Jain, Pranav Gangrade, Ayaaz Khan

机构 * Ericsson R&D Bangalore Karnataka India(爱立信研发部班加罗尔卡纳塔克邦印度)

专题命中 跨模态检索 :multi-modal(title);分类 cs.CL、cs.AI

Comments Accepted for publication at the KDD 2025 Workshop on Structured Knowledge for Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15875 2025-07-23 cs.AI cs.MM 76%

Differential Multimodal Transformers

Jerry Li, Timothy Oh, Joseph Hoang, Vardhit Veeramachaneni

机构 * University of California, Riverside(加州大学河滨分校)

专题命中 跨模态检索 :multimodal(title);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14035 2025-06-18 cs.CV cs.AI 76%

SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement

Chelsi Jain, Yiran Wu, Yifan Zeng, Jiale Liu, S hengyu Dai, Zhenwen Shao, Qingyun Wu, Huazheng Wang

机构 * Oregon State University(俄勒冈州立大学) Pennsylvania State University(宾夕法尼亚州立大学) AG2AI, Inc.(AG2AI公司) Johnson & Johnson(强生公司)

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08854 2025-06-11 cs.CV cs.AI 76%

Spatial Transcriptomics Expression Prediction from Histopathology Based on Cross-Modal Mask Reconstruction and Contrastive Learning

Junzhuo Liu, Markus Eckstein, Zhixiang Wang, Friedrich Feuerhake, Dorit Merhof

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV、cs.AI

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08023 2025-06-11 q-bio.BM cs.AI cs.CE cs.CV cs.LG 76%

Aligning Proteins and Language: A Foundation Model for Protein Retrieval

Qifeng Wu, Zhengzhe Liu, Han Zhu, Yizhou Zhao, Daisuke Kihara, Min Xu

机构 * Carnegie Mellon University(卡内基梅隆大学) Purdue University(普渡大学)

专题命中 跨模态检索 :multimodal(abstract,comments);multimodal foundation model(abstract,comments);分类 cs.CV、cs.AI

Comments 4 pages for body, 3 pages for appendix, 11 figures. Accepted to CVPR 2025 Workshop on Multimodal Foundation Models for Biomedicine: Challenges and Opportunities(MMFM-BIOMED)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09519 2024-10-15 cs.CV cs.AI 76%

Pic@Point: Cross-Modal Learning by Local and Global Point-Picture Correspondence

Vencia Herzog, Stefan Suwelack

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV、cs.AI

Comments Accepted at ACML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08544 2024-08-19 cs.CV cs.MM 76%

Scaling up Multimodal Pre-training for Sign Language Understanding

Wengang Zhou, Weichao Zhao, Hezhen Hu, Zecheng Li, Houqiang Li

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.MM

Comments Sign language recognition; Sign language translation; Sign language retrieval

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06357 2024-08-14 cs.CV cs.AI 76%

Algorithm Research of ELMo Word Embedding and Deep Learning Multimodal Transformer in Image Description

Xiaohan Cheng, Taiyuan Mei, Yun Zi, Qi Wang, Zijun Gao, Haowei Yang

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00599 2024-08-02 cs.CV cs.MM eess.IV 76%

Learned Compression of Point Cloud Geometry and Attributes in a Single Model through Multimodal Rate-Control

Michael Rudolph, Aron Riemenschneider, Amr Rizk

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.MM

Comments 20 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06201 2024-06-11 cs.CV cs.AI 76%

2DP-2MRC: 2-Dimensional Pointer-based Machine Reading Comprehension Method for Multimodal Moment Retrieval

Jiajun He, Tomoki Toda

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.AI

Comments Accepted by INTERSPEECH 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08851 2024-03-15 astro-ph.IM cs.CL cs.CV cs.IR cs.LG 76%

PAPERCLIP: Associating Astronomical Observations and Natural Language with Multi-Modal Models

Siddharth Mishra-Sharma, Yiding Song, Jesse Thaler

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV、cs.CL

Comments 17+6 pages, 3+1 figures, 5+2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12846 2024-02-21 cs.CV cs.AI 76%

ConVQG: Contrastive Visual Question Generation with Multimodal Guidance

Li Mi, Syrielle Montariol, Javiera Castillo-Navarro, Xianjie Dai, Antoine Bosselut, Devis Tuia

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.AI

Comments AAAI 2024. Project page at https://limirs.github.io/ConVQG

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12258 2023-08-29 cs.SD cs.CL cs.IR cs.LG eess.AS 76%

Data leakage in cross-modal retrieval training: A case study

Benno Weck, Xavier Serra

专题命中 跨模态检索 :cross-modal(title);分类 cs.CL、eess.AS

Comments 5 pages. Accepted at ICASSP2023

详情

展开后加载摘要…

URL PDF HTML 收藏