arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-24 至 2026-03-24 共收录 18 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 18 篇

2511.01946 2026-03-24 cs.LG cond-mat.mtrl-sci cs.AI physics.chem-ph 88%

COFAP: A Universal Framework for COFs Adsorption Prediction through Designed Multi-Modal Extraction and Cross-Modal Synergy

COFAP:通过设计的多模态提取和跨模态协同的通用COFs吸附预测框架

Zihan Li, Mingyang Wan, Mingyu Gao, Xishi Tai, Zhongshan Chen, Xiangke Wang, Feifan Zhang

机构 * College of Science, College of Information and Electrical Engineering(科学学院,信息与电气工程学院) China Agricultural University(中国农业大学) Qingdao Institute of Software, College of Computer Science and Technology(软件研究所,计算机科学与技术学院) China University of Petroleum (East China)(中国石油大学(华东)) Weifang university(潍坊大学) College of Environmental Science and Engineering(环境科学与工程学院) North China Electric Power University(华北电力大学) College of Science(科学学院)

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 本文提出COFAP框架,通过深度学习提取多模态结构和化学特征,并利用跨模态注意力机制融合特征,实现高效COFs吸附预测,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19961 2026-03-24 cs.CL cs.IR 83%

Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval

解锁多模态文档智能:从当前成就到视觉文档检索的未来前沿

Yibo Yan, Jiahao Huo, Guanbo Feng, Mingdong Ou, Yi Cao, Xin Zou, Shuliang Liu, Yuanhuiyi Lyu, Yu Huang, Jungang Li, Kening Zheng, Xu Zheng, Philip S. Yu, James Kwok, Xuming Hu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Alibaba Cloud Computing(阿里云计算) Hong Kong University of Science and Technology(香港科技大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

AI总结 本文综述了视觉文档检索领域,探讨了多模态大语言模型时代下的方法演进与挑战,提出未来发展方向。

Comments Under review. This version updates the relevant works released before 15 March, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23306 2026-03-24 cs.CV 83%

ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding

ThinkOmni: 通过指导解码将文本推理提升到多模态场景

Yiran Guan, Sifan Tu, Dingkang Liang, Linghao Zhu, Jianzhong Ju, Zhenbo Luo, Jian Luan, Yuliang Liu, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 跨模态检索 :omni-modal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 ThinkOmni提出一种无需训练和数据的框架,通过指导解码将文本推理扩展到多模态场景,实验显示在多个多模态推理基准上取得显著提升。

Comments Accept by ICLR 2026, Code: https://github.com/1ranGuan/thinkomni

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14408 2026-03-24 cs.CV cs.AI 81%

Feature Recalibration Based Olfactory-Visual Multimodal Model for Enhanced Rice Deterioration Detection

基于特征重校准的嗅觉-视觉多模态模型用于增强稻米劣变检测

Rongqiang Zhao, Hengrui Hu, Yijing Wang, Mingchun Sun, Jie Liu

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) National Key Laboratory of Smart Farm Technologies and Systems(国家智能农业技术与系统重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出基于特征重校准的嗅觉-视觉多模态模型,通过改进的特征嵌入构造器和重校准注意力网络,提升稻米劣变检测的准确性和效率,相比传统方法提升11.51%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04846 2026-03-24 cs.CV 79%

Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language Models

多范式协同对抗攻击多模态大语言模型

Yuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou, Xiao-Jun Wu, Josef Kittler

机构 * School of Artificial Intelligence and Computer Science, Jiangnan University(江南大学人工智能与计算机科学学院) Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(Surrey 大学视觉、语音和信号处理中心)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 针对多模态大语言模型的多范式协同对抗攻击方法,通过聚合视觉和语言特征进行联合优化,提升对抗示例的可转移性,实验表明优于现有方法。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02258 2026-03-24 cs.CV 79%

Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning

病理代理RAG:通过强化学习实现多模态代理检索增强生成用于病理学视觉语言模型

Wenchuan Zhang, Jingru Guo, Hengzhe Zhang, Penghao Zhang, Jie Chen, Shuwan Zhang, Zhang Zhang, Yuhao Yi, Hong Bu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出Patho-AgenticRAG,通过强化学习实现多模态代理检索增强生成,解决病理学视觉语言模型在高分辨率、复杂组织结构和临床语义上的挑战,提升诊断准确性。

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(35): 29921-29929, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20970 2026-03-24 cs.CV 79%

GraPHFormer: A Multimodal Graph Persistent Homology Transformer for the Analysis of Neuroscience Morphologies

GraPHFormer:一种多模态图持久同调变换器,用于神经科学形态学分析

Uzair Shah, Marco Agus, Mahmoud Gamal, Mahmood Alzubaidi, Corrado Cali, Pierre J. Magistretti, Abdesselam Bouzerdoum, Mowafa Househ

机构 * Hamad Bin Khalifa University(哈马德·本·卡西姆大学) University of Turin(都灵大学) BESE, King Abdullah University of Science and Technology(贝赛,国王阿卜杜勒阿齐兹大学科学与技术学院) University of Wollongong(沃林根大学) Neuroscience Institute Cavalieri Ottolenghi(卡瓦利埃-奥托伦奇神经科学研究所) Université Grenoble-Alpes(格勒诺布尔阿尔卑斯大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 GraPHFormer通过CLIP式对比学习统一拓扑和图结构分析,利用持久图像编码和树状LSTM编码器,实现对神经形态的高精度识别与分类,优于传统方法。

Comments Accepted to IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06638 2026-03-24 cs.CV cs.AI 73%

StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering

StaR-KVQA:用于隐式知识视觉问答的结构化推理轨迹

Zhihao Wen, Wenkang Wei, Yuan Fang, Xingtong Yu, Hui Zhang, Weicheng Zhu, Xin Zhang

机构 * Ant International, Ant Group(蚂蚁集团国际部,蚂蚁集团) School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院) School of Computing and Information Systems, Singapore Management University(新加坡管理学院计算与信息系统学院) Anhui Provincial Key Laboratory of High Performance Computing(安徽省高性能计算重点实验室)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 StaR-KVQA通过引入双路径结构化推理轨迹提升隐式知识视觉问答的准确性与推理透明度,采用自蒸馏方法构建轨迹增强数据集,无需外部检索工具,在OK-VQA基准上实现11.3%的精度提升。

Comments 8+3+3 pages, code: https://github.com/jianyingzhihe/StaR-KVQA

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21886 2026-03-24 cs.IR cs.CV 70%

ADaFuSE: Adaptive Diffusion-generated Image and Text Fusion for Interactive Text-to-Image Retrieval

ADaFuSE: 适应性扩散生成图像与文本融合用于交互式文本到图像检索

Zhuocheng Zhang, Xingwu Zhang, Kangheng Liang, Guanxuan Li, Richard Mccreadie, Zijun Long

机构 * Hunan University(湖南大学) University of Glasgow(格拉斯哥大学)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 ADaFuSE通过引入双分支融合机制,结合自适应门控和语义感知专家混合,提升交互式文本到图像检索的性能,实现更稳健的多模态融合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21135 2026-03-24 cs.CV cs.AI 62%

One Pool Is Not Enough: Multi-Cluster Memory for Practical Test-Time Adaptation

一个池子不够:多簇内存用于实用测试时适应

Yu-Wen Tseng, Xingyi Zheng, Ya-Chen Wu, I-Bin Liao, Yung-Hui Li, Hong-Han Shuai, Wen-Huang Cheng

机构 * National Taiwan University(国立台湾大学) National Yang Ming Chiao Tung University(国立阳明交通大学) Hon Hai Research Institute(宏海研究院)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出多簇内存(MCM)以解决实用测试时适应中单池存储的不足,通过多簇组织提升适应稳定性,实验证明在多个数据集上均取得显著提升。

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20818 2026-03-24 cs.CV cs.AI 62%

PlanaReLoc: Camera Relocalization in 3D Planar Primitives via Region-Based Structure Matching

PlanaReLoc:通过基于区域的结构匹配实现3D平面原语的相机重定位

Hanqiao Ye, Yuzhou Liu, Yangdong Liu, Shuhan Shen

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出PlanaReLoc,利用3D平面原语和地图进行轻量级6自由度相机重定位,通过深度匹配和统一嵌入空间实现可靠的跨模态结构对应。

Comments Accepted by CVPR 2026. 20 pages, 15 figures. Code at https://github.com/3dv-casia/PlanaReLoc

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21925 2026-03-24 cs.AI 57%

Guideline-grounded retrieval-augmented generation for ophthalmic clinical decision support

基于指南的检索增强生成用于眼科临床决策支持

Shuying Chen, Sen Cui, Zhong Cao

机构 * University of International Business and Economics(国际商务经济大学) Tsinghua University(清华大学) Heidelberg Institute of Global Health(海德堡全球健康研究院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 本文提出Oph-Guid-RAG系统,通过多模态视觉RAG方法提升眼科临床问答与决策支持,通过可控检索框架和多模态推理提升证据基础和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01049 2026-03-24 cs.CV cs.RO 57%

KeySG: Hierarchical Keyframe-Based 3D Scene Graphs

KeySG:基于关键帧的层次化3D场景图

Abdelrhman Werby, Dennis Rotondi, Fabio Scaparro, Kai O. Arras

机构 * Socially Intelligent Robotics Lab, Institute for Artificial Intelligence University of Stuttgart, Germany(社会智能机器人实验室,人工智能研究所,斯图加特大学,德国)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

AI总结 KeySG通过层次化3D场景图结构,结合关键帧提取多模态信息,提升复杂环境中的语义推理与规划能力,优于现有方法。

Comments Code and video are available at https://keysg-lab.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13876 2026-03-24 cs.CV 57%

Scene Prior Filtering for Depth Super-Resolution

场景先验过滤用于深度超分辨率

Zhengxue Wang, Zhiqiang Yan, Ming-Hsuan Yang, Jinshan Pan, Guangwei Gao, Ying Tai, Jian Yang

机构 * PCA Lab, Nanjing University of Science and Technology, China(南京理工大学科学与工程学院实验室) National University of Singapore(新加坡国立大学) University of California, USA, and Yonsei University, South Korea(美国加州大学和韩国延世大学) PCA Lab, Nanjing University, China(南京大学实验室)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出SPFNet网络,利用大尺度模型的表面法线和语义图作为先验信息,通过多模态先验传播和一对一先验嵌入减少纹理干扰并提升边缘表示,实现在真实和合成数据集上的超分辨率性能。

Comments Accepted to IJCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21083 2026-03-24 cs.CV 57%

Hierarchical Text-Guided Brain Tumor Segmentation via Sub-Region-Aware Prompts

基于子区域感知提示的层次化文本引导脑肿瘤分割

Bahram Mohammadi, Ta Duc Huy, Afrouz Sheikholeslami, Qi Chen, Vu Minh Hieu Phan, Sam White, Minh-Son To, Xuyun Zhang, Amin Beheshti, Luping Zhou, Yuankai Qi

机构 * Macquarie University(麦考瑞大学) Adelaide University(阿德莱德大学) Flinders University(弗林德斯大学) University of Sydney(悉尼大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 本文提出TextCSP框架,通过文本引导和子区域感知提示提升脑肿瘤分割精度,实验显示在Dice和HD95指标上优于现有方法。

Comments 10 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12985 2026-03-24 cs.LG cs.CV 57%

Angular Gradient Sign Method: Uncovering Vulnerabilities in Hyperbolic Networks

角梯度符号法:揭示双曲网络中的漏洞

Minsoo Jo, Dongyoon Yang, Taesup Kim

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种基于双曲空间几何性质的对抗攻击方法,通过分解梯度为径向和角向分量,生成高影响的对抗样本,提升分类和检索任务的欺骗率。

Comments Accepted by AAAI 2026. Code available at: https://github.com/J-Minsoo/AGSM

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(7), 5566-5574, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20513 2026-03-24 cs.IR cs.AI 57%

ReBOL: Retrieval via Bayesian Optimization with Batched LLM Relevance Observations and Query Reformulation

ReBOL:通过批量LLM相关性观察和查询重述进行检索

Anton Korikov, Scott Sanner

机构 * University of Toronto(多伦多大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 ReBOL通过引入多模态贝叶斯优化和查询重述技术,改进检索阶段的召回率和排名质量,优于现有LLM重排序基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20336 2026-03-24 cs.IR cs.AI cs.DB 57%

GEM: A Native Graph-based Index for Multi-Vector Retrieval

GEM:一种基于图的多向量检索索引

Yao Tian, Zhoujin Tian, Xi Zhao, Ruiyuan Zhang, Xiaofang Zhou

机构 * The Hong Kong University of \ Hong Kong Generative AI Research \& Development Center Hong Kong China The Hong Kong University of Science Hong Kong Generative AI Research \& Development Center

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

AI总结 GEM提出一种针对多向量表示的原生索引框架,通过构建向量集的接近图,保留细粒度语义并实现高效导航,提升检索效率和准确性。

Comments This paper has been accepted by SIGMOD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏