arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-24 至 2026-02-24 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 7 篇

2602.19212 2026-02-24 cs.CL 83%

Retrieval Augmented Enhanced Dual Co-Attention Framework for Target Aware Multimodal Bengali Hateful Meme Detection

基于检索增强的增强双注意力框架的目标感知多模态孟加拉语仇恨表情识别

Raihan Tanvir, Md. Golam Rabiul Alam

机构 * Department of Computer Science and Engineering, BRAC University(计算机科学与工程系,BRAC大学) Department of Computer Science and Engineering, Ahsanullah University of Science and Technology(计算机科学与工程系,Ahsanullah科技大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文提出基于检索增强的增强双注意力框架,用于目标感知多模态孟加拉语仇恨表情识别,通过语义对齐提升类别平衡和多样性,实验表明xDORA和RAG-Fused DORA在识别和检测任务中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19091 2026-02-24 cs.CV 83%

CREM: Compression-Driven Representation Enhancement for Multimodal Retrieval and Comprehension

CREM: 为多模态检索与理解的压缩驱动表示增强

Lihao Liu, Yan Wang, Biao Yang, Da Li, Jiangxia Cao, Yuxiao Luo, Xiang Chen, Xiangyu Wu, Wei Yuan, Fan Yang, Guiguang Ding, Tingting Gao, Guorui Zhou

机构 * Tsinghua University(清华大学) Kuaishou Technology(快手科技)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 CREM通过压缩驱动的表示增强方法,在提升多模态检索性能的同时保持生成能力,实现生成与嵌入任务的统一优化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19735 2026-02-24 cs.CV 79%

VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments

VGGT-MPR:基于VGGT的多模态地点识别在自动驾驶环境中的应用

Jingyi Xu, Zhangshuo Qi, Zhongmiao Yan, Xuyu Gao, Qianyun Jiao, Songpengcheng Xia, Xieyuanli Chen, Ling Pei

机构 * Shanghai Jiao Tong University(上海交通大学) Beijing Institute of Technology(北京理工大学) National University of Defense Technology(国防科技大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 VGGT-MPR通过视觉几何基础的Transformer实现多模态地点识别,在自动驾驶中表现出强鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19987 2026-02-24 cs.LG cs.IR 78%

Counterfactual Understanding via Retrieval-aware Multimodal Modeling for Time-to-Event Survival Prediction

基于检索感知多模态建模的反事实理解用于生存预测

Ha-Anh Hoang Nguyen, Tri-Duc Phan Le, Duc-Hoang Pham, Huy-Son Nguyen, Cam-Van Thi Nguyen, Duc-Trong Le, Hoang-Quynh Le

机构 * University of Engineering and Technology(工程大学) Hanoi Vietnam National University(河内越南国家大学) Delft University of Technology(代尔夫特理工大学)

专题命中 跨模态检索 :multimodal(title,abstract)

AI总结 CURE通过多模态建模和潜在子组检索提升生存预测,优于现有基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03098 2026-02-24 cs.LG cs.AI 70%

TextME: Bridging Unseen Modalities Through Text Descriptions

TextME: 通过文本描述弥合未见模态

Soyeon Hong, Jinchan Kim, Jaegook You, Seungtaek Choi, Suha Kwak, Hyunsouk Cho

机构 * Department of Artificial Intelligence, Ajou University, Suwon, South Korea Division of Language \& AI, Hankuk University of Foreign Studies, Seoul, Korea Graduate School of AI, POSTECH, Pohang, Korea Department of Software, Ajou University, Suwon, South Korea

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 TextME通过仅使用文本描述实现跨模态扩展,无需配对监督,有效弥合不同模态间的差距。

Comments Code available at https://github.com/SoyeonHH/TextME

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00523 2026-02-24 cs.AI cs.CV 62%

VIRTUE: Visual-Interactive Text-Image Universal Embedder

VIRTUE: 视觉-交互文本-图像通用嵌入器

Wei-Yao Wang, Kazuya Tateishi, Qiyu Wu, Shusuke Takahashi, Yuki Mitsufuji

机构 * Sony Group Corporation(索尼集团有限公司) Sony AI(索尼人工智能)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 VIRTUE通过引入视觉交互能力,提升图像和文本的联合表示学习,实现更精确的实体级信息处理和应用拓展。

Comments ICLR 2026. 25 pages. Project page: https://sony.github.io/virtue/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18476 2026-02-24 q-bio.BM cs.AI cs.LG 57%

BioLM-Score: Language-Prior Conditioned Probabilistic Geometric Potentials for Protein-Ligand Scoring

BioLM-Score:基于语言先验的概率几何势用于蛋白质-配体评分

Zhangfan Yang, Baoyun Chen, Dong Xu, Jia Wang, Ruibin Bai, Junkai Ji, Zexuan Zhu

机构 * School of Computer Science, University of Nottingham Ningbo(计算机科学学院,诺丁汉大学宁波分校) School of Artificial Intelligence, Shenzhen University(人工智能学院,深圳大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 BioLM-Score结合几何建模与表征学习,提供一种高效、可泛化且可解释的蛋白质-配体评分方法,提升药物发现效率。

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏