arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-04-17 至 2026-04-17 共收录 10 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 10 篇

2604.15233 2026-04-17 cs.AI cs.DB 83%

Blue Data Intelligence Layer: Streaming Data and Agents for Multi-source Multi-modal Data-Centric Applications

蓝色数据智能层:面向多源多模态数据导向应用的流数据与智能体

Moin Aminnaseri, Farima Fatahi Bayat, Nikita Bhutani, Jean-Flavien Bussotti, Kevin Chan, Rafael Li Chen, Yanlin Feng, Jackson Hassell, Estevam Hruschka, Eser Kandogan, Hannah Kim, James Levine, Seiji Maekawa, Jalal Mahmud, Kushan Mitra, Naoki Otani, Pouya Pezeshkpour, Nima Shahbazi, Chen Shen, Dan Zhang

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出Blue数据智能层,通过整合异构数据源、多模态信息及上下文数据,解决多源多模态数据导向应用中的复杂查询需求,结合智能体与数据处理实现自然语言到SQL的扩展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15210 2026-04-17 cs.AI cs.CL 81%

Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding

学习像卡通标题作家一样思考:用于多模态幽默理解的不一致-解决监督

Hatice Merve Vural, Doga Kukul, Ege Erdem Ozlu, Demir Ekin Arikan, Bob Mankoff, Erkut Erdem, Aykut Erdem

机构 * Koç University(科克大学) KUIS AI Center(KUIS人工智能中心) Air Mail and Cartoon Collections(Air Mail和卡通收藏) Hacettepe University(哈恰特佩大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出IRS框架,通过分解幽默理解为不一致建模、解决建模和偏好对齐,提升多模态幽默理解能力,在NYCC基准上优于现有模型,且在零样本迁移中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15946 2026-04-17 cs.LG cs.AI cs.CR 77%

Fall into a Pit, Gain in a Wit: Cognitive-Guided Harmful Meme Detection via Misjudgment Risk Pattern Retrieval

跌入陷阱,获得智慧:通过误判风险模式检索的认知引导有害迷因检测

Wenshuo Wang, Ziyou Jiang, Junjie Wang, Mingyang Li, Jie Huang, Yuekai Huang, Zhiyuan Chang, Feiyan Duan, Qing Wang

机构 * State Key Laboratory of Complex System Modeling and Simulation Technology(复杂系统建模与仿真技术国家重点实验室) Science and Technology on Integrated Information System Laboratory(集成信息系统技术研究所) Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 跨模态检索 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.AI

AI总结 本文提出PatMD方法,通过学习并主动缓解潜在误判风险,识别有害迷因的深层误判风险模式,提升多模态大语言模型的检测能力,实验显示在5项有害检测任务中,F1-score和准确率均显著提升。

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14204 2026-04-17 cs.SD cs.AI eess.AS 73%

Disentangled Dual-Branch Graph Learning for Conversational Emotion Recognition

解耦双分支图学习用于对话情感识别

Chengling Guo, Yuntao Shou, Tao Meng, Wei Ai, Yun Tan, Keqin Li

机构 * College of Computer and Mathematics(计算机与数学学院) Central South University of Forestry and Technology(林业科技大学) Changsha Hospital for Maternal and Child Health Care(长沙妇幼保健医院) Department of Computer Science, State University of New York(纽约州立大学计算机科学系)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI、eess.AS

AI总结 本文提出解耦双分支图学习框架,通过分离模态不变和特定特征,结合傅里叶图神经网络和说话人感知超图,提升对话情感识别性能。

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14710 2026-04-17 cs.CV 70%

G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval

G-MIXER:基于测地混合的隐式语义扩展和显式语义重新排序用于零样本复合图像检索

Jiyoung Lim, Heejae Yang, Jee-Hyong Lee

机构 * Sungkyunkwan University(顺天大学)

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 G-MIXER通过测地混合生成隐式语义特征并重新排序显式语义,提升零样本复合图像检索的多样性和准确性,实现跨多个基准的最优性能。

Comments CVPR 2026 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10166 2026-04-17 cs.MM 57%

Fact-Checking with Contextual Narratives: Leveraging Retrieval-Augmented LLMs for Social Media Analysis

基于上下文叙述的事实核查:利用检索增强的LLM进行社交媒体分析

Arka Ujjal Dey, Muhammad Junaid Awan, Georgia Channing, Christian Schroeder de Witt, John Collomosse

专题命中 跨模态检索 :multimodal(abstract);分类 cs.MM

AI总结 CRAVE框架整合检索增强的大语言模型与聚类技术,通过多模态证据检索和聚类分析,提升社交媒体事实核查的准确性和解释性。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10410 2026-04-17 cs.AI 57%

CWCD: Category-Wise Contrastive Decoding for Structured Medical Report Generation

CWCD:用于结构化医学报告生成的类别级对比解码

Shantam Srivastava, Mahesh Bhosale, David Doermann, Mingchen Gao

机构 * The Department of Computer Science and Engineering(计算机科学与工程系) University at Buffalo, The State University of New York, NY, USA(布法罗大学,纽约州立大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出CWCD框架,通过类别特定参数化和对比正常与遮蔽X光片生成结构化放射报告,提升临床效果和生成质量。

Comments Accepted to MIDL 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08952 2026-04-17 cs.CL cs.IR 57%

MAB-DQA: Addressing Query Aspect Importance in Document Question Answering with Multi-Armed Bandits

MAB-DQA: 通过多臂老虎机解决文档问答中的查询方面重要性

Yixin Xiang, Yunshan Ma, Xiaoyu Du, Yibing Chen, Yanxin Zhang, Jinhui Tang

机构 * Nanjing University of Science and Technology(南京理工大学) Singapore Management University(新加坡国立大学) Nanjing Pami Intelligent Technology Co., Ltd.(南京帕米智能科技有限公司) University of Wisconsin - Madison(威斯康星大学麦迪逊分校) Nanjing Forestry University(南京林业大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

AI总结 本文提出MAB-DQA框架,通过多臂老虎机模型解决文档问答中查询方面的重要性问题,通过分解查询为方面感知的子查询并动态分配检索预算,提升文档理解性能。

Comments Accepted by ACL 2026. 20 pages, 9 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21262 2026-04-17 cs.CL 57%

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding

CausalEmbed: 在潜在空间中实现自动回归多向量生成用于视觉文档嵌入

Jiahao Huo, Yu Huang, Yibo Yan, Ye Pan, Kening Zheng, Wei-Chieh Huang, Yi Cao, Mingdong Ou, Philip S. Yu, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Alibaba Cloud Computing(阿里云计算) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

AI总结 本文提出CausalEmbed方法,通过对比训练中的迭代边际损失生成紧凑且结构良好的多向量嵌入,实现高效的视觉文档检索,减少30-155倍的token数量同时保持高性能。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13168 2026-04-17 cs.AI cs.CE cs.IR cs.MA 57%

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows

Finch:跨电子表格中心企业工作流的金融与会计基准测试

Haoyu Dong, Pengkun Zhang, Yan Gao, Xuanyu Dong, Yilin Cheng, Mingzhe Lu, Zikun Zhu, Adina Yakefu, Shuxin Zheng

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 Finch通过真实企业工作流评估AI代理,包含数据录入、结构化、格式化等任务,涵盖预算、交易等多领域,展示了AI在复杂企业工作流中的挑战。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏