arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-03 至 2026-03-03 共收录 23 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 23 篇

2602.20089 2026-03-03 cs.CV cs.AI 84%

StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues

StructXLIP: 通过多模态结构线索增强视觉-语言模型

Zanxi Ruan, Songqun Gao, Qiuyu Kong, Yiming Wang, Marco Cristani

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 StructXLIP通过引入多模态结构线索提升视觉-语言模型的跨模态检索性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01416 2026-03-03 cs.AI 83%

Securing the Floor and Raising the Ceiling: A Merging-based Paradigm for Multi-modal Search Agents

保障地板,提升天花板:一种基于合并的多模态搜索代理范式

Zhixiang Wang, Jingxuan Xu, Dajun Chen, Yunfang Wu, Wei Jiang, Yong Li

机构 * Peking University, Beijing, China(北京大学)

专题命中 图文多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出一种无需训练的多模态搜索代理范式,通过跨模态模型合并提升VLMs的自主搜索能力,OBM算法在搜索任务中表现出更高的效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01783 2026-03-03 cs.CV 83%

Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-Spoofing

在多模态大语言模型中利用链式推理进行人脸反伪装

Honglu Zhang, Zhiqin Fang, Ningning Zhao, Saihui Hou, Long Ma, Renwang Pei, Zhaofeng He

机构 * Didi Chuxing(滴滴出行) Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Normal University(北京师范大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出FaceCoT数据集和CEPL策略,通过链式推理提升多模态大语言模型在人脸反伪装中的鲁棒性和性能。

Comments Accepted to CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01816 2026-03-03 cs.MM 79%

Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health Understanding

声音、面孔与情感:多模态情绪-认知描述用于心理健康理解

Zhiyuan Zhou, Yanrong Guo, Shijie Hao

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.MM

AI总结 ECMC通过多模态数据生成情绪-认知描述,提升心理健康评估的准确性和可解释性。

Comments Accepted at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01055 2026-03-03 cs.AI 79%

MMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Contextual Reasoning

MMCOMET:一种大规模多模态常识知识图谱用于上下文推理

Eileen Wang, Hiba Arnaout, Dhita Pratama, Shuo Yang, Dangyang Liu, Jie Yang, Josiah Poon, Jeff Pan, Caren Han

机构 * University of Sydney(悉尼大学) University of Melbourne(墨尔本大学) University of Edinburgh(爱丁堡大学) The University of Sydney(悉尼大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 MMCOMET是一种大规模多模态常识知识图谱,通过整合视觉、物理和社会知识,提升了复杂推理任务如图像描述和叙事生成的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00694 2026-03-03 cs.RO cs.AI 79%

Wild-Drive: Off-Road Scene Captioning and Path Planning via Robust Multi-modal Routing and Efficient Large Language Model

Wild-Drive: 通过鲁棒多模态路由和高效大语言模型实现越野场景描述与路径规划

Zihang Wang, Xu Li, Benwu Wang, Wenkai Zhu, Xieyuanli Chen, Dong Kong, Kailin Lyu, Yinan Du, Yiming Peng, Haoyang Che

机构 * School of Instrument Science and Engineering, Southeast University(东南大学仪器科学与工程学院) Southeast University Nanjing Jiangbei New Area Innovation Research Institute(东南大学南京江滨新区创新研究院) National Key Laboratory of Equipment State Sensing and Smart Support, National University of Defense Technology(国防科技大学装备状态感知与智能支撑国家重点实验室) School of Transportation, Shandong University of Science and Technology(山东科技大学交通学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

AI总结 Wild-Drive通过鲁棒多模态路由和高效大语言模型实现越野场景描述与路径规划,提升复杂环境下的可解释性和稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06223 2026-03-03 cs.LG cs.CV stat.ML 79%

Beyond DAGs: A Latent Partial Causal Model for Multimodal Learning

超越DAGs:一种用于多模态学习的潜在部分因果模型

Yuhang Liu, Zhen Zhang, Dong Gong, Erdun Gao, Biwei Huang, Mingming Gong, Anton van den Hengel, Kun Zhang, Javen Qinfeng Shi

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种用于多模态学习的潜在部分因果模型,通过解耦表示提升模型在少量样本学习和领域泛化中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05606 2026-03-03 cs.CV cs.CL 79%

Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision

Uni-cot: 向跨文本和视觉的统一链式推理迈进

Luozheng Qin, Jia Gong, Yuqing Sun, Tianjiao Li, Mengping Yang, Xiaomeng Yang, Chao Qu, Zhiyu Tan, Hao Li

机构 * Shanghai Academy of AI for Science(上海人工智能科学研究院) Fudan University(复旦大学) Nanyang Technological University(南洋理工大学)

专题命中 图文多模态 :multimodal(abstract);multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

AI总结 Uni-CoT通过统一的链式推理框架实现跨文本和视觉的连贯多模态推理,采用宏级和微级推理范式,提升多模态推理的效率和性能。

Comments Accepted by ICLR 2026, Project Page: https://sais-fuxi.github.io/projects/uni-cot/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00854 2026-03-03 cs.LG cs.IR 78%

GeMi: A Graph-based, Multimodal Recommendation System for Narrative Scroll Paintings

GeMi:基于图的多模态推荐系统用于叙事卷轴绘画

Haimonti Dutta, Pruthvi Moluguri, Jin Dai, Saurabh Amarnath Mahindre

专题命中 图文多模态 :multimodal(title,abstract)

AI总结 GeMi是一种基于图的多模态推荐系统,专门用于叙事卷轴绘画,结合多模态内容和用户偏好,为艺术保护和数据存储提供解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10067 2026-03-03 cs.CV 77%

FiLo++: Zero-/Few-Shot Anomaly Detection by Fused Fine-Grained Descriptions and Deformable Localization

FiLo++: 通过融合细粒度描述和变形定位实现零/少样本异常检测

Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(基础模型研究中心,自动化研究所,中国科学院) School of Artifcial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Wuhan AI Research(武汉AI研究所) Peng Cheng Laboratory(鹏城实验室) Guangdong Provincial Key Laboratory of Intellectual Property & Big Data, Guangdong Polytechnic Normal University(广东省知识产权与大数据重点实验室,广东工业大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 FiLo++通过融合细粒度描述和变形定位技术,提升零/少样本异常检测的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17132 2026-03-03 cs.CV cs.CL 73%

Dynamic Token Reweighting for Robust Vision-Language Models

动态令牌重加权用于鲁棒的视觉-语言模型

Tanqiu Jiang, Jiacheng Liang, Rongyi Zhu, Jiawei Zhou, Fenglong Ma, Ting Wang

机构 * Stony Brook University(石溪大学) Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 图文多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL

AI总结 DTR通过优化KV缓存动态调整视觉令牌权重,提升视觉-语言模型对多模态劫持攻击的鲁棒性,同时保持模型性能。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16612 2026-03-03 cs.CL cs.AI cs.CV 67%

MemeIntel: Explainable Detection of Propagandistic and Hateful Memes

MemeIntel: 可解释性地检测宣传性及仇恨言论的迷因

Mohamed Bayan Kmainasi, Abul Hasnat, Md Arid Hasan, Ali Ezzat Shahroor, Firoj Alam

机构 * Qatar Computing Research Institute(卡塔尔计算研究 institute)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 MemeIntel通过多阶段优化方法和视觉-语言模型,提升了阿拉伯语宣传性迷因和英语仇恨言论的检测与解释生成性能,准确率提升显著。

Comments disinformation, misinformation, factuality, harmfulness, fake news, propaganda, hateful meme, multimodality, text, images

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01124 2026-03-03 cs.CV cs.AI 62%

ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models

ClinCoT:面向医学视觉语言模型的临床感知视觉推理链

Xiwei Liu, Yulong Li, Xinlin Zhuang, Xuhui Li, Jianxu Chen, Haolin Yang, Imran Razzak, Yutong Xie

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) East China Normal University(东华大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 ClinCoT通过引入临床感知的视觉推理链框架,提升医学视觉语言模型在临床决策中的事实性支撑与多模态推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00529 2026-03-03 cs.CV cs.AI 62%

CaptionFool: Universal Image Captioning Model Attacks

CaptionFool: 针对最新Transformer图像描述模型的通用图像描述模型攻击

Swapnil Parekh

机构 * Intuit

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

AI总结 CaptionFool通过修改少量图像块,成功生成任意目标描述,揭示了视觉-语言模型的安全漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06292 2026-03-03 cs.CV cs.AI 62%

ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations

ChainMPQ: 交错文本图像推理链用于缓解关系幻觉

Yike Wu, Yiwei Wang, Yujun Cai

机构 * University of Queensland(昆士兰大学) University of California, Merced(加州大学梅尔德分校) Ant Group(蚂蚁集团)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 ChainMPQ通过交错文本和图像推理链,利用积累的记忆减少大型视觉语言模型的关系幻觉,提升关系推理能力。

Comments Accepted by ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12678 2026-03-03 cs.CV 57%

$β$-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment

$β$-CLIP:多粒度视觉-语言对齐的文本条件对比学习

Fatimah Zohra, Chen Zhao, Hani Itani, Bernard Ghanem

机构 * King Abdullah University of Science and Technology (KAUST)(卡布斯大学科学与技术研究院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

AI总结 $β$-CLIP通过多粒度文本条件对比学习提升视觉-语言对齐,实现从完整描述到短语的层次化对齐,改进细粒度和长文本检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08919 2026-03-03 cs.CV cs.LG 57%

PHyCLIP: $\ell_1$-Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation Learning

PHyCLIP通过$\ell_1$-Product度量结合双曲因子,统一了视觉-语言表示学习中的层次结构与组合性

Daiki Yoshikawa, Takashi Matsubara

机构 * Hokkaido University(北海道大学) CyberAgent

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

AI总结 PHyCLIP通过$\ell_1$-Product度量结合双曲因子,统一了视觉-语言表示学习中的层次结构与组合性。

Comments 24 pages. Codes are available at https://github.com/tksmatsubara/PHyCLIP

Journal ref International Conference on Learning Representations (ICLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01029 2026-03-03 cs.CV 57%

Vision-Language Feature Alignment for Road Anomaly Segmentation

视觉-语言特征对齐用于道路异常分割

Zhuolin He, Jiacheng Tang, Jian Pu, Xiangyang Xue

机构 * School of Computer Science, Fudan University(复旦大学计算机科学学院) Institute of Science and Technology for Brain-Inspired Intelligence, Fudan University(复旦大学脑启发智能科学与技术研究院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

AI总结 VL-Anomaly通过结合视觉-语言模型的语义先验,提出了一种新的道路异常分割框架,有效提升异常检测的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00842 2026-03-03 cs.CL 57%

MedGPT-oss: Training a General-Purpose Vision-Language Model for Biomedicine

MedGPT-oss: 为生物医学训练一个通用的视觉-语言模型

Kai Zhang, Zhengqing Yuan, Cheng Peng, Songlin Zhao, Mengxian Lyu, Ziyi Chen, Yanfang Ye, Wei Liu, Ying Zhang, Kaleb E Smith, Lifang He, Lichao Sun, Yonghui Wu

机构 * Department of Computer Science and Engineering, Lehigh University(莱斯大学计算机科学与工程系) Department of Computer Science and Engineering, University of Notre Dame(圣母大学计算机科学与工程系) Department of Health Outcomes & Biomedical Informatics, University of Florida(佛罗里达大学健康结果与生物医学信息学系) Department of Radiation Oncology, Mayo Clinic(梅奥诊所放射肿瘤科) Research Computing, University of Florida(佛罗里达大学研究计算中心) AI Technology Center, NVIDIA(NVIDIA人工智能技术中心)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

AI总结 MEDGPT-oss是一款开放式权重的200亿参数视觉-语言模型,通过优化训练课程实现多模态和文本推理能力,适用于隐私保护的临床AI研究。

Comments Technical report, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23607 2026-03-03 cs.CV 57%

Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations

Concerto:联合2D-3D自监督学习产生空间表示

Yujia Zhang, Xiaoyang Wu, Yixing Lao, Chengyao Wang, Zhuotao Tian, Naiyan Wang, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

AI总结 Concerto通过联合2D-3D自监督学习,产生更连贯的信息空间表示,优于现有SOTA模型并在多个基准上取得新成就。

Comments NeurIPS 2025, produced by Pointcept, project page: https://pointcept.github.io/Concerto

Journal ref Neural Information Processing Systems 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14362 2026-03-03 cs.CV 57%

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

DeepEyes: 通过强化学习激励'图像思考'

Ziwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao, Guohai Xu, Le Yang, Chao Shen, Xing Yu

机构 * Xiaohongshu Inc.(小红书公司) Xi’an Jiaotong University(西安交通大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 DeepEyes通过强化学习实现图像引导的推理,无需预训练数据,提升多模态任务性能。

Comments Accepted by ICLR2026. Ziwei, Michael, Jack, and Chenxiao are equal-contribution. The list order is random

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03566 2026-03-03 cs.CV cs.LG 57%

CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally

CLIP 表现得像一个词袋模型但不是单模态的

Darina Koishigarina, Arnas Uselis, Seong Joon Oh

机构 * Tübingen AI Center, University of Tübingen(图宾根人工智能中心,图宾根大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

AI总结 CLIP表现出词袋模型特性,但其属性-对象绑定信息已单模态编码,通过简单线性变换可提升性能。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01813 2026-03-03 cs.RO 50%

SSMG-Nav: Enhancing Lifelong Object Navigation with Semantic Skeleton Memory Graph

SSMG-Nav: 通过语义骨架记忆图增强终身物体导航

Haochen Niu, Lantao Zhang, Xingwu Ji, Rendong Ying, Peilin Liu, Fei Wen

机构 * Brain-Inspired Application Technology Center (BATC), Shanghai Jiao Tong University(脑启发应用技术中心(BATC),上海交通大学)

专题命中 图文多模态 :multimodal(abstract)

AI总结 SSMG-Nav通过语义骨架记忆图提升终身物体导航,整合多模态信息与持久记忆,优化路径效率与任务成功率。

Comments Accepted by 2026 ICRA

详情

展开后加载摘要…

URL PDF HTML 收藏