arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4865 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4865 篇

2507.15692 2026-07-02 cs.HC cs.CL cs.CV 交叉投稿 91%

Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions

表面化变异以校准MLLM生成图像描述的可信度感知

Meng Chen, Akhil Iyer, Amy Pavel

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of California, Berkeley(加州大学伯克利分校)

专题命中 其他多模态 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CV、cs.CL

AI总结 通过系统性地展示多个MLLM响应间的变异,帮助盲人或低视力用户无需视觉检查即可检测不可靠信息,实验表明该方法将识别不可靠声明的能力提升4.9倍。

Comments 18 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28546 2026-06-30 cs.LG 89%

NIVA: A Multimodal Foundation Model for Actionable Earth System Intelligence

NIVA:面向可操作地球系统智能的多模态基础模型

Anisha Pal, Aodhan Sweeney, Kyle Heyblom, Kalai Ramea

机构 * Planette AI

专题命中 其他多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract)

AI总结 提出多模态基础模型NIVA,通过联合学习海洋与大气模态,捕获耦合地球系统动力学,为次季节至季节预测提供基础,并准确预测主要气候指数。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16022 2026-05-18 cs.CV 89%

EndoGSim: Physics-Aware 4D Dynamic Endoscopic Scene Simulations via MLLM-Guided Gaussian Splatting

EndoGSim: 基于多模态大语言模型的物理感知4D动态内窥镜场景模拟

Changjing Liu, Yiming Huang, Long Bai, Beilei Cui, Hongliang Ren

机构 * Department of Electronic Engineering, The Chinese University of Hong Kong (CUHK)(香港中文大学电子工程系)

专题命中 其他多模态 :MLLM(title,summary_cn);multi-modal(abstract);分类 cs.CV

AI总结 本文提出EndoGSim框架,通过MLLM引导的高斯点散布实现内窥镜场景的物理感知重建与模拟,结合预训练分割和深度估计,提升手术模拟的真实性和准确性。

Comments Early Accepted by MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03229 2026-07-07 cs.DC 新提交 88%

HyperParallel-Mpipe: A Composable Algebra System for Optimizing MLLM Training over Supernode Clusters

HyperParallel-Mpipe:用于在超级节点集群上优化多模态大语言模型训练的可组合代数系统

Chong Li, Zhengdao Yu, Nelson Lossing, Thibaut Tachon, Pierre Leca, Etienne Filhol, Yujie Yuan, Chong Bao, Teng Su

专题命中 其他多模态 :MLLM(title,summary_cn);multimodal(abstract)

AI总结 分析多模态大语言模型(MLLM)训练痛点,引入Mpipe,用调度代数从紧凑调度规范推导运行时行为,得出多模态感知异构并行调度transpose,在Ascend 910C NPU集群上加速显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05575 2025-07-09 cs.CV 88%

Multi-Modal Face Anti-Spoofing via Cross-Modal Feature Transitions

Jun-Xiong Chong, Fang-Yu Hsu, Ming-Tsung Hsu, Yi-Ting Lin, Kai-Heng Chien, Chiou-Ting Hsu, Pei-Kai Huang

机构 * College of Computer and Cyber Security, Fujian Normal University(计算机与网络安全学院,福建师范大学) Department of Computer Science, National Tsing Hua University(计算机科学系,国立清华大学)

专题命中 其他多模态 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15727 2025-01-28 cs.HC cs.AI 88%

Gensors: Authoring Personalized Visual Sensors with Multimodal Foundation Models and Reasoning

Michael Xieyang Liu, Savvas Petridis, Vivian Tsai, Alexander J. Fiannaca, Alex Olwal, Michael Terry, Carrie J. Cai

专题命中 其他多模态 :multimodal(title,abstract);multimodal foundation model(title);MLLM(abstract);分类 cs.AI

Journal ref 30th International Conference on Intelligent User Interfaces (IUI'25), March 24-27, 2025, Cagliari, Italy. ACM, New York, NY, USA, 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.08073 2018-09-13 stat.ML cs.AI cs.LG q-bio.QM 88%

Cross-modal Recurrent Models for Weight Objective Prediction from Multimodal Time-series Data

Petar Veličković, Laurynas Karazija, Nicholas D. Lane, Sourav Bhattacharya, Edgar Liberis, Pietro Liò, Angela Chieh, Otmane Bellahsen, Matthieu Vegreville

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

Comments To appear in NIPS ML4H 2017 and NIPS TSW 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.10111 2018-07-31 cs.CV 88%

MRI to FDG-PET: Cross-Modal Synthesis Using 3D U-Net For Multi-Modal Alzheimer's Classification

Apoorva Sikka, Skand Vishwanath Peri, Deepti. R. Bathula

专题命中 其他多模态 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15828 2025-04-18 cs.LG 88%

Measuring Cross-Modal Interactions in Multimodal Models

Laura Wenderoth, Konstantin Hemker, Nikola Simidjievski, Mateja Jamnik

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(title,abstract)

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, Volume 39, February 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14504 2025-08-21 cs.CV cs.AI 87%

PB-IAD: Utilizing multimodal foundation models for semantic industrial anomaly detection in dynamic manufacturing environments

Bernd Hofmann, Albert Scheck, Joerg Franke, Patrick Bruendl

专题命中 其他多模态 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04201 2025-03-07 cs.CL cs.AI 87%

Knowledge-Decoupled Synergetic Learning: An MLLM based Collaborative Approach to Few-shot Multimodal Dialogue Intention Recognition

Bin Chen, Yu Zhang, Hongfei Ye, Ziyi Huang, Hongyang Chen

专题命中 其他多模态 :multimodal(title,abstract);MLLM(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.14378 2022-06-09 cs.AI 86%

Towards artificial general intelligence via a multimodal foundation model

Nanyi Fei, Zhiwu Lu, Yizhao Gao, Guoxing Yang, Yuqi Huo, Jingyuan Wen, Haoyu Lu, Ruihua Song, Xin Gao, Tao Xiang, Hao Sun, Ji-Rong Wen

专题命中 其他多模态 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.AI

Comments Published by Nature Communications, see https://www.nature.com/articles/s41467-022-30761-2

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.02790 2020-09-22 eess.IV cs.CV 86%

Adversarial Uni- and Multi-modal Stream Networks for Multimodal Image Registration

Zhe Xu, Jie Luo, Jiangpeng Yan, Ritvik Pulya, Xiu Li, William Wells, Jayender Jagadeesan

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(title);分类 cs.CV

Comments accepted by MICCAI 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02871 2026-04-21 cs.CL cs.AI 86%

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning

位置:多模态大语言模型可以显著推动科学推理

Yibo Yan, Shen Wang, Jiahao Huo, Jingheng Ye, Zhendong Chu, Xuming Hu, Philip S. Yu, Carla Gomes, Bart Selman, Qingsong Wen

机构 * Squirrel AI HKUST(GZ)(香港科技大学(广州)) HKUST(香港科技大学) Tsinghua University(清华大学) University of Illinois at Chicago(伊利诺伊大学香槟分校) Cornell University(康奈尔大学)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本文探讨多模态大语言模型在科学推理中的应用,提出四阶段研究路线,指出当前模型在跨领域推理中的潜力与挑战,为实现通用人工智能提供新视角。

Comments Accepted by The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026, Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03113 2026-03-03 cs.CV cs.CL 86%

Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection

通过基于梯度的自我反思缓解多模态幻觉

Shan Wang, Maying Shen, Nadine Chang, Chuong Nguyen, Hongdong Li, Jose M. Alvarez

机构 * NVIDIA Australian National University(澳大利亚国立大学) Data61, CSIRO(Data61,CSIRO)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出GACD方法,通过梯度分析缓解多模态模型的幻觉问题,提升输出的视觉基础性。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17648 2024-07-09 cs.CV cs.AI 86%

Bridging Modality Gap for Visual Grounding with Effecitve Cross-modal Distillation

Jiaxi Wang, Wenhui Hu, Xueyang Liu, Beihu Wu, Yuting Qiu, YingYing Cai

专题命中 其他多模态 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14986 2025-06-19 cs.LG 86%

Early Prediction of Multiple Sclerosis Disability Progression via Multimodal Foundation Model Benchmarks

Maxime Usdin, Lito Kriara, Licinio Craveiro

机构 * Genentech, Inc(基因泰克公司) F. Hoffmann-La Roche Ltd(Hoffmann-La Roche公司)

专题命中 其他多模态 :multimodal(title,abstract);multimodal foundation model(title)

Comments Accepted to IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00111 2025-03-13 cs.LG cond-mat.mtrl-sci 86%

Multimodal Foundation Models for Material Property Prediction and Discovery

Viggo Moro, Charlotte Loh, Rumen Dangovski, Ali Ghorashi, Andrew Ma, Zhuo Chen, Samuel Kim, Peter Y. Lu, Thomas Christensen, Marin Soljačić

专题命中 其他多模态 :multimodal(title,abstract);multimodal foundation model(title)

Comments 12 pages, 4 figures

Journal ref Newton, Volume 1, Issue 1, 100016 (2025) Newton, Volume 1, Issue 1, 100016 (2025) Newton, Volume 1, Issue 1, 100016 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16586 2026-07-30 cs.CV 版本更新 85%

LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models

LOCUS: 局部视觉线索搜索增强多模态大语言模型的细粒度感知

Zhou Tao, Fang Zhang, Zewen Ding, Shida Wang, Xiaokun Sun, YongXiang Hua, Haoyu Cao, Linli Xu

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(summary_cn);分类 cs.CV

AI总结 提出LOCUS训练框架,通过可验证的局部线索搜索代理任务,使MLLM内化细粒度证据选择,提升定位敏感视觉理解而不改变推理接口。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02020 2026-07-03 cs.AI 新提交 85%

Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails

持续多模态学习中的隐藏遗忘:当准确率幸存但基础失效时

Qianyu Chen, Canran Xiao, Runxuan Tang

机构 * Nanyang Technological University(南洋理工大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.AI

AI总结 针对持续多模态学习中模型答案准确但证据使用方式改变的问题,提出无重放依赖约束框架RCL,通过冻结旧模型、反事实干预估计证据依赖并联合优化,有效降低隐藏遗忘。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20705 2026-04-23 cs.CV 85%

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models

SSL-R1: 多模态大语言模型的自监督视觉强化后训练

Jiahao Xie, Alessio Tonioni, Nathalie Rauschmayr, Federico Tombari, Bernt Schiele

机构 * Max Planck Institute for Informatics, SIC(马克斯·普朗克信息研究所,SIC) VIA Research Center(VIA研究中心) Google(谷歌)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 SSL-R1通过自监督学习生成可验证奖励,提升多模态大语言模型的推理能力,无需人工或外部监督,有效增强视觉理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20503 2026-04-20 cs.CL 85%

Protecting multimodal large language models against misleading visualizations

保护多模态大语言模型免受误导性可视化的影响

Jonathan Tonglet, Tinne Tuytelaars, Marie-Francine Moens, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, TU Darmstadt and National Research Center for Applied Cybersecurity ATHENE(普遍知识处理实验室(UKP实验室)、计算机科学系,德累斯顿技术大学及应用网络安全国家研究中心ATHENE) Department of Electrical Engineering, KU Leuven(电气工程系,鲁文大学) Department of Computer Science, KU Leuven(计算机科学系,鲁文大学)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CL

AI总结 研究发现多模态大语言模型在面对误导性可视化时回答准确性显著下降,提出两种有效方法提升其处理能力,同时保持非误导性可视化下的准确性。

Comments Preprint. Code and data available at https://github.com/UKPLab/arxiv2025-misleading-visualizations

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04800 2026-03-06 cs.CV 85%

MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models

MASQuant: 多模态大语言模型的模态感知平滑量化

Lulu Hu, Wenhu Xiao, Xin Chen, Xinhua Xu, Bowen Xu, Kun Li, Yongliang Tao

机构 * Alibaba Cloud Computing, Alibaba Group(阿里巴巴云 computing,阿里巴巴集团)

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 MASQuant通过模态感知平滑和跨模态补偿技术,解决多模态大语言模型的量化问题,实现稳定且高效的量化性能。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16863 2025-04-03 cs.CV cs.AI cs.CL cs.MM 85%

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

Federico Cocchi, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10663 2025-03-12 cs.CV cs.LG 85%

Cross-Modal Few-Shot Learning: a Generative Transfer Learning Framework

Zhengwei Yang, Yuke Li, Qiang Sun, Basura Fernando, Heng Huang, Zheng Wang

专题命中 其他多模态 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments 15 pages, 9 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05240 2025-02-13 cs.CV 85%

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM

Yueying Zou, Peipei Li, Zekun Li, Huaibo Huang, Xing Cui, Xuannan Liu, Chenghanyu Zhang, Ran He

专题命中 其他多模态 :MLLM(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00056 2023-03-08 cs.LG cs.AI cs.CL cs.CV cs.MM 85%

MultiViz: Towards Visualizing and Understanding Multimodal Models

Paul Pu Liang, Yiwei Lyu, Gunjan Chhablani, Nihal Jain, Zihao Deng, Xingbo Wang, Louis-Philippe Morency, Ruslan Salakhutdinov

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICLR 2023. Code available at: https://github.com/pliang279/MultiViz

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12724 2026-08-14 cs.LG 新提交 85%

MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

MAG:流形引导的半监督多模态上下文学习

Zirui Cheng, Xun Xu, Tiankai Chen, Fady Rezk, Bowen Zheng, Xiaodong Shi, Shijie Li, Kangkang Lu, Bharadwaj Veeravalli, Nancy F. Chen

专题命中 其他多模态 :multi-modal(title,abstract);MLLM(abstract,abstract_cn)

AI总结 本研究提出MAG框架,通过将未标记多模态数据用于半监督传播的演示选择,提升多模态大语言模型的上下文学习性能,在8个基准的标签稀缺场景下优于强基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25451 2026-07-08 cs.LG 版本更新 85%

BigMac: Breaking the Pareto Frontier of Compute and Memory in Multimodal LLM Training

BigMac: 打破多模态大语言模型训练中的计算与内存帕累托前沿

Zili Zhang, Chengxu Yang, Shenglong Zhang, Chenyu Wang, Yufan Zhang, Tuo Dai, Zhouyang Li, Yuhong Ge, Chao Jin, Xin Jin, Yuliang Liu

机构 * Peking University(北京大学) Independent Researcher(独立研究员) Xiaohongshu, Inc(小红书公司)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn)

AI总结 提出BigMac训练流水线,通过嵌套编码器和生成器计算到LLM流水线中,同时优化计算效率和内存使用,打破帕累托前沿。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08962 2026-05-12 cs.DC 85%

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production

MegaScale-Omni: 一种面向多模态大语言模型生产训练的超大规模、工作负载容错系统

Chunyu Xue, Yangrui Chen, Jianyu Jiang, Ningxin Zheng, Junda Feng, Jingji Chen, Shixiong Zhao, Shen Yan, Yi Lin, Lei Shi, Zanbo Wang, Lishu Luo, Faming Wu, Haibin Lin, Xin Liu, Yanghua Peng, Quan Chen

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn)

AI总结 本文提出MegaScale-Omni系统,通过解耦并行策略、统一表示和负载平衡技术,提升多模态大语言模型在动态工作负载下的训练效率,实现1.27倍至7.57倍的吞吐量提升。

详情

展开后加载摘要…

URL PDF HTML 收藏