arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4868 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4868 篇

2606.22565 2026-06-23 cs.CL cs.AI cs.CV 新提交 82%

Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do

轻看,重思:多模态思维链推理能做什么和不能做什么

Zhuoran Jin, Kejian Zhu, Hongbang Yuan, Yupu Hao, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao

机构 * The Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所复杂系统认知与决策智能实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文系统探究多模态思维链推理在12个感知与推理任务上的表现,发现CoT对感知任务有副作用,但对数学、科学等多图像推理有效,且现有开源多模态推理模型提升有限,视觉推理仍是瓶颈。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02153 2026-06-15 math.NA cs.NA math.OC 版本更新 82%

A Joint Variational Framework for Multimodal X-ray Ptychography and Fluorescence Reconstruction

多模态X射线ptychography与荧光重建的联合变分框架

Chengru Eric Zou, Elle Buser, Zichao Wendy Di, Yuanzhe Xi

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文提出一种联合变分框架,整合X射线ptychography和荧光重建,通过共享空间变量解决非线性逆问题,提升稳定性和重建精度。

Comments Keywords: inverse problems, x-ray imaging science, ill-posedness, joint reconstruction. This work is sponsored by NSF DMS-2338904

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09301 2026-06-09 cs.LG 新提交 82%

PRISM: Topology-Aware Cross-Modal Imputation for Modality-Deficient Federated Graph Learning

PRISM: 面向模态缺失联邦图学习的拓扑感知跨模态插补

Zekai Chen, Miao Zhang, Jiayang Xing, Xunkai Li, Xun Wu, Rong-Hua Li, Guoren Wang

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 其他多模态 :cross-modal(title,abstract);multimodal(abstract)

AI总结 针对联邦图学习中客户端级模态缺失问题,提出拓扑感知跨模态插补框架PRISM,通过联邦检索缺失模态语义并利用拓扑控制注入局部图传播,在六个多模态图数据集上平均提升4.48%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02842 2026-06-03 cs.LG 82%

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning

光谱渐进式思维流:轻量级多模态推理

Yixian Shen, Zhiheng Yang, Qi Bi, Changshuo Wang, Shuai Wang, Jia-Hong Huang, George Floros, Prayag Tiwari, Anuj Pathania

机构 * Informatics Institute, University of Amsterdam, Amsterdam, The Netherlands(阿姆斯特丹大学信息学院) Department of Computer Science, University College London(伦敦大学学院计算机科学系) University of Thessaly, Volos, Greece(塞萨洛尼基大学) Department of Electronic and Electrical Engineering, Trinity College Dublin, Dublin, Ireland(都柏林信任学院电子与电气工程系) School of Information Technology, Halmstad University, Halmstad, Sweden(哈姆斯塔德大学信息科技学院) Amazon AGI, Seattle, USA(亚马逊人工智能研究部)

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract)

AI总结 提出光谱渐进式思维流(SpecFlow),通过在固定大小离散余弦空间中表示中间视觉思维,并利用无分类器引导将视觉状态更新与文本意图对齐,实现轻量级多模态空间推理,在保持竞争性能的同时将计算和KV缓存成本降低高达2.1倍。

Comments Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27918 2026-05-28 cs.DC 82%

Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain

解决分布式多模态训练中的变量异构性问题:Entrain

Insu Jang, Mosharaf Chowdhury

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract)

AI总结 提出Entrain框架,通过宏观批次的静态模型并行配置和层次化微批次分配算法,解决多模态大语言模型训练中的数据异构性和变量波动,提升吞吐量达1.40倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22726 2026-05-22 eess.SY cs.SY 82%

Dynamic Lane Allocation in UAM Corridors for Efficient Multimodal Door-to-Door Mobility

动态车道分配用于高效多模式门到门移动

Jung Ho Park, Jordan Kam, Vishwanath Bulusu, Alexandre Bayen, Raja Sengupta

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract)

AI总结 本文提出了一种动态方向车道分配方法,用于城市空中交通(UAM)走廊,通过离散时间混合整数线性规划(MILP)模型来动态激活、停用和反转车道方向,以应对双向空域需求的变化。研究通过分解每个行程为多模式序列,并通过垂直机场侧调度模型路由UAM服务的中段,利用旧金山湾区作为案例研究,发现动态策略可减少未使用空域容量5倍,提高车道利用率至67%,并减少通勤人口的平均出行时间高达21.6%。

Comments Submitted to AIAA Aviation Forum

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08539 2026-04-21 cs.CV cs.AI cs.CL 82%

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks

OpenVLThinkerV2: 一个多领域视觉任务的通用多模态推理模型

Wenbo Hu, Xin Chen, Yan Gao-Tian, Yihe Deng, Nanyun Peng, Kai-Wei Chang

机构 * University of California, Los Angeles (UCLA)(加州大学洛杉矶分校)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出G$^2$RPO训练目标,通过非线性分布匹配解决多视觉任务中的奖励拓扑差异和感知与推理平衡问题,构建了鲁棒的OpenVLThinkerV2模型,在18个基准测试中表现优异。

Comments code at: https://github.com/uclanlp/openvlthinker

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12271 2026-04-15 cs.LG 82%

RoleMAG: Learning Neighbor Roles in Multimodal Graphs

RoleMAG: 在多模态图中学习邻居角色

Yilong Zuo, Xunkai Li, Zhihan Zhang, Ronghua Li, Guoren Wang

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract)

AI总结 RoleMAG通过区分邻居在传播中的角色,改进多模态图的处理,实验表明其在RedditS和Bili_Dance上表现最佳,且在Toys上保持竞争力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19067 2026-03-20 cs.LG eess.SP 82%

Communication-Efficient and Robust Multi-Modal Federated Learning via Latent-Space Consensus

通过潜在空间共识实现高效的多模态联邦学习

Mohamed Badi, Chaouki Ben Issaid, Mehdi Bennis

机构 * Center for Wireless Communications, University of Oulu(奥卢大学无线通信中心)

专题命中 其他多模态 :multi-modal(title,abstract);cross-modal(abstract)

AI总结 本文提出CoMFed框架,通过可学习的投影矩阵生成压缩的潜在表示,并利用潜在空间正则化器对齐客户端间的表示,提升跨模态一致性和鲁棒性,实验显示在人体活动识别基准上具有竞争力的准确率与低开销。

Comments Accepted for publication in IEEE Wireless Communications Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22171 2026-03-03 cs.HC 82%

A Taxonomy of Human--MLLM Interaction in Early-Stage Sketch-Based Design Ideation

早期阶段基于草图的设计构想中人类与大语言模型交互的分类

Weiyan Shi, Kenny Tsu Wei Choo

专题命中 其他多模态 :MLLM(title,abstract);multimodal(abstract)

AI总结 本文提出了一种分类方法,用于描述人类与大语言模型在早期阶段基于草图的设计构想中的交互模式,揭示了人类与AI角色的动态变化。

Comments Accepted at CHI 2026 Posters

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.04903 2026-02-06 eess.IV cs.LG 82%

Normative Modeling using Multimodal Variational Autoencoders to Identify Abnormal Brain Structural Patterns in Alzheimer Disease

基于多模态变分自编码器的规范建模用于识别阿尔茨海默病异常脑结构模式

Sayantan Kumar, Philip Payne, Aristeidis Sotiras

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract)

AI总结 本文提出基于多模态变分自编码器的规范建模框架,用于识别阿尔茨海默病中异常脑结构模式,通过联合分布建模提高疾病阶段检测的敏感性和准确性。

Comments Medical Imaging Meets NeurIPS workshop in NeurIPS 2022

Journal ref Proc. SPIE 12465, Medical Imaging 2023: Computer-Aided Diagnosis, 1246503 (7 April 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11304 2026-01-23 cs.AI cs.CL cs.CV 82%

Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring

利用实例分割辅助的多模态大语言模型进行智能交通监控

Murat Arda Onsu, Poonam Lohan, Burak Kantarci, Aisha Syed, Matthew Andrews, Sean Kennedy

机构 * University of Ottawa(渥太华大学) Nokia Bell Labs(诺基亚贝尔实验室)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文利用多模态大语言模型和实例分割技术,实现高准确率的交通监控系统,提升交通管理效率和安全性。

Comments 6 pages, 7 figures, submitted to 30th IEEE International Symposium on Computers and Communications (ISCC) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03181 2026-01-07 cs.NI cs.AI cs.CL cs.CV 82%

Multi-Modal Data-Enhanced Foundation Models for Prediction and Control in Wireless Networks: A Survey

多模态数据增强的基础模型用于无线网络中的预测与控制:综述

Han Zhang, Mohammad Farzanullah, Mohammad Ghassemi, Akram Bin Sediq, Ali Afana, Melike Erol-Kantarci

机构 * School of Electrical Engineering and Computer Science, University of Ottawa(Ottawa 大学电子工程与计算机科学学院) Ericsson(埃里克森公司)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文综述了多模态数据增强的基础模型在无线网络预测与控制中的应用,探讨了其在无线网络管理中的关键任务和未来发展方向。

Comments 5 figures, 7 tables, IEEE COMST

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18215 2025-12-23 cs.LG cs.AI cs.CL cs.CV 82%

Stable and Efficient Single-Rollout RL for Multimodal Reasoning

稳定且高效的多模态推理单次迭代强化学习

Rui Liu, Dian Yu, Lei Ke, Haolin Liu, Yujun Zhou, Zhenwen Liang, Haitao Mi, Pratap Tokekar, Dong Yu

机构 * Tencent AI Lab(腾讯AI实验室) University of Maryland(马里兰大学) University of Virginia(弗吉尼亚大学) University of Notre Dame(诺特大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 MSSR通过熵基优势塑造机制,实现多模态推理中的稳定高效单次迭代强化学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15691 2025-12-18 cs.LG cs.IT cs.SY eess.SP eess.SY math.IT 82%

Multi-Modal Semantic Communication

多模态语义通信

Matin Mortaheb, Erciyes Karakaya, Sennur Ulukus

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) University of Maryland College Park(马里兰大学学院市分校)

专题命中 其他多模态 :multi-modal(title,abstract);cross-modal(abstract)

AI总结 本文提出多模态语义通信框架,通过融合文本查询与视觉特征,实现复杂场景下的高效信息传输与任务关键信息保留。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12305 2025-11-18 cs.LG 82%

MMSense: Adapting Vision-based Foundation Model for Multi-task Multi-modal Wireless Sensing

Zhizhen Li, Xuanhao Luo, Xueren Ge, Longyu Zhou, Xingqin Lin, Yuchen Liu

机构 * North Carolina State University, USA(北卡罗来纳州立大学) University of Virginia, USA(弗吉尼亚大学) Singapore University of Technology(新加坡科技设计大学) NVIDIA Corporation, USA(英伟达公司)

专题命中 其他多模态 :multi-modal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03328 2025-11-06 cs.CL cs.AI cs.CV cs.LG 82%

Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks

Jindong Hong, Tianjie Chen, Lingjie Luo, Chuanyang Zheng, Ting Xu, Haibao Yu, Jianing Qiu, Qianzhong Chen, Suning Huang, Yan Xu, Yong Gui, Yijun He, Jiankai Sun

机构 * Bytedance(字节跳动) Peking University(北京大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Mohamed bin Zayed University of Artificial Intelligence(马尔代夫人工智能大学) Stanford University(斯坦福大学) University of Michigan(密歇根大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.16075 2025-10-31 cs.CL cs.AI cs.CV 82%

Impact of Visual Context on Noisy Multimodal NMT: An Empirical Study for English to Indian Languages

Baban Gain, Dibyanayan Bandyopadhyay, Samrat Mukherjee, Chandranath Adak, Asif Ekbal

机构 * Indian Institute of Technology Patna India Indian Institute of Technology Jodhpur \& Indian Institute of Technology Patna India Indian Institute of Technology Patna Indian Institute of Technology Jodhpur \& Indian Institute of Technology Patna

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18673 2025-10-14 cs.CL cs.AI cs.MM 82%

Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum

Xinglong Yang, Quan Feng, Zhongying Pan, Xiang Chen, Yu Tian, Wentong Li, Shuofei Qiao, Yuxia Geng, Xingyu Zhao, Sheng-Jun Huang

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Hunan Vanguard Group Corporation Co., Ltd.(湖南 Vanguard集团有限公司) Huaneng Information Technology Co., Ltd.(华能信息技术有限公司) Tsinghua University(清华大学) Zhejiang University(浙江大学) PowerChina Huadong Engineering Co., Ltd.(中国电建华东工程有限公司)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07335 2025-10-02 cs.LG cs.AI cs.CV cs.GT cs.MM 82%

Balancing Multimodal Training Through Game-Theoretic Regularization

Konstantinos Kontras, Thomas Strypsteen, Christos Chatzichristos, Paul Pu Liang, Matthew Blaschko, Maarten De Vos

机构 * Department of Electrical Engineering, KU Leuven(KU Leuven 电子工程系) Department of Development and Regeneration, KU Leuven(KU Leuven 发展与再生系) Media Lab and EECS, MIT(MIT 媒体实验室和电子工程与计算机科学系)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments 23 pages, 7 figures, 6 tables, 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25831 2025-10-01 cs.LG 82%

MIDAS: Misalignment-based Data Augmentation Strategy for Imbalanced Multimodal Learning

Seong-Hyeon Hwang, Soyoung Choi, Steven Euijong Whang

机构 * KAIST(韩国科学技术院)

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract)

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12521 2025-09-17 cs.LG 82%

Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time

Yifan Lan, Yuanpu Cao, Weitong Zhang, Lu Lin, Jinghui Chen

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) The University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 其他多模态 :multi-modal(title);multimodal(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07436 2025-09-10 eess.SP 82%

SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding

Feifan Zhang, Yuyang Du, Yifan Xiang, Xiaoyan Liu, Soung Chang Liew

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06312 2025-09-09 eess.SY cs.LG cs.SY 82%

Enhancing Low-Altitude Airspace Security: MLLM-Enabled UAV Intent Recognition

Guangyu Lei, Tianhao Liang, Yuqi Ping, Xinglin Chen, Longyu Zhou, Junwei Wu, Xiyuan Zhang, Huahao Ding, Xingjian Zhang, Weijie Yuan, Tingting Zhang, Qinyu Zhang

机构 * School of Information Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)信息科学与技术学院) Guangdong Provincial Key Laboratory of Space-Aerial Networking and Intelligent Sensing(广东省空间-航空网络与智能感知重点实验室) Information Systems Technology and Design, Singapore University of Technology and Design(新加坡科技设计大学信息系统技术与设计) School of System Design and Intelligent Manufacturing, Southern University of Science and Technology(南方科技大学系统设计与智能制造学院)

专题命中 其他多模态 :MLLM(title,abstract);multimodal(abstract)

Comments The paper has been submitted to IEEE Internet of Things Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01337 2025-09-03 cs.MM cs.AI cs.CL 82%

LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition

Qianrui Zhou, Hua Xu, Yifan Wang, Xinzhi Dong, Hanlei Zhang

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) School of Information Science and Engineering, Hebei University of Science and Technology(河北科技大学信息科学与工程学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted by EMNLP 2025 (Main Track, Long Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21801 2025-09-01 cs.IR 82%

DMGIN: How Multimodal LLMs Enhance Large Recommendation Models for Lifelong User Post-click Behaviors

Zhuoxing Wei, Qingchen Xie, Qi Liu

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract)

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11940 2025-08-05 cs.CE 82%

MLLM-based Discovery of Intrinsic Coordinates and Governing Equations from High-Dimensional Data

Ruikun Li, Yan Lu, Shixiang Tang, Biqing Qi, Wanli Ouyang

专题命中 其他多模态 :MLLM(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22607 2025-08-01 cs.CV cs.AI cs.CL 82%

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning

Ruifeng Yuan, Chenghao Xiao, Sicong Leng, Jianyu Wang, Long Li, Weiwen Xu, Hou Pong Chan, Deli Zhao, Tingyang Xu, Zhongyu Wei, Hao Zhang, Yu Rong

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab(虎盘实验室) Fudan University(复旦大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 21 pages, 5 figures, 6 tables. Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00045 2025-07-02 cs.CV cs.AI cs.CL 82%

CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning

Ming Li, Chenguang Wang, Yijun Liang, Xiyao Wang, Yuhang Zhou, Xiyang Wu, Yuqing Zhang, Ruiyi Zhang, Tianyi Zhou

机构 * University of Maryland, College Park(马里兰大学学院 park)

专题命中 其他多模态 :MLLM(title);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17451 2025-06-09 cs.CL cs.AI cs.CV cs.LG 82%

Diving into Self-Evolving Training for Multimodal Reasoning

Wei Liu, Junlong Li, Xiwen Zhang, Fan Zhou, Yu Cheng, Junxian He

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICML 2025, Project Page: https://mstar-lmm.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏