arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-18 至 2026-03-18 共收录 82 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 6 篇

2603.16664 2026-03-18 cs.CV cs.AI 62%

Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation

Kestrel: 为降低LVLM幻觉而引入自反思

Jiawei Mao, Hardy Chen, Haoqin Tu, Yuhan Wang, Letian Zhang, Zeyu Zheng, Huaxiu Yao, Zirui Wang, Cihang Xie, Yuyin Zhou

机构 * UC Santa Cruz(加州大学圣克ruz分校) UC Berkeley(加州大学伯克利分校) UNC-Chapel Hill(北卡罗来纳大学教堂山分校) Apple(苹果公司)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Kestrel提出一种无需训练的框架,通过显式视觉 grounding 与证据验证自反思机制减少LVLM幻觉,实验显示在POPE和MME-Hallucination基准上性能提升,同时提供透明的验证轨迹。

Comments 16 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21676 2026-03-18 cs.RO cs.NI 50%

Real-World Deployment of Cloud-based Autonomous Mobility Systems for Outdoor and Indoor Environments

云原生自主移动系统在户外和室内环境中的实际部署

Yufeng Yang, Minghao Ning, Keqi Shu, Aladdin Saleh, Ehsan Hashemi, Amir Khajepour

机构 * Department of Mechanical and Mechatronics Engineering, University of Waterloo(滑铁卢大学机械与机电工程系) Technology Partnerships and Innovations, Rogers Communications, Canada Inc.(罗杰斯通讯加拿大有限公司技术伙伴关系与创新部) Mechanical Engineering Department, University of Alberta(阿尔伯塔大学机械工程系)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出云原生自主移动框架,通过基础设施智能传感与云计算协调提升自主操作能力,实验证明在城市环岛和医院类室内环境中的感知鲁棒性和安全性提升。

Comments This paper has been submitted to IEEE Robotics and Automation Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18373 2026-03-18 cs.RO cs.HC 50%

UGotMe: An Embodied System for Affective Human-Robot Interaction

UGotMe: 一种用于情感人机交互的具身系统

Peizhen Li, Longbing Cao, Xiao-Ming Wu, Xiaohan Yu, Runze Yang

机构 * School of Computing, Macquarie University(麦考瑞大学计算学院) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) Department of Automation, Shanghai Jiao Tong University(上海交通大学自动化学院)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出UGotMe系统,解决多对话场景中视觉噪声和实时响应问题,通过去噪策略和高效数据传输提升情感识别能力。

Comments Accepted to the 2025 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态训练与对齐 12 篇

2603.15818 2026-03-18 cs.CV 83%

Conflict-Aware Multimodal Fusion for Ambivalence and Hesitancy Recognition

具有冲突意识的多模态融合用于矛盾与犹豫识别

Salah Eddine Bekhouche, Hichem Telli, Azeddine Benlamoudi, Salah Eddine Herrouz, Abdelmalik Taleb-Ahmed, Abdenour Hadid

机构 * University of the Basque Country UPV/EHU(巴斯克大学UPV/EHU) Laboratory of LESIA, University of Biskra(贝斯克拉大学LESIA实验室) Lab. de Génie Electrique (LAGE), University Kasdi Merbah Ouargla(奥尔加拉大学LAGE实验室) Institute of Electronics, Microelectronics and Nanotechnology (IEMN), Polytechnic University of Hauts-de-France, University of Lille(电子、微电子与纳米技术研究所(IEMN),法国 Hauts-de-France 工业大学,里尔大学) Sorbonne Center for Artificial Intelligence, Sorbonne University Abu Dhabi, UAE(索邦人工智能中心,索邦大学阿布扎比校区,阿联酋)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出ConflictAwareAH框架,通过多模态融合识别矛盾与犹豫状态,提升F1指标,采用冲突特征作为双向线索,改进模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16259 2026-03-18 cs.MM 79%

Hyperbolic Multimodal Generative Representation Learning for Generalized Zero-Shot Multimodal Information Extraction

双曲多模态生成表示学习用于广义零样本多模态信息提取

Baohang Zhou, Kehui Song, Rize Jin, Yu Zhao, Xuhui Sui, Xinying Qian, Xingyue Guo, Ying Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

AI总结 本文提出双曲多模态生成表示学习框架HMGRL,解决零样本多模态信息提取中见与不见类别共存的问题,通过双曲空间建模多级语义关联,提升模型泛化能力。

Comments Accepted by WWW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16143 2026-03-18 eess.SP cs.AI 79%

Structure-Aware Multimodal LLM Framework for Trustworthy Near-Field Beam Prediction

面向可信近场波束预测的结构感知多模态大语言模型框架

Mengyuan Li, Qianfan Lu, Jiachen Tian, Hongjun Hu, Yu Han, Xiao Li, Chao-kai Wen, Shi Jin

机构 * School of Information Science and Engineering, Southeast University(信息科学与工程学院,东南大学) Institute of Communications Engineering, National Sun Yat-sen University(通讯工程学院,国立中山大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出基于大语言模型的多模态框架,融合历史GPS数据、RGB图像、LiDAR数据及任务特定文本提示,以提升复杂低空环境中的近场波束对齐能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19596 2026-03-18 eess.SP cs.LG 78%

Towards Robust Multimodal Physiological Foundation Models: Handling Arbitrary Missing Modalities

迈向稳健的多模态生理基础模型:处理任意缺失模态

Wei-Bang Jiang, Xi Fu, Yi Ding, Cuntai Guan

机构 * Nanyang Technological University(南洋理工大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出PhysioOmni模型,通过解耦多模态信号提取通用表示,解决现有方法在跨数据集泛化和缺失模态处理上的不足,实验显示其在情绪识别等任务中表现优异。

Comments 19 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16482 2026-03-18 cs.CV cs.AI 62%

DST-Net: A Dual-Stream Transformer with Illumination-Independent Feature Guidance and Multi-Scale Spatial Convolution for Low-Light Image Enhancement

DST-Net:一种双流Transformer,具有光照无关特征引导和多尺度空间卷积的低光照图像增强

Yicui Shi, Yuhan Chen, Xiangfei Huang, Zhenguo Wang, Wenxuan Yu, Ying Fang

机构 * College of Mechanical and Vehicle Engineering, Chongqing University(重庆大学机械与车辆工程学院) Shanghai Zhenhua Heavy Industries Co.,Ltd.(上海震华重工有限公司)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出DST-Net,通过光照无关特征引导和多尺度空间卷积提升低光照图像增强效果,解决传统方法在信号先验丢失的问题,实验表明其在主观和客观指标上均表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16245 2026-03-18 cs.CV cs.CL 62%

How to Utilize Complementary Vision-Text Information for 2D Structure Understanding

如何利用互补的视觉-文本信息进行2D结构理解

Jiancheng Dong, Pengyue Jia, Derong Xu, Jiawei Cheng, Jingyu Peng, Chao Zhang, Bowen Liu, Xin Sun, Lixin Su, Shuaiqiang Wang, Dawei Yin, Xiangyu Zhao

机构 * City University of Hong Kong(香港城市大学) Baidu Inc(百度公司)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出DiVA-Former架构,通过动态查询视觉标记将长文本序列压缩为摘要向量,有效整合视觉与文本信息,在13个表格基准测试中提升23.9%。

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15965 2026-03-18 cs.CL cs.AI 62%

MoLoRA: Composable Specialization via Per-Token Adapter Routing

MoLoRA:通过每token适配器路由实现可组合的专精

Shrey Shah, Justin Wagle

机构 * Microsoft(微软)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 MoLoRA通过每token适配器路由实现可组合专精,优于规模扩展:MoLoRA使Qwen3-1.7B在四个推理基准上超越Qwen3-8B,且小4.7倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13000 2026-03-18 cs.CV 61%

Omni Survey for Multimodality Analysis in Visual Object Tracking

多模态分析在视觉目标跟踪中的全方位调研

Zhangyong Tang, Tianyang Xu, Xuefeng Zhu, Hui Li, Shaochuan Zhao, Tao Zhou, Chunyang Cheng, Xiaojun Wu, Josef Kittler

机构 * School of Artificial Intelligence and Computer Science, Jiangnan University(江南大学人工智能与计算机科学学院) School of Computer Science and Technology, China University of Mining and Technology(中国矿业大学计算机科学与技术学院) School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) Centre for Vision, Speech and Signal Processing, University of Surrey(Surrey 大学视觉、语音与信号处理中心)

专题命中 多模态训练与对齐 :multi-modal(abstract,comments);分类 cs.CV

AI总结 本文从多模态分析角度调研了多模态视觉目标跟踪(MMVOT)的关键问题,涵盖数据收集、模态对齐、标注、模型设计与评估,分析了现有方法的分类及挑战,并首次探讨了多模态跟踪的适用性及数据集中的类别分布。

Comments The first comprehensive survey for multi-modal visual object tracking; 6 multi-modal tasks; 338 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16671 2026-03-18 cs.CV 57%

$x^2$-Fusion: Cross-Modality and Cross-Dimension Flow Estimation in Event Edge Space

$x^2$-Fusion:事件边缘空间中的跨模态和跨维度流估计

Ruishan Guo, Ciyu Ruan, Haoyang Wang, Zihang Gong, Jingao Xu, Xinlei Chen

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Harbin Institute of Technology(哈尔滨工业大学) The University of Hong Kong(香港大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出$x^2$-Fusion方法,通过事件边缘空间实现跨模态和跨维度流的统一表示,提升动态场景理解的准确性和鲁棒性。

Comments This version is the camera-ready version accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21637 2026-03-18 cs.CV 57%

CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image Analysis

CARE:一种具有自适应区域建模的分子引导基础模型用于全切片图像分析

Di Zhang, Zhangpeng Gong, Xiaobo Pang, Jiashuai Liu, Junbo Lu, Hao Cui, Jiusong Ge, Zhi Zeng, Kai Yi, Yinghua Li, Si Liu, Tingsong Yu, Haoran Wang, Mireia Crispin-Ortuzar, Weimiao Yu, Chen Li, Zeyu Gao

机构 * Xi’an Jiaotong University(西安交通大学) University of Cambridge(剑桥大学) KingMed(康方生物) BGI Research(贝登基因研究院) A ⋆ STAR

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 CARE通过自适应区域建模和分子引导,提升全切片图像分析的性能,实现对病理区域的精准识别与分类,优于现有基础模型。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18444 2026-03-18 cs.CV 57%

SineProject: Machine Unlearning for Stable Vision Language Alignment

SineProject:用于稳定视觉语言对齐的机器反遗忘

Arpit Garg, Hemanth Saratchandran, Simon Lucey

机构 * Australian Institute for Machine Learning(澳大利亚机器学习研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 SineProject通过在冻结的投影器中加入正弦调制的可训练参数,提升Jacobian的谱条件数,稳定反遗忘过程中的视觉语言对齐,减少无害查询拒绝并实现目标信息遗忘。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21524 2026-03-18 cs.AI cs.IT math.IT 57%

Large Language Models for Wireless Communications: From Adaptation to Autonomy

大语言模型在无线通信中的应用:从适应到自主

Le Liang, Hao Ye, Yucheng Sheng, Ouya Wang, Jiacheng Wang, Shi Jin, Geoffrey Ye Li

机构 * Southeast University(东南大学) University of California, Santa Cruz(加州大学圣克鲁兹分校) Imperial College London(伦敦帝国理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 本文探讨大语言模型在无线通信中的应用,包括适应预训练模型、开发专用基础模型以及实现自主推理的代理模型,总结其优势与挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他多模态 7 篇

2508.03715 2026-03-18 eess.SP cs.AI cs.HC cs.LG 79%

Detection of Autonomic Dysreflexia in Individuals With Spinal Cord Injury Using Multimodal Wearable Sensors

利用多模态可穿戴传感器检测脊髓损伤患者自主神经反射症

Bertram Fuchs, Mehdi Ejtehadi, Ana Cisnal, Jürgen Pannek, Anke Scheel-Sailer, Robert Riener, Inge Eriks-Hoogland, Diego Paez-Granados

机构 * SCAI-Lab, Department of Health Sciences and Technology (D-HEST), ETH Zurich(SCAI实验室,健康科学与技术系(D-HEST),苏黎世联邦理工学院) School of Computation, Information and Technology, Technical University of Munich(计算、信息与技术学院,慕尼黑技术大学) Swiss Paraplegic Research (SPF)(瑞士瘫痪研究机构(SPF)) Institute of Advanced Production Technologies, University of Valladolid(先进生产技术研究所,瓦尔达利德大学) Swiss Paraplegic Centre (SPZ)(瑞士瘫痪中心(SPZ)) Balgrist University Hospital(巴尔格里斯大学医院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出一种非侵入性的机器学习框架,通过多模态可穿戴传感器检测自主神经反射症,利用BorutaSHAP进行特征选择,并通过堆叠集成模型提高检测性能,结果显示心率和心电图特征在检测中表现最佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19553 2026-03-18 math.ST cs.LG math.PR stat.ML stat.TH 78%

Convergence Bounds for Sequential Monte Carlo on Multimodal Distributions using Soft Decomposition

基于软分解的多模分布序贯蒙特卡洛算法收敛性界

Holden Lee, Matheau Santana-Gijzen

机构 * Department of Applied Mathematics \& Statistics, Johns Hopkins University e1,e2

专题命中 其他多模态 :multimodal(title);multi-modal(abstract)

AI总结 本文研究了序贯蒙特卡洛算法在多模分布上的收敛性界,通过软分解方法,利用局部马尔可夫链混合动态而非全局混合动态,提出新的收敛性界分析方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16590 2026-03-18 cs.CL cs.AI 73%

BATQuant: Outlier-resilient MXFP4 Quantization via Learnable Block-wise Optimization

BATQuant: 通过可学习的分块优化实现抗异常的MXFP4量化

Ji-Fu Li, Manyi Zhang, Xiaobo Xia, Han Bao, Haoli Bai, Zhenhua Dong, Xianzhi Yu

机构 * Huawei Technologies University of Science(华为技术大学科学)

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出BATQuant方法,通过可学习的分块优化解决MXFP4量化中异常传播问题,实现高性能量化方案。

Comments 30 pages, 13 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14610 2026-03-18 cs.CV eess.IV 57%

Make it SING: Analyzing Semantic Invariants in Classifiers

使分类器具备语义不变性:分析分类器中的语义不变性

Harel Yadid, Meir Yossef Levi, Roy Betser, Guy Gilboa

机构 * Viterbi Faculty of Electrical and Computer Engineering(电气与计算机工程学院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出SING方法,通过构建等价图像并赋予语义解释,揭示分类器中语义不变性,分析不同模型在保持语义方面的差异。

Comments Accepted to the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17018 2026-03-18 cs.CV 57%

SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward

SophiaVL-R1: 通过思考奖励强化大语言模型的推理能力

Kaixuan Fan, Kaituo Feng, Haoming Lyu, Dongzhan Zhou, Xiangyu Yue

机构 * MMLab, The Chinese University of Hong Kong(MMLab,香港中文大学) Shanghai Artifcial Intelligence Laboratory(上海人工智能实验室)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 SophiaVL-R1通过引入思考奖励信号提升多模态大语言模型的推理能力,采用信任GRPO方法和退火训练策略,有效缓解奖励黑客问题,实验表明其在多个基准测试中表现优异。

Comments ICLR 2026, Project page:https://github.com/kxfan2002/SophiaVL-R1

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15767 2026-03-18 cs.CV 57%

CLRNet: Targetless Extrinsic Calibration for Camera, Lidar and 4D Radar Using Deep Learning

CLRNet:基于深度学习的无目标外校准方法用于摄像头、激光雷达和4D雷达

Marcell Kegl, Andras Palffy, Csaba Benedek, Dariu M. Gavrila

机构 * Hungarian Research Network Institute for Computer Science and Control (HUN-REN SZTAKI)(匈牙利研究网络计算机科学与控制研究所) Pázmány Péter Catholic University(帕兹曼·佩斯特天主教大学) Intelligent Vehicles Group, TU Delft(智能车辆组,代尔夫特理工大学) Perciv AI

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出CLRNet,一种多模态端到端深度学习网络,用于摄像头、激光雷达和4D雷达的联合校准或配对校准,通过等距投影、基于摄像头的深度图像预测和雷达通道等方法,提升校准精度。

Comments Submitted to IEEE Transactions on Intelligent Vehicles

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.07716 2026-03-18 math.DS 50%

Existence of Large deviations rate function for any $S$-unimodal map

任意S-单峰映射的大偏差率函数的存在性

Hiroki Takahasi, Masato Tsujii

专题命中 其他多模态 :multimodal(abstract)

AI总结 研究了任意负Schwarzian单峰映射的水平2大偏差原理,并给出了多峰映射不满足该原理的例子。

Comments 47 pages, 2 figures, to appear in Ergodic Theory and Dynamical Systems

详情

展开后加载摘要…

URL PDF HTML 收藏