arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-24 至 2026-03-24 共收录 148 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 20 篇

2603.20325 2026-03-24 cs.CV 57%

DCG-Net: Dual Cross-Attention with Concept-Value Graph Reasoning for Interpretable Medical Diagnosis

DCG-Net:双交叉注意力与概念-值图推理用于可解释的医学诊断

Getamesay Dagnaw, Xuefei Yin, Muhammad Hassan Maqsood, Yanming Zhu, Alan Wee-Chung Liew

机构 * School of Information and Communication Technology(信息与通信技术学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 DCG-Net通过双交叉注意力和概念-值图推理,提升医学诊断的可解释性,实现白血球形态和皮肤病变诊断的高精度分类。

Journal ref ICME 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20289 2026-03-24 cs.CV 57%

Remote Sensing Image Dehazing: A Systematic Review of Progress, Challenges, and Prospects

遥感图像去雾:进展、挑战与前景的系统综述

Heng Zhou, Xiaoxiong Liu, Zhenxi Zhang, Jieheng Yun, Chengyang Li, Yunchu Yang, Dongyi Xia, Chunna Tian, Xiao-Jun Wu

机构 * School of Artificial Intelligence and Computer Science(人工智能与计算机科学学院) Josef Kittler Research Institute on Artificial Intelligence(乔塞夫·基特勒人工智能研究所) School of Electronic Engineering(电子工程学院) College of Artificial Intelligence(人工智能学院) Aerospace Information Research Institute(航天信息研究所) Department of Computer Science(计算机科学系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文系统综述了遥感图像去雾的研究进展、挑战与前景,总结了超过30种方法,分析了不同模型在性能上的差异,并提出了未来研究方向。

Comments 82 pages, 23 figures,

Journal ref ISPRS P&RS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21263 2026-03-24 cs.SE 50%

From Natural Language to Executable Properties for Property-based Testing of Mobile Apps

从自然语言到可执行属性:用于移动应用基于属性测试的可执行属性

Yiheng Xiong, Ting Su, Jingling Sun, Jue Wang, Qin Li, Geguang Pu, Zhendong Su

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文提出了一种新的结构化属性合成方法,将自然语言属性描述自动转换为可执行属性,并通过iPBT工具验证,展示了其在提升测试效率和鲁棒性方面的贡献。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20721 2026-03-24 cs.LG math.ST stat.ML stat.TH 50%

Scaling Laws are Redundancy Laws

缩放定律是冗余定律

Yuda Bi, Vince D Calhoun

专题命中 多模态评测 :multi-modal(abstract)

AI总结 本文通过核回归证明缩放定律可形式化为冗余定律,揭示学习曲线斜率依赖数据冗余,统一了经验观察与理论基础。

Comments This is not a serious research at this time

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03467 2026-03-24 cs.IT cs.CR cs.LG eess.SP math.IT stat.ME 50%

Differentially Private Distribution Release of Gaussian Mixture Models via KL-Divergence Minimization

基于KL散度最小化的高斯混合模型差分隐私分布发布

Hang Liu, Anna Scaglione, Sean Peisert

机构 * State Key Laboratory of Internet of Things for Smart City and the Department of Electrical and Computer Engineering, University of Macau(物联网智能城市国家重点实验室和澳门大学电子与计算机工程系) Department of Electrical and Computer Engineering, Cornell Tech, Cornell University(电气与计算机工程系,康奈尔科技,康奈尔大学) Computing Sciences Research, Lawrence Berkeley National Laboratory(计算科学研究所,劳伦斯伯克利国家实验室)

专题命中 多模态评测 :multi-modal(abstract)

AI总结 本文提出通过KL散度度量高斯混合模型发布精度,结合差分隐私机制,在保证隐私安全的同时保持模型效用。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 16 篇

2601.10744 2026-03-24 cs.AI cs.CV 81%

Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration

探索与长期记忆:一个基准和基于多模态LLM的强化学习框架用于具身探索

Sen Wang, Bangwei Liu, Zhenkun Gao, Lizhuang Ma, Xuhong Wang, Yuan Xie, Xin Tan

机构 * East China Normal University(东华大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出LMEE框架,通过多模态LLM强化学习促进终身学习,构建LMEE-Bench基准评估具身探索过程与结果,采用MemoryExplorer方法提升记忆检索与主动探索能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20198 2026-03-24 cs.CR cs.CV cs.LG 79%

Visual Exclusivity Attacks: Automatic Multimodal Red Teaming via Agentic Planning

视觉排斥攻击:通过代理规划实现自动多模态红队行动

Yunbei Zhang, Yingqiang Ge, Weijie Xu, Yuhui Xu, Jihun Hamm, Chandan K. Reddy

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出视觉排斥攻击,通过多模态多轮代理规划框架MM-Plan,实现基于视觉内容的攻击,展示了对前沿模型的高攻击成功率,揭示了当前安全对齐的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22507 2026-03-24 cs.NI cs.MA eess.SP 78%

A Unified Cloud-Edge-Terminal Framework for Multimodal Integrated Sensing and Communication

多模态感知与通信一体化的统一云-边-终端框架

Yubo Peng, Luping Xiang, Kun Yang, Feibo Jiang, Kezhi Wang, Christos Masouros

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出统一云-边-终端框架,通过多模态感知与通信融合,解决异构融合、通信开销和系统扩展性等挑战,提升任务导向的多模态感知能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01013 2026-03-24 cs.LG 78%

TimeXL: Explainable Multi-modal Time Series Prediction with LLM-in-the-Loop

TimeXL:基于LLM的多模态时间序列预测可解释方法

Yushan Jiang, Wenchao Yu, Geon Lee, Dongjin Song, Kijung Shin, Wei Cheng, Yanchi Liu, Haifeng Chen

机构 * School of Computing, University of Connecticut(大学计算机学院) Data Science & System Security Department, NEC Labs America(数据科学与系统安全部,NEC美国实验室) Kim Jaechul Graduate School of AI, KAIST(金 Jaechul人工智能研究生院,韩国科学技术院)

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 TimeXL通过集成原型时间序列编码器与三个协作LLM,提升时间序列预测的准确性与可解释性,实验证明在四个真实数据集上AUC提升达8.9%。

Comments NeurIPS 2025 camera ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21577 2026-03-24 cs.AI 74%

Mind over Space: Can Multimodal Large Language Models Mentally Navigate?

心灵超越空间:多模态大语言模型能否进行心理导航?

Qihui Zhu, Shouwei Ruan, Xiao Yang, Hao Jiang, Yao Huang, Shiji Zhao, Hanwei Fan, Hang Su, Xingxing Wei

机构 * Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院) Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University(清华大学-博世联合机器学习中心、THBI实验室、BNRist中心、清华大学计算机科学与技术系) School of Automation Science and Electrical Engineering , Beihang University(北京航空航天大学自动化科学与电气工程学院) college of AI, Tsinghua University(清华大学人工智能学院) Department of Computer Science and Technology , Tsinghua University(清华大学计算机科学与技术系)

专题命中 多模态Agent :multimodal(title);分类 cs.AI

AI总结 本文提出Video2Mental基准测试,评估多模态大语言模型的空间导航能力,发现标准预训练模型无法自然生成空间表示,NavMind通过显式认知地图提升导航性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21410 2026-03-24 cs.RO 71%

Bayesian Active Object Recognition and 6D Pose Estimation from Multimodal Contact Sensing

基于多模态接触传感的贝叶斯主动物体识别与6D位姿估计

Haodong Zheng, Gabriele M. Caddeo, Andrei C. Jalba, Wijnand A. IJsselsteijn, Lorenzo Natale, Raymond H. Cuijpers

机构 * Eindhoven University of Technology(埃因霍温理工大学) Italian Institute of Technology(意大利理工学院) Humanoid Sensing and Perception Group(人形感知与感知小组)

专题命中 多模态Agent :multimodal(title)

AI总结 本文提出一种结合触觉与自由空间约束的贝叶斯框架,用于联合物体识别和6D位姿估计,通过定制粒子滤波器提升推理效率,并利用推理结果指导主动探索,实验表明触觉信息显著提升了识别和位姿估计的准确性与稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21341 2026-03-24 cs.AI 70%

RoboAlign: Learning Test-Time Reasoning for Language-Action Alignment in Vision-Language-Action Models

RoboAlign:学习测试时推理以改进视觉-语言-动作模型中的语言-动作对齐

Dongyoung Kim, Sumin Park, Woomin Song, Seungku Kim, Taeyoung Kim, Huiwon Jang, Jinwoo Shin, Jaehyung Kim, Younggyo Seo

机构 * Yonsei University(延世大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI

AI总结 本文提出RoboAlign框架,通过零样本自然语言推理生成动作标记并结合强化学习提升动作准确性,从而在视觉-语言-动作模型中实现语言与低级动作之间的对齐,提升模型性能。

Comments 15 pages, 7 figures, 9 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20200 2026-03-24 cs.RO cs.AI cs.CV 62%

Your Robot Will Feel You Now: Empathy in Robots and Embodied Agents

你的机器人将能感知你:机器人与具身代理中的共情

Angelica Lim, Ö. Nilay Yalçin

机构 * School of Computing Science, Simon Fraser University(计算科学学院,西蒙弗雷泽大学) School of Interactive Arts and Technology, Simon Fraser University(交互艺术与技术学院,西蒙弗雷泽大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文综述了人类机器人交互和具身对话代理领域中共情行为和模型的研究,探讨如何将这些经验应用于当今基于语言的代理系统。

Comments Accepted manuscript. Chapter in "Empathy and Artificial Intelligence: Challenges, Advances and Ethical Considerations" edited by Anat Perry; C. Daryl Cameron

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22280 2026-03-24 cs.CV cs.RO 57%

DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models

DualCoT-VLA: 通过并行推理实现视觉-语言链式思维的视觉-语言-动作模型

Zhide Zhong, Junfeng Li, Junjie He, Haodong Yan, Xin Gong, Guanyi Zhao, Yingjie Cai, Jiantao Gao, Xu Yan, Bingbing Liu, Yingcong Chen, Liuqing Yang, Haoang Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Huawei Foundation Model Department(华为基础模型部门)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

AI总结 DualCoT-VLA通过并行推理机制,结合视觉和语言链式思维,解决传统VLA模型在复杂多步骤任务和精细空间感知中的不足,实现更高效的视觉-语言-动作处理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21029 2026-03-24 cs.AI 57%

KLDrive: Fine-Grained 3D Scene Reasoning for Autonomous Driving based on Knowledge Graph

KLDrive: 基于知识图谱的细粒度3D场景推理用于自动驾驶

Ye Tian, Jingyi Zhang, Zihao Wang, Xiaoyuan Ren, Xiaofan Yu, Onat Gungor, Tajana Rosing

机构 * University of California San Diego, La Jolla, California, USA(加州大学圣地亚哥分校) University of California Merced, Merced, California, USA(加州大学默塞德分校)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 KLDrive通过构建可靠场景知识图谱和LLM代理进行事实驱动推理,提升自动驾驶中的细粒度问答性能,实验表明其在NuScenes-QA和GVQA上均取得最佳准确率和SPICE分数。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21706 2026-03-24 physics.med-ph 50%

Comprehensive Dosimetric Verification and Positional Sensitivity Analysis in Brachytherapy: A Unified ESAPI Tool for HDR and LDR Treatments

放射治疗中的全面剂量学验证与位置敏感性分析:一种用于 HDR 和 LDR 治疗的统一 ESAPI 工具

J. A. Valgoma

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出了一种基于 Varian Eclipse Scripting API 的独立软件工具,用于验证 HDR 和 LDR 放射治疗的 QA,通过比较点源和线源模型,分析位置不确定性,并提高临床工作流程的安全性。

Comments 13 pages, 2 tables, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01834 2026-03-24 cs.RO 50%

Concept-Based Dictionary Learning for Inference-Time Safety in Vision Language Action Models

基于概念的词典学习用于推理时的安全性在视觉语言动作模型中

Siqi Wen, Shu Yang, Shaopeng Fu, Jingfeng Zhang, Lijie Hu, Di Wang

机构 * Beijing Jiaotong University(北京交通大学) Provable Responsible AI and Data Analytics (PRADA) Lab(可证责任AI与数据分析实验室) King Abdullah University of Science and Technology(国王 Abdullah 科学技术大学) University of Auckland(奥克兰大学) RIKEN Center for Advanced Intelligence Project (AIP)(理化学研究所高级智能项目中心) Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出基于概念的词典学习框架,用于提升视觉语言动作模型推理时的安全性,通过学习稀疏可解释词典识别有害概念方向并抑制风险组件,实验表明其在多个基准上有效降低攻击成功率70%以上。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19988 2026-03-24 stat.ML cs.LG q-bio.QM 50%

BioBO: Biology-informed Bayesian Optimization for Perturbation Design

BioBO:结合生物学知识的贝叶斯优化用于扰动设计

Yanke Li, Tianyu Cui, Tommaso Mansi, Mangal Prakash, Rui Liao

机构 * Johnson & Johnson Innovative Medicine(强生创新医药) ETH Zurich(苏黎世联邦理工学院)

专题命中 多模态Agent :multimodal(abstract)

AI总结 BioBO结合多模态基因嵌入和富集分析,提升代理建模和获取策略,提高标签效率25-40%,并提供路径级解释,链接设计与生物调控电路。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16963 2026-03-24 cs.RO cs.SY eess.SY 50%

A Tactile-based Interactive Motion Planner for Robots in Unknown Cluttered Environments

基于触觉的交互式运动规划器用于未知杂乱环境中的机器人

Chengjin Wang, Yanmin Zhou, Zheng Yan, Feng Luan, Runjie Shen, Hongrui Sang, Zhipeng Wang, Bin He

机构 * Shanghai Research Institute for Intelligent Autonomous Systems(上海智能自主系统研究院) State Key Laboratory of Autonomous Intelligent Unmanned Systems(自主智能无人系统国家重点实验室) Frontiers Science Center for Intelligent Autonomous Systems(智能自主系统前沿科学中心) College of Electronics and Information Engineering, Tongji University(同济大学电子与信息学院)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出了一种基于触觉的交互式运动规划框架,通过多模态触觉感知实时构建接触模型,从而在未知杂乱环境中安全扩展自由运动空间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20560 2026-03-24 cs.HC cs.GR 50%

Nevis Digital Twin: Photogrammetry and Immersive Visualization of Historical Sites

Nevis数字孪生:历史遗址的摄影测量与沉浸式可视化

Alex Apffel, Huy Tran, Vuthea Chheang

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出了一种多模态数据采集流程,用于保护受威胁的历史遗址,通过摄影测量和3D高斯点散布实现虚拟重建,以提供可扩展的数字遗产民主化模型。

Comments ARCHERIX Workshop - IEEE VR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20355 2026-03-24 eess.IV 50%

CaroTo: A Tool for Fast Comprehensive Analysis of Carotid Artery Stenosis in 4D PC- and 3D BB-MRI Data

CaroTo:一种用于快速全面分析4D PC-和3D BB-MRI数据颈动脉狭窄的工具

Hinrich Rahlfs, Markus Hüllebrand, Sebastian Schmitter, Jonathan Andrae, Christoph Strecker, Andreas Harloff, Anja Hennemuth

专题命中 多模态Agent :multimodal(abstract)

AI总结 CaroTo工具通过多模态和多维分割、生物标志物提取和可视化,实现颈动脉动脉粥样斑块的标准化评估,提升颈动脉狭窄分析的精度和一致性。

Comments VCBM 2024, Poster Honorable Mention

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 27 篇

2505.11404 2026-03-24 cs.CV cs.AI 86%

Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner

Patho-R1: 基于多模态强化学习的病理专家推理器

Wenchuan Zhang, Penghao Zhang, Jingru Guo, Tao Cheng, Jie Chen, Shuwan Zhang, Zhang Zhang, Yuhao Yi, Hong Bu

机构 * Department of Pathology, West China Hospital, Sichuan University(四川大学华西医院病理科部门) Institute of Clinical Pathology, West China Hospital, Sichuan University(四川大学华西医院临床病理科研究所) University of Toronto(多伦多大学) Business School, Sichuan University(四川大学商学院) Department of Pathology, Shengjing Hospital of China Medical University(中国医科大学盛京医院病理科部门)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文提出Patho-R1,通过构建高质量推理导向数据集,结合三阶段训练流程提升病理推理能力,实现跨模态任务的鲁棒性能。

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(33): 28418-28426, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20808 2026-03-24 cs.CV cs.LG 85%

Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models

预测正则化对抗多模态大语言模型中的视觉表征退化

Enguang Wang, Qiang Wang, Yuanchen Wu, Ke Yan, Xinbin Yuan, Shouhong Ding, Xialei Liu, Ming-Ming Cheng

机构 * NKIARI VCIP, CS, Nankai University(VCIP计算机科学系,南开大学) AAIS, Nankai University(AAIS,南开大学) Tencent Youtu Lab(腾讯优设实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文研究多模态大语言模型中的视觉表征退化问题,提出预测正则化方法以维持视觉表征,提升视觉语言性能。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21100 2026-03-24 cs.CV cs.AI 84%

Learning Progressive Adaptation for Multi-Modal Tracking

多模态跟踪的渐进适应学习

He Wang, Tianyang Xu, Zhangyong Tang, Xiao-Jun Wu, Josef Kittler

机构 * School of Artificial Intelligence and Computer Science, Jiangnan University(江南大学人工智能与计算机科学学院) Centre for Vision, Speech and Signal Processing, University of Surrey(Surrey大学视觉、语音和信号处理中心)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出PATrack方法,通过引入模态依赖、模态交织和任务级适配器,解决多模态跟踪中预训练RGB模型适应问题,提升跨模态交互和预测头的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21584 2026-03-24 cs.LG cs.CV 83%

SSAM: Singular Subspace Alignment for Merging Multimodal Large Language Models

SSAM:奇异子空间对齐用于融合多模态大语言模型

Md Kaykobad Reza, Ameya Patil, Edward Ayrapetian, M. Salman Asif

机构 * University of California Riverside(加州大学河滨分校) Amazon(亚马逊)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 SSAM通过参数空间对齐融合多模态大语言模型,无需训练数据实现跨模态统一,提升性能并降低资源消耗。

Comments 25 Pages, 9 Figures, 5 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14188 2026-03-24 cs.CV 83%

Joint Segmentation and Grading with Iterative Optimization for Multimodal Glaucoma Diagnosis

多模态青光眼诊断的联合分割与分级迭代优化方法

Zhiwei Wang, Yuxing Li, Meilu Zhu, Defeng He, Edmund Y. Lam

机构 * Department of Electrical and Electronic Engineering, The University of Hong Kong, Hong Kong, China(香港大学电子与电气工程系) College of Information Engineering, Zhejiang University of Technology, Hangzhou, China(浙江工业大学信息工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种迭代多模态优化模型,通过中层融合策略整合眼底和OCT特征,并利用跨模态特征对齐模块减少模态差异,实现青光眼的精确分割与分级。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03521 2026-03-24 cs.MM cs.LG 83%

Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation

跨空间协同:一种用于对话中多模态情感识别的统一框架

Xiaosen Lyu, Jiayu Xiong, Yuren Chen, Wanlong Wang, Xiaoqing Dai, Jing Wang

机构 * Xiaosen Lyu 1,2(李绍森 1,2) Jiayu Xiong 1,2(熊佳宇 1,2) Yuren Chen 1,2(陈远人 1,2) Wanlong Wang 1,2(王万龙 1,2) Xiaoqing Dai 1,2(戴晓青 1,2) Jing Wang 1,2(王婧 1,2)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

AI总结 本文提出Cross-Space Synergy框架,通过协同多项式融合和帕累托梯度调节器有效提升多模态情感识别的准确性和训练稳定性。

Comments Accepted to AAAI 2026

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(29), 24226-24234 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22862 2026-03-24 cs.LG cs.CV 83%

Bridging Modalities via Progressive Re-alignment for Multimodal Test-Time Adaptation

通过渐进重对齐桥接模态以实现多模态测试时适应

Jiacheng Li, Songhe Feng

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出BriMPR框架,通过分治策略解决多模态测试时适应中的模态间分布偏移和语义对齐问题,通过提示调优和跨模态对比学习提升多模态特征对齐效果。

Comments Accepted by AAAI 2026 (Oral)

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence. 2026, 40(27): 22931-22939

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21612 2026-03-24 cs.LG 82%

Towards Multimodal Time Series Anomaly Detection with Semantic Alignment and Condensed Interaction

迈向多模态时间序列异常检测的语义对齐与压缩交互

Shiyan Hu, Jianxin Jin, Yang Shu, Peng Chen, Bin Yang, Chenjuan Guo

机构 * East China Normal University(华东师范大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文提出MindTS模型,通过语义对齐和压缩交互解决多模态时间序列异常检测中的关键问题,实验表明其性能优于现有方法。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20729 2026-03-24 cs.CV cs.AI physics.geo-ph 81%

Weakly supervised multimodal segmentation of acoustic borehole images with depth-aware cross-attention

弱监督多模态分割:基于深度感知的跨注意力机制用于声学钻孔图像

Jose Luis Lima de Jesus Silva

机构 * Federal University of Bahia, Institute of Geosciences, Department of Geophysics(巴伊亚联邦大学,地质科学学院,地球物理学系) Grupo de Estudos e Aplicação de Inteligência Artificial em Geofísica (GAIA)(地质物理中人工智能研究与应用小组)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出一种弱监督多模态分割框架,通过深度感知的跨注意力机制提升钻孔图像分割性能,结合二维图像纹理与一维井下数据,实现无监督的高精度分割。

详情

展开后加载摘要…

URL PDF HTML 收藏