arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2603.01758 2026-03-03 cs.CV 79%

Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining

通过语言枢轴预训练统一异构多模态遥感检测

Yuxuan Li, Yuming Chen, Yunheng Li, Ming-Ming Cheng, Xiang Li, Jian Yang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 BabelRS通过语言枢轴预训练框架统一异构多模态遥感检测,解耦模态对齐与任务学习,提升训练稳定性与检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09285 2026-03-03 cs.CV 79%

Spotlight on Token Perception for Multimodal Reinforcement Learning

多模态强化学习中的token感知聚焦

Siyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo, Zefeng He, Daizong Liu, Yu Cheng

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学) Nanjing University(南京大学) Wuhan University(武汉大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出VPPO算法,通过token感知优化提升多模态强化学习的视觉推理能力。

Comments Accepted by ICLR 2026, project page: https://github.com/huaixuheqing/VPPO-RL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03214 2026-03-03 cs.CV 79%

RTGMFF: Enhanced fMRI-based Brain Disorder Diagnosis via ROI-driven Text Generation and Multimodal Feature Fusion

RTGMFF:基于ROI驱动文本生成和多模态特征融合的增强型fMRI脑部疾病诊断

Junhao Jia, Yifei Sun, Yunyou Liu, Cheng Yang, Changmiao Wang, Feiwei Qin, Yong Peng, Wenwen Min

机构 * Hangzhou Dianzi University(杭州电子科技大学) Zhejiang University(浙江大学) Shenzhen Research Institute of Big Data(深圳大数据研究院) Yunnan University(云南大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 RTGMFF通过结合ROI驱动文本生成和多模态特征融合,提升fMRI在脑部疾病诊断中的准确性。

Comments The paper has been accepted by BIBM 2025

Journal ref 2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). IEEE, 2025, pp. 2301-2308

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22283 2026-03-03 cs.CV 79%

Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment

重新审视在跨模态不匹配下的LVLMs视觉标记减少

Rui Xu, Yunke Wang, Yong Luo, Bo Du

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出VisionDrop方法,通过视觉-only修剪框架减少LVLMs中的视觉标记,无需额外训练,提升推理效率并保持性能。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00157 2026-03-03 cs.CV 79%

FujiView: Multimodal Late-Fusion for Predicting Scenic Visibility

FujiView: 多模态晚期融合用于预测风景可见性

Bryceton Bible, Shah Md Nehal Hasnaeen, Hairong Qi

机构 * University of Tennessee, Knoxville(田纳西大学,科文克顿)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 FujiView通过融合摄像头图像与气象数据,实现风景可见性的多模态预测,展示了在短期和长期预测中的不同方法效果。

Comments 9 pages (including references), 8 figures, 2 tables. Accepted to the IEEE/CVF WACV 2026 proceedings. Introduces a large human-labeled Mount Fuji visibility dataset; public release forthcoming

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24041 2026-03-02 cs.CV 79%

Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation

仔细观察:多模态大语言模型中的自适应视觉增强以缓解幻觉

Xingyu Zhu, Kesen Zhao, Liang Yi, Shuo Wang, Zhicai Wang, Beier Zhu, Hanwang Zhang

机构 * MoE Key Lab of BIPC, University of Science and Technology of China(信息与电子技术联合实验室,中国科学技术大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究提出自适应视觉增强框架AIR,通过减少冗余标记和选择性整合补丁来缓解多模态大语言模型中的幻觉问题。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04576 2026-03-02 cs.CV 79%

TARDis: Time Attenuated Representation Disentanglement for Incomplete Multi-Modal Tumor Segmentation and Classification

TARDis: 用于不完整多模态肿瘤分割与分类的时间衰减表示解耦

Zishuo Wan, Qinqin Kang, Na Li, Yi Huang, Qianru Zhang, Le Lu, Yun Bian, Dawei Ding, Ke Yan

机构 * School of Automation and Electrical Engineering, University of Science and Technology Beijing(北京科技大学自动化与电气工程学院) Alibaba Group DAMO Academy(阿里巴巴集团DAMO学院) Hupan Lab(湖畔实验室) Departments of Radiology, Changhai Hospital(上海长海医院放射科)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 TARDis通过时间衰减表示解耦框架,解决多模态肿瘤分割与分类中缺失模态问题,提升诊断精度并降低辐射暴露。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17966 2026-03-02 cs.IR cs.CV 79%

LLM-Enhanced Multimodal Fusion for Cross-Domain Sequential Recommendation

基于大语言模型的跨域序列推荐多模态融合

Wangyu Wu, Zhenhong Chen, Wenqiao Zhang, Xianglin Qiu, Siqi Song, Xiaowei Huang, Fei Ma, Jimin Xiao

机构 * Xi'an Jiaotong-Liverpool University(西安交通大学利物浦大学) University of Liverpool(利物浦大学) Microsoft(微软公司) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出LLM-EMF方法,通过融合视觉和文本数据提升跨域序列推荐性能,利用CLIP模型生成多模态嵌入并引入多重注意力机制,实验证明其在多领域推荐中的有效性。

Comments arXiv admin note: substantial text overlap with arXiv:2504.15085

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01728 2026-03-02 cs.CV 79%

Shuffle Mamba: State Space Models with Random Shuffle for Multi-Modal Image Fusion

Shuffle Mamba:基于随机洗牌的态空间模型用于多模态图像融合

Ke Cao, Xuanhua He, Tao Hu, Chengjun Xie, Man Zhou, Jie Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Institute of Intelligent Machines(智能机器研究所) Hefei Institutes of Physical Science, Chinese Academy of Sciences(中国科学院合肥物质科学研究院) Intelligent Agriculture Engineering Laboratory of Anhui Province, Institute of Intelligent Machines(安徽省智能农业工程实验室,智能机器研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 Shuffle Mamba通过引入随机洗牌策略和逆洗牌,解决多模态图像融合中固定扫描策略带来的偏见问题,提升融合质量。

Comments Accepted by IEEE Transactions on Circuits and Systems for Video Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22236 2026-02-27 q-bio.GN cs.CV cs.LG 79%

CrossLLM-Mamba: Multimodal State Space Fusion of LLMs for RNA Interaction Prediction

CrossLLM-Mamba: LLMs多模态状态空间融合用于RNA相互作用预测

Rabeya Tus Sadia, Qiang Ye, Qiang Cheng

机构 * Department of Computer Science, University of Kentucky, Lexington, KY, USA(计算机科学系,肯塔基大学,路易斯维尔,KY,美国) Department of Mathematics, University of Kentucky, Lexington, KY, USA(数学系,肯塔基大学,路易斯维尔,KY,美国) Institute for Biomedical Informatics, University of Kentucky, Lexington, KY, USA(生物医学信息学研究所,肯塔基大学,路易斯维尔,KY,美国)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV

AI总结 CrossLLM-Mamba通过多模态状态空间融合实现RNA相互作用预测,采用双向Mamba编码器和动态序列转换模型,达到高精度性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20723 2026-02-27 cs.AI 79%

Modality-Guided Mixture of Graph Experts with Entropy-Triggered Routing for Multimodal Recommendation

模态引导的图专家混合网络与熵触发路由用于多模态推荐

Ji Dai, Quan Fang, Dengsheng Cai

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tianjin University of Technology(天津理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 MAGNET通过模态引导的图专家混合网络与熵触发路由,提升多模态推荐中融合的可控性、稳定性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21952 2026-02-26 cs.CV 79%

MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous Driving

MindDriver: 引入渐进多模态推理用于自动驾驶

Lingjun Zhang, Yujian Yuan, Changjie Wu, Xinyuan Chang, Xin Cai, Shuang Zeng, Linzhe Shi, Sijin Wang, Hang Zhang, Mu Xu

机构 * Amap, Alibaba Group(阿里集团蚂巴公司) The Hong Kong University of Science and Technology(香港科学与技术大学) The Chinese University of Hong Kong(香港中文大学) Xi’an Jiaotong University(西安交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 MindDriver通过渐进多模态推理框架提升自动驾驶系统的推理能力,实现语义到物理空间的转化与轨迹规划,展现优异的性能表现。

Comments CVPR2026; Yujian Yuan and Lingjun Zhang contributed equally with random order

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21154 2026-02-25 cs.AI 79%

CG-DMER: Hybrid Contrastive-Generative Framework for Disentangled Multimodal ECG Representation Learning

CG-DMER:一种用于解耦多模态ECG表示学习的对比-生成框架

Ziwei Niu, Hao Sun, Shujun Bian, Xihong Yang, Lanfen Lin, Yuxin Liu, Yueming Jin

机构 * Department of Biomedical Engineering, National University of Singapore(新加坡国立大学生物医学工程系) Department of Electrical and Computer Engineering, National University of Singapore(新加坡国立大学电气与计算机工程系) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) College of information science and engineering, Ritsumeikan university(立命馆大学信息科学与工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 CG-DMER通过对比-生成框架解耦多模态ECG表示,提升ECG信号在心血管疾病诊断中的建模能力。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02276 2026-02-25 cs.AI 79%

BioX-Bridge: Model Bridging for Unsupervised Cross-Modal Knowledge Transfer across Biosignals

BioX-Bridge:生物信号跨模态知识迁移的模型桥接

Chenqi Li, Yu Liu, Timothy Denison, Tingting Zhu

机构 * Department of Engineering Science University of Oxford(工程科学系 奥克大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.AI

AI总结 BioX-Bridge通过轻量级桥接网络实现生物信号跨模态无监督知识迁移,显著减少可训练参数并提升迁移性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23479 2026-02-24 cs.CV 79%

MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding

MergeMix:一种用于视觉和多模态理解的统一增强范式

Xin Jin, Siyuan Li, Siyong Jian, Kai Yu, Huan Wang

机构 * Westlake University(西拉克大学) Zhejiang University(浙江大学) College of Computer Science and Technology(计算机科学与技术学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 MergeMix通过高效的令牌合并Mixup增强方法,平衡了SFT和RL的优劣,提升多模态大语言模型的分类准确率和对齐能力。

Comments ICLR 2026, Web link: https://jinxins.github.io/MergeMix_Web/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19702 2026-02-24 cs.IR cs.AI 79%

DReX: An Explainable Deep Learning-based Multimodal Recommendation Framework

DReX:一种基于深度学习的可解释多模态推荐框架

Adamya Shyam, Venkateswara Rao Kagita, Bharti Rana, Vikas Kumar

机构 * University of Delhi, Delhi, India(德里大学) National Institute of Technology, Warangal, India(印度战争格尔国家理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 DReX通过多模态反馈的交互级特征逐步细化用户和物品表示,提升推荐系统的可解释性和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19140 2026-02-24 cs.CV cs.LG 79%

CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusion

CaReFlow:循环适应修正流用于多模态融合

Sijie Mai, Shiqin Han

机构 * South China Normal University(华南师范大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 CaReFlow通过循环适应修正流实现多模态融合,解决模态间隙问题,提升分布对齐和特征转换的鲁棒性。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19058 2026-02-24 cs.CL 79%

Do LLMs and VLMs Share Neurons for Inference? Evidence and Mechanisms of Cross-Modal Transfer

大型语言模型和视觉语言模型是否共享神经元进行推理?证据和跨模态转移的机制

Chenhang Cui, An Zhang, Yuxin Chen, Gelei Deng, Jingnan Zheng, Zhenkai Liang, Xiang Wang, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) University of Science and Technology of China(中国科学技术大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态训练与对齐 :cross-modal(title);multimodal(abstract);分类 cs.CL

AI总结 本文研究了大型语言模型和视觉语言模型在推理过程中共享神经元的现象,提出SNRF框架实现低秩融合,提升多模态推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06450 2026-02-24 cs.CV cs.LG 79%

Countering Multi-modal Representation Collapse through Rank-targeted Fusion

通过秩目标融合对抗多模态表示崩溃

Seulgi Kim, Kiran Kokilepersaud, Mohit Prabhushankar, Ghassan AlRegib

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出Rank-enhancing Token Fuser框架,通过提升有效秩对抗多模态表示崩溃,验证了深度与RGB融合的平衡性,并在动作预测任务中取得显著性能提升。

Comments Accepted in 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10652 2026-02-24 cs.CL 79%

ViTextVQA: A Large-Scale Visual Question Answering Dataset and a Novel Multimodal Feature Fusion Method for Vietnamese Text Comprehension in Images

ViTextVQA: 一个大规模视觉问答数据集和一种新的多模态特征融合方法用于图像中的越南语文本理解

Quan Van Nguyen, Dan Quang Tran, Huy Quang Pham, Thang Kien-Bao Nguyen, Nghia Hieu Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

机构 * Faculty of Information Science and Engineering, University of Information Technology(信息科学与工程学院,信息科技大学) Vietnam National University(越南国家大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 ViTextVQA是一个大规模视觉问答数据集,提出了一种新的多模态特征融合方法,用于提升图像中越南语文本的理解能力。

Comments International Journal of Expert Systems with Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18752 2026-02-24 cs.CV cs.GR 79%

Optimizing ID Consistency in Multimodal Large Models: Facial Restoration via Alignment, Entanglement, and Disentanglement

多模态大模型中ID一致性优化:通过对齐、纠缠与解纠缠进行面部修复

Yuran Dong, Hang Dai, Mang Ye

机构 * National Engineering Research Center for Multimedia Software(多媒体软件国家工程研究中心) School of Computer Science, Wuhan University(武汉大学计算机学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 EditedID通过引入对齐、解纠缠和纠缠机制,提升多模态大模型中面部身份一致性的修复能力。

Comments ICLR 26

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15162 2026-02-20 eess.SP cs.AI cs.LG 79%

Multimodal Wireless Foundation Models

多模态无线基础模型

Ahmed Aboulfotouh, Hatem Abou-Zeid

机构 * Department of Electrical and Software Engineering(电气与软件工程系) University of Calgary(卡尔加里大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出首个多模态无线基础模型,能处理IQ流和图像类无线模态,支持多种无线任务,展示了其在不同模态上的广泛应用潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16245 2026-02-19 cs.CV 79%

HyPCA-Net: Advancing Multimodal Fusion in Medical Image Analysis

HyPCA-Net:在医学图像分析中推进多模态融合

J. Dhar, M. K. Pandey, D. Chakladar, M. Haghighat, A. Alavi, S. Mistry, N. Zaidi

机构 * Indian Institute of Technology Ropar(印度理工学院罗帕尔分校) RoentGen Health(RoentGen健康公司) Lulea University of Technology(卢勒奥大学) QUT(昆士兰科技大学) RMIT University(皇家墨尔本理工大学) Curtin University(Curtin大学) Deakin University(德肯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 HyPCA-Net通过高效残差注意力模块和双视角级联注意力模块,提升多模态医学图像分析的性能与效率。

Comments Accepted at the IEEE/CVF Winter Conference on Applications of Computer Vision 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15896 2026-02-19 cs.CL 79%

Every Little Helps: Building Knowledge Graph Foundation Model with Fine-grained Transferable Multi-modal Tokens

每一点帮助都有价值:通过细粒度可迁移多模态标记构建知识图谱基础模型

Yichi Zhang, Zhuo Chen, Lingbing Guo, Wen Zhang, Huajun Chen

机构 * Zhejiang University, Zhejiang, China(浙江大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

AI总结 本文提出TOFU模型,通过细粒度可迁移多模态标记提升多模态知识图谱推理的跨图谱迁移能力。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15740 2026-02-18 cs.LG cs.AI q-bio.QM 79%

MRC-GAT: A Meta-Relational Copula-Based Graph Attention Network for Interpretable Multimodal Alzheimer's Disease Diagnosis

MRC-GAT:一种基于元关系耦合的图注意力网络用于可解释的多模态阿尔茨海默病诊断

Fatemeh Khalvandi, Saadat Izadi, Abdolah Chalechale

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 MRC-GAT通过元关系耦合和多关系注意力机制,实现多模态阿尔茨海默病诊断的高精度与可解释性。

Comments 27 pages, 10 figures, 10 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00168 2026-02-18 cs.CV 79%

SSL4EO-S12 v1.1: A Multimodal, Multiseasonal Dataset for Pretraining, Updated

SSL4EO-S12 v1.1:一种用于预训练的多模态、多季节数据集,更新版

Benedikt Blumenstiel, Nassim Ait Ali Braham, Conrad M Albrecht, Stefano Maurogiovanni, Paolo Fraccaro

机构 * IBM Research Europe(IBM欧洲研究中心) German Aerospace Center(德国航空航天中心) Julich Supercomputing Centre University of Iceland(朱利奇超级计算中心爱沙尼亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 SSL4EO-S12 v1.1通过增加多模态数据和改进数据结构,为预训练大规模基础模型提供了更高效、更全面的地球观测数据集。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15461 2026-02-18 cs.CV 79%

Emergent Morphing Attack Detection in Open Multi-modal Large Language Models

开放多模态大语言模型中的涌现形变攻击检测

Marija Ivanovska, Vitomir Štruc

机构 * Faculty of Electrical Engineering, University of Ljubljana(卢布尔雅那大学电气工程学院)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

AI总结 本文提出利用开源多模态大语言模型进行零样本人脸形变攻击检测,通过实验表明其在无微调情况下具备优异的判别能力,性能超越传统基线23%。

Comments This manuscript is currently under review at Pattern Recognition Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15346 2026-02-18 cs.CV 79%

Effective and Robust Multimodal Medical Image Analysis

高效且鲁棒的多模态医学图像分析

Joy Dhar, Nayyar Zaidi, Maryam Haghighat

机构 * Indian Institute of Technology Ropar(印度理工学院罗帕尔分校) Deakin University(德克萨斯大学) Queensland University of Technology(昆士兰理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出MAIL网络,通过高效残差学习和多模态交叉注意力模块,提升多模态医学图像分析的效率与鲁棒性,实现实验性能提升9.34%并降低计算成本78.3%。

Comments Accepted at Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14518 2026-02-17 cs.AI 79%

Diagnosing Knowledge Conflict in Multimodal Long-Chain Reasoning

多模态长链推理中的知识冲突诊断

Jing Tang, Kun Wang, Haolang Lu, Hongjin Chen, KaiTao Chen, Zhongxiang Sun, Qiankun Li, Lingjuan Lyu, Guoshun Nan, Zhigang Zeng

机构 * Huazhong University of Science and Technology(华中科技大学) Nanyang Technological University(南洋理工大学) Beijing University of Posts and Telecommunications(北京邮电大学) Renmin University of China(中国人民大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本研究提出了一种统一的知识冲突概念,揭示了多模态长链推理中不同冲突类型的特征和处理机制,为诊断和控制推理失败提供了原理性方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10685 2026-02-17 cs.CV 79%

GaussianFormer3D: Multi-Modal Gaussian-based Semantic Occupancy Prediction with 3D Deformable Attention

GaussianFormer3D: 多模态基于高斯的语义占用预测与3D可变形注意力

Lingjun Zhao, Sizhe Wei, James Hays, Lu Gan

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 GaussianFormer3D通过3D可变形注意力机制,结合激光雷达与摄像头数据,实现高效且精确的语义占用预测。

详情

展开后加载摘要…

URL PDF HTML 收藏