arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-05-26 至 2026-05-26 共收录 39 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 39 篇

2601.21670 2026-05-26 cs.CV cs.LG 87%

Diverse via bounded Agreement: Geometric Regularization for Multimodal Fusion

通过有界一致性实现多样性:多模态融合的几何正则化

Zixuan Xia, Hao Wang, Pengcheng Weng, Yanyu Qian, Yangxin Xu, William Dan, Fei Wang

机构 * Department of Informatics University of Bern(伯尔尼大学信息学院) College of Computing and Data Science Nanyang Technological University(南洋理工大学计算机与数据科学学院) School of Software Engineering Xi’an Jiaotong University(西安交通大学软件工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);audio-visual(abstract)

AI总结 提出一种轻量级即插即用的几何正则化框架,通过有界一致性原则在保持模态特异多样性的同时约束跨模态漂移,提升多模态融合性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26110 2026-05-26 cs.LG cs.CL cs.CV 86%

Prism: A Plug-in Reproducible Infrastructure for Scalable Multimodal Continual Instruction Tuning

Prism:面向可扩展多模态持续指令微调的插件式可复现基础设施

Jun-Tao Tang, Yu-Cheng Shi, Zhen-Hao Xie, Da-Wei Zhou

机构 * School of Artificial Intelligence, Nanjing University, China(南京大学人工智能学院) National Key Laboratory for Novel Software Technology, Nanjing University, China(南京大学新型软件技术国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.CL

AI总结 针对多模态持续指令微调中工程瓶颈问题,提出Prism插件式代码库,通过轻量级插件注册机制分离算法开发与骨干实现,支持大规模训练流水线,实现可复现、可扩展的实验。

Comments Code is available at https://github.com/LAMDA-CL/Prism

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12374 2026-05-26 cs.CV cs.AI cs.LG 86%

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models

填补GAP:多模态大语言模型中视觉推理的粒度对齐范式

Yanting Miao, Yutao Sun, Dexin Wang, Mengyu Zhou, Pascal Poupart, Lei Lv, Qi Zhao, Li Wang, Hao Li, Xiaoxi Jiang, Guanjun Jiang

机构 * Qwen Large Model Application Team, Alibaba(阿里云大模型应用团队) Alibaba University of Waterloo(阿里大学水力学院) Vector Institute(向量研究所) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 提出GAP(粒度对齐范式),通过特征级、上下文级和能力引导级对齐,解决多模态大语言模型中视觉潜在推理的特征空间不匹配问题,提升感知与推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00857 2026-05-26 cs.LG cs.AI 85%

MultiPUFFIN: A Multimodal Domain-Constrained Foundation Model for Molecular Property Prediction of Small Molecules

MultiPUFFIN:用于小分子性质预测的多模态领域约束基础模型

Idelfonso B. R. Nogueira, Carine M. Rebello, Mumin Enis Leblebici, Erick Giovani Sperandio Nascimento

机构 * Department of Chemical Engineering, Norwegian University of Science and Technology (NTNU)(挪威科学与技术大学化学工程系) Faculty of Industrial Engineering, KU Leuven(鲁文大学工业工程学院) University of Surrey(萨里大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 提出多模态基础模型MultiPUFFIN,融合SMILES、2D图、3D构象及实验条件,通过条件感知精炼和热力学约束头,在小样本下优于ChemBERTa-2,预测小分子热物理性质。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00682 2026-05-26 cs.IR cs.AI 83%

RecGOAT: Graph Optimal Adaptive Transport for LLM-Enhanced Multimodal Recommendation with Dual Semantic Alignment

RecGOAT: 用于LLM增强多模态推荐的图最优自适应传输与双语义对齐

Yuecheng Li, Hengwei Ju, Zeyu Song, Wei Yang, Chi Lu, Peng Jiang, Kun Gai

机构 * Fudan University(复旦大学) University of Southern California(南加州大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 针对生成式语言模型表示与ID协同信号之间的语义异质性,提出基于图神经网络和最优传输理论的双粒度语义对齐框架RecGOAT,通过实例级和分布级对齐实现统一特征空间,理论证明其表示误差更低,实验达到最优性能。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10012 2026-05-26 cs.LG 82%

PID-Guided Partial Alignment for Multimodal Decentralized Federated Learning

PID引导的多模态去中心化联邦学习部分对齐

Yanhang Shi, Xiaoyu Wang, Houwei Cao, Jian Li, Yong Liu

机构 * Department of Electrical and Computer Engineering, Stony Brook University(石溪大学电气与计算机工程系) Department of Applied Mathematics and Statistics and the Department of Computer Science, Stony Brook University(石溪大学应用数学与统计系和计算机科学系) Department of Electrical and Computer Engineering, New York University(纽约大学电气与计算机工程系) Department of Computer Science, New York Institute of Technology(纽约理工学院计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 针对多模态去中心化联邦学习中异构代理间更新不兼容的问题,提出基于部分信息分解的PARSE框架,通过特征分裂和部分对齐实现高效通信与协作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17044 2026-05-26 cs.LG cs.AI cs.CV 82%

Do Understanding and Generation Fight? A Diagnostic Study of DPO for Unified Multimodal Models

理解与生成相冲突吗?统一多模态模型DPO的诊断研究

Abinav Rao, Sujan Rachuri

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 通过系统实验发现,在统一多模态模型上应用DPO时,生成质量难以对齐,主要原因是理解和生成梯度近乎正交且存在11-14倍的幅度不平衡,源于VQ token数量不对称。

Comments Experiments are inconclusive: The claim that architectures such as Chameleon or Emu would exhibit stronger gradient conflict is not supported by experiments or analysis, and all experiments are conducted on Janus-Pro without evaluation on other unified multimodal architectures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26026 2026-05-26 cs.CV cs.AI cs.LG 81%

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

一种用于光片荧光显微镜的多模态3D基础模型实现少样本分割、分类和去模糊

Adina Scheinfeld, Haotan Zhang, Shang Mu, Rudolf L. M. van Herten, Lucas Stoffl, Ali Erturk, Zhuhao Wu, Johannes C. Paetzold

机构 * Tri-Institutional Program in Computational Biology \& Medicine, Weill Cornell Medicine, New York, NY, USA Department of Radiology, Weill Cornell Medicine, New York, NY, USA Helen Robert Appel Alzheimers Disease Research Institute, Feil Family Brain Mind Research Institute, Weill Cornell Medicine, New York, NY, USA Graduate Program in Physiology, Biophysics Systems Biology, Weill Cornell Medicine, New York, NY, USA Cornell Tech, New York, NY, USA Institute for Intelligent Biotechnologies (iBIO), Helmholtz Center Munich, Neuherberg, Germany Institute for Stroke Dementia Research, Klinikum der Universität München, Ludwig-Maximilians University Munich, Munich, Germany

专题命中 多模态训练与对齐 :multimodal(title);image-text(abstract);分类 cs.CV、cs.AI

AI总结 提出一种基于掩码重建与图像-文本对齐联合优化的3D基础模型,在光片荧光显微镜数据上预训练,通过少样本适应显著降低标注成本并提升分割、分类和去模糊性能。

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26004 2026-05-26 cs.CV cs.CL 81%

MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models

MAGIC: 面向视觉语言模型的多模态对齐与接地感知指令核心集

Shristi Das Biswas, Kaushik Roy

机构 * Purdue University(普渡大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 提出MAGIC方法,利用预训练VLM中的多模态增益、桥接相关性和技能神经元签名三种内在信号,通过无训练、前向传播的核心集选择,构建紧凑且行为保真的子集用于多模态指令微调,在20%预算下达到甚至超越全微调性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25952 2026-05-26 cs.CV cs.AI 81%

VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding

VEN-VL: 一种用于高效多模态理解的视觉集成MoE框架

Yinghao Wu, Zhuoyan Luo, Yiyao Yu, Zhaojian Yu, Yujiu Yang, Xiao-Ping Zhang

机构 * Tsinghua University(清华大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出VEN-VL框架,通过先丰富后压缩的策略,利用视觉集成MoE和自适应路由增强视觉令牌的信息容量与密度,在少量压缩令牌下实现复杂视觉任务的性能与效率平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25328 2026-05-26 cs.CV cs.MM 81%

DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement

DIVA: 利用统一多模态模型中的表示差异实现相互增强

Renjie Lu, Xulong Zhang, Xiaoyang Qu, Shangfei Wang, Jianzong Wang

机构 * Ping An Technology (Shenzhen) Co., Ltd., Shenzhen, China(平安科技(深圳)有限公司,深圳,中国) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.MM

AI总结 针对统一多模态模型中理解与生成任务因监督信号差异导致相互干扰的问题,提出DIVA框架,通过分解视觉表示为共享和独有成分并利用互信息估计实现内部协同,在理解与生成任务上分别提升7.82%和8.46%。

Comments Accepted to the 43rd International Conference on Machine Learning (ICML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24799 2026-05-26 cs.CV cs.AI 81%

Divide-and-Conquer Inference for Large-Scale Visual Recognition with Multimodal Large Language Models

面向大规模视觉识别的多模态大语言模型分治推理

Zhipeng Ye, Jiaqi Huang, Feng Jiang, Qiufeng Wang, Yikang Duan, Dawei Wang, Xihang Zhou, Qian Qiao

机构 * Taizhou Institute of Science and Technology, Nanjing University of Science and Technology(泰州科技学院、南京理工大学) Department of Intelligence Science, Xi’an Jiaotong-Liverpool University(智能科学系,西安交通大学利物浦大学) School of Computer Science and Technology, Soochow University(计算机科学与技术学院,苏州大学) Department of Statistical Sciences, University of Toronto(统计科学系,多伦多大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 针对多模态大语言模型在长序列识别中性能崩溃的问题,提出分治推理(DCI)策略,通过递归分解任务和动态剪枝提升信噪比与分类精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25571 2026-05-26 cs.CV 79%

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution

AnE: 通过锚点进化推动多模态大语言模型的推理前沿

Zehao Wang, Yihan Zeng, Zidong Gong, Yuanfan Guo, Feng Zhu, Hongzhi Zhang, Wei Zhang, Wangmeng Zuo

机构 * Harbin Institute of Technology(哈尔滨工业大学) Huawei Noah's Ark Lab(华为诺亚实验室) Independent Researcher(独立研究员)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 提出锚点进化(AnE)范式,通过真值锚点数据策展和脚手架剥离机制,解决多模态大模型推理中的认知漂移和幻觉路径问题,显著提升推理性能。

Comments 34 pages,10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25262 2026-05-26 cs.CV 79%

Semantics-Guided Multimodal Masked Autoencoder Pretraining for 3D BEV Object Detection

语义引导的多模态掩码自编码器预训练用于3D BEV目标检测

Prabuddhi Wariyapperuma, Rajitha de Silva, Marc Hanheide, Thomas Bohné, Leonardo Guevara

机构 * University of Lincoln, Lincoln Centre for Autonomous Systems(林肯大学,林肯自主系统中心) University of Cambridge, Institute for Manufacturing, Department of Engineering(剑桥大学,制造研究所,工程系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 提出语义引导的多模态掩码自编码器框架,通过语义引导的LiDAR体素掩码和辅助点语义解码分支,在预训练中注入语义信息,提升3D BEV目标检测性能。

Comments Accepted at the ICRA 2026 Workshop on Semantics for Reliable Robot Autonomy (SRRA) as a lightning talk and poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23076 2026-05-26 cs.LG cs.AI cs.HC 79%

Multimodal Functional Maximum Correlation for Emotion Recognition

多模态功能最大相关用于情感识别

Deyang Zheng, Tianyi Zhang, Wenming Zheng, Shujian Yu

机构 * Key Laboratory of Child Development and Learning Science (Ministry of Education), School of Biological Sciences and Medical Engineering, Southeast University(儿童发展与学习科学重点实验室(教育部)、生物科学与医学工程学院、东南大学) Department of Artificial Intelligence, Westlake University(人工智能学院、西湖大学) Department of Artificial Intelligence, Vrije Universiteit Amsterdam(人工智能学院、阿姆斯特丹自由大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 提出多模态功能最大相关(MFMC)框架,通过双重总相关目标最大化高阶多模态依赖,在情感识别基准上取得最先进性能。

Comments manuscript accepted by IEEE Transactions on Affective Computing. Code is available at https://github.com/DY9910/MFMC

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25288 2026-05-26 cs.IT cs.AI cs.ET cs.LG eess.SP math.IT 79%

CSI-tuples-based 3D Channel Fingerprints Construction Assisted by MultiModal Learning

基于CSI元组的多模态学习辅助3D信道指纹构建

Chenjie Xie, Li You, Ruirong Chen, Gaoning He, Xiqi Gao

机构 * National Mobile Communications Research Laboratory, Southeast University(东南大学国家移动通信研究中心) Purple Mountain Laboratories(紫金山实验室) Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 针对低空通信中的3D信道指纹构建问题,提出一种基于CSI元组的多模态回归框架,通过融合位置、通信测量和地理环境地图,实现高效高精度的信道状态信息估计。

Comments 14 pages, 9 figures

Journal ref IEEE Transactions on Wireless Communications, vol. 25, pp. 17369-17383, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10548 2026-05-26 cs.CV 79%

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

Blink: 动态视觉令牌分辨率增强多模态理解

Yuchen Feng, Zhenyu Zhang, Naibin Gu, Yilong Chen, Peng Fu, Zheng Lin, Shuohuan Wang, Yu Sun, Hua Wu, Weiping Wang, Haifeng Wang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Baidu Inc(百度公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 提出Blink框架,通过注意力引导的令牌超分辨率和动态丢弃机制,在单次前向传播中模拟人类眨眼式扫描,提升多模态大语言模型的视觉感知能力。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23961 2026-05-26 q-bio.BM cs.AI cs.LG 79%

Multimodal Alignment and Preference Optimization for Zero-Shot Conditional RNA Generation

多模态对齐与偏好优化用于零样本条件RNA生成

Roman Klypa, Alberto Bietti, Sergei Grudinin

机构 * Univ. Grenoble Alpes, CNRS, Grenoble INP, LJK(格勒诺布尔阿尔卑斯大学、法国国家科学研究中心、格勒诺布尔INP、LJK实验室) Center for Computational Mathematics, Flatiron Institute(计算数学中心、Flatiron研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 提出Moirain框架,通过多模态监督微调和直接偏好优化实现条件RNA序列生成,在零样本条件下生成具有高结合亲和力的生物合理RNA序列。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25726 2026-05-26 cs.IR 78%

SIREN: Unified Multi-Granularity Semantic Interaction for Multi-Modal Lifelong User Interest Modeling

SIREN:面向多模态终身用户兴趣建模的统一多粒度语义交互

Yaqian Zhang, Ruyi Yu, Tianyi Li, Bohan Liu, Maoquan Ye, Ke Wang, Shifeng Wen, Junwei Pan, Lijie Wang, Qi Zhou, Yeshou Cai, Chengguo Yin, Lifeng Wang, Hui Li, Lei Xiao, Haijie Gu

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

AI总结 提出SIREN框架,通过统一多粒度语义交互(包括基于多模态相似性的软检索和基于语义ID的硬检索)解决多模态与协同空间的对齐问题,在离线数据集上达到最优GAUC,并在腾讯广告平台多个场景中取得GMV提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25540 2026-05-26 cs.SD cs.LG 78%

A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

基于语言和声学表征学习的多模态痴呆检测框架

Loukas Ilias, Dimitris Askounis

机构 * Decision Support Systems Laboratory, School of Electrical and Computer Engineering, National Technical University of Athens(决策支持系统实验室,电气与计算机工程学院,国家技术大学雅典)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 提出一个端到端可训练的多模态深度学习框架,通过预训练模型提取声学和文本特征,结合注意力融合与互信息最大化,实现自动痴呆检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25161 2026-05-26 physics.geo-ph 78%

FMSIM: A Multimodal Flow Matching Framework for Conditional Geomodeling

FMSIM: 一种用于条件地质建模的多模态流匹配框架

Jiayuan Huang, Suihong Song, Tapan Mukerji

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract)

AI总结 提出FMSIM多模态条件流匹配框架,通过深度学习速度场、语义表示、迭代投影和时间引导门控机制,融合概念地质描述、稀疏井观测和空间先验约束,实现高效稳定的地下相模型生成。

Comments Preprint version of a manuscript submitted to Water Resources Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20772 2026-05-26 cs.CV 77%

VIHD: Visual Intervention-based Hallucination Detection for Medical Visual Question Answering

VIHD: 基于视觉干预的医学视觉问答幻觉检测

Jiayi Chen, Benteng Ma, Zehui Liao, Winston Chong, Yasmeen George, Jianfei Cai

机构 * Department of Data Science \& AI, Faculty of Information Technology, Monash University, Melbourne, VIC 3800, Australia Alfred Health Radiology, Alfred Health, Melbourne, VIC 3004, Australia School of Translational Medicine, Faculty of Medicine, Nursing Health Sciences, Monash University, Melbourne, VIC 3800, Australia Hong Kong Polytechnic University, Hong Kong SAR, China

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract_cn);cross-modal(abstract);分类 cs.CV

AI总结 提出VIHD方法,通过视觉依赖探测和视觉干预解码校准语义熵,有效检测医学多模态大语言模型中的幻觉响应。

Comments Early accepted by MICCAI 2026. This version of the contribution has been accepted for publication, after peer review (when applicable) but is not the Version of Record and does not reflect post-acceptance improvements, or any corrections

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24532 2026-05-26 cs.CV 70%

Image-Conditioned Instance Prompt Network for Referring Remote Sensing Image Segmentation

图像条件实例提示网络用于遥感图像指代分割

Biaoyu Ren, Qingsheng Wang, Cun Xu, Dingkang Yang, Wenxuan Wang

机构 * School of Computer Science, Northwestern Polytechnical University, Xi'an, China(西北工业大学计算机科学学院,西安,中国) College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai, China(复旦大学智能机器人与先进制造学院,上海,中国) Shenzhen Research Institute of Northwestern Polytechnical University, Shenzhen, China(西北工业大学深圳研究院,深圳,中国)

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 提出图像条件实例提示网络(ICIPNet),通过自适应视觉语义表示和双边信息融合模块,缓解跨模态特征融合瓶颈,提升遥感图像指代分割性能。

Comments 6 pages, 3 figures. Equal contribution: Biaoyu Ren and Qingsheng Wang. Corresponding authors: Dingkang Yang and Wenxuan Wang

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01328 2026-05-26 cs.CV 70%

CSFNet: A Cosine Similarity Fusion Network for Real-Time RGB-X Semantic Segmentation of Driving Scenes

CSFNet: 用于驾驶场景实时RGB-X语义分割的余弦相似度融合网络

Danial Qashqai, Emad Mousavian, Shahriar Baradaran Shokouhi, Sattar Mirzakuchaki

机构 * Department of Electrical Engineering, Iran University of Science and Technology(伊朗科学技术大学电气工程系)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出CSFNet,通过余弦相似度注意力融合模块(CS-AFM)高效融合双模态特征,实现实时且高精度的RGB-X语义分割。

Journal ref Engineering Applications of Artificial Intelligence, 174, 114362 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24675 2026-05-26 cs.CV cs.AI 62%

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation

VaaWIT: 面向多语言网页图像翻译的大语言模型视觉感知适配

Bo Li, Ronghao Chen, Ningyuan Deng, Huacan Wang, Shaolin Zhu, Lijie Wen

机构 * The Hong Kong University of Science(香港科技大学) Tianjin University(天津大学) Tsinghua University(清华大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 针对网页图像翻译中视觉表示差距问题,提出VaaWIT框架,通过双流注意力模块和视觉感知适配器,实现大语言模型对细粒度视觉特征的动态融合,在多个基准上超越开源模型并接近闭源模型性能。

Comments Accepted by KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24639 2026-05-26 cs.CV cs.AI 62%

DisDop: Distillation with Domain Priors for Open-Vocabulary Aerial Object Detection

DisDop: 基于领域先验蒸馏的开放词汇航空目标检测

Ruihao Xu, Yong Liu, Yansong Tang, Sule Bai, Xubing Ye, Bingyao Yu, Yutao Guo, Jiwen Lu, Jie Zhou

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Tsinghua University(清华大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 提出DisDop框架,通过从遥感基础模型(RemoteCLIP和DINOv3)中系统蒸馏多级领域先验知识到轻量级检测器,实现开放词汇航空目标检测的最新性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21417 2026-05-26 cs.CV cs.AI 62%

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition

排序重要:面向混合情感识别的排名感知选择性融合

Junghyun Lee, Hyunseo Kim, Hanna Jang, Junhyug Noh

机构 * Department of Artificial Intelligence and Software(人工智能与软件系)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出一种排名感知的多编码器框架,通过注意力门控模块选择最有效的编码器进行融合,并解耦预测为存在性和显著性头,结合无监督域适应,在混合情感识别任务中取得第二名成绩。

Comments Accepted at IEEE FG 2026 Workshops. Final system ranked 2nd in the BlEmoRE Challenge. 9 pages including appendix, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23997 2026-05-26 cs.CV cs.AI cs.LG 62%

IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning

IVR-R1:通过强化学习中的迭代视觉基础推理优化轨迹

Chenghao Li, Fusheng Hao, Xikai Zhang, Likang Xiao, Yanwei Ren, Fuxiang Wu, Quan Chen, Liu Liu

机构 * Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院) School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) Kuaishou Technology(快手科技) Shenzhen Institute of Advanced Integration Technology, Shenzhen(深圳先进集成技术研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出IVR-R1框架,利用奖励驱动的筛选机制和迭代再推理循环,在强化学习中动态校正多模态推理轨迹,以解决视觉幻觉和逻辑错误问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25922 2026-05-26 cs.CV 57%

Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models

闭环双向提示用于视觉语言模型的对抗鲁棒性

Xiao Liu, Jiaxiang Liu, Boci Peng, Boren Hu, Yusong Wang, Xiwen Chen, Prayag Tiwari, Liming Zhang, Mingkun Xu

机构 * University of Macau(澳门大学) Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) Peking University(北京大学) Independent Researcher(独立研究员) Institute of Science Tokyo(东京科学研究院) Morgan Stanley(摩根大通) Halmstad University(哈马碧大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 针对视觉语言模型在对抗扰动下跨模态语义对齐脆弱的问题,提出闭环双向提示方法,通过动态反馈循环恢复跨模态一致性,并引入语义锚点约束循环更新,实现实例自适应保护,在11个数据集上达到最先进的鲁棒性和泛化性能。

Comments 24 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25541 2026-05-26 cs.CG cs.AI cs.HC cs.LG 57%

TopoAlign: Topology-Aware Visual Representation Alignment

TopoAlign:拓扑感知的视觉表示对齐

Xinyuan Yan, Rita Sevastjanova, Mennatallah El-Assady, Bei Wang

机构 * University of Utah(犹他大学) ETH Zürich(苏黎世联邦理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 提出TopoAlign框架,利用拓扑数据分析中的mapper图,通过联合力导向优化、自动结构匹配区域检测和基序查询,从拓扑角度比较不同模型或层的表示结构对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏