arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-06-26 至 2026-06-26 共收录 91 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 12 篇

2606.26196 2026-06-26 cs.CL cs.AI cs.CV cs.LG cs.MM 新提交 89%

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

从结构到协同:多模态大语言模型中视觉-语言感知范式演进综述

Haoxiang Sun, Tao Wang, Li Yuan, Jian Zhao, Jiancheng Lv

机构 * School of Computer Science, Sichuan University(四川大学计算机学院) School of Electronic and Computer Engineering, Peking University Shenzhen Graduate School(北京大学深圳研究生院电子与计算机工程学院) Institute of Artificial Intelligence (TeleAI), China Telecom and Northwestern Polytechnical University(中国电信与西北工业大学人工智能研究院(TeleAI))

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文系统综述多模态大语言模型中统一视觉-语言感知的范式演进,提出五阶段分类法,梳理各阶段代表性方法,并指出开放挑战与未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27373 2026-06-26 cs.CV 新提交 79%

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

在自进化大型多模态模型中更加关注视觉标记

Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar, Rao Muhammad Anwer, Hisham Cholakkal, Salman Khan, Fahad Khan

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Aalto University(阿尔托大学) Australian National University(澳大利亚国立大学) Linköping University(林雪平大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 提出VISE框架,通过几何不变性和语义不变性奖励直接正则化视觉条件策略,解决自进化LMMs中视觉欠条件问题,在无监督设置下提升视觉语言理解。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26922 2026-06-26 cs.RO cs.AI 新提交 79%

Risk-Aware Selective Multimodal Driver Monitoring with Driver-State World Modeling

风险感知的选择性多模态驾驶员监控与驾驶员状态世界建模

Daosheng Qiu, Haozhuang Chi, Hao Su, Shu Long, Xinyue Miao, Yongle Dong, Wei Zhang

机构 * Hubei University(湖北大学) Nanyang Technological University(南洋理工大学) Osaka University(大阪大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 提出成本感知的选择性推理框架,结合轻量级RGB-生理学生网络和门控机制,在低延迟下减少不安全决策,实现可部署的多模态驾驶员监控。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26794 2026-06-26 cs.CV cs.AI 新提交 73%

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP

ReasonCLIP-58M: CLIP的视觉基础常识推理监督

Sicheng Zhang, Muzammal Naseer, Binzhu Xie, Naufal Suryanto, Shi Qiu, Jamal Bentahar, Naveed Akhtar, Mubarak Shah

机构 * Khalifa University(卡利法大学) University of Western Australia(西澳大学) The Chinese University of Hong Kong(香港中文大学) University of Melbourne(墨尔本大学) University of Central Florida(佛罗里达中央大学)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 提出ReasonCLIP-58M框架,通过两阶段策略将大规模推理监督集成到CLIP模型中,构建ReasonLite-42M和ReasonPro-16M数据集及RCLIP-Bench基准,提升视觉基础常识推理和零样本检索性能。

Comments Accepted to ECCV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27023 2026-06-26 cs.LG cs.CL cs.CV 新提交 62%

Just how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQA

你到底有多确定?改进医学视觉问答中的口头不确定性校准

Eren Senoglu, Federico Toschi, Nicolo Brunello, Andrea Sassella, Mark James Carman

机构 * Politecnico di Milano(米兰理工大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 针对多模态大语言模型在医学VQA中过度自信的问题,提出基于复合损失函数的微调框架,通过校准项、锚定正则化、对比对齐和KL稳定项,将校准误差降低60%以上,判别力提升26%以上。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26891 2026-06-26 cs.CV cs.AI 新提交 62%

Bridging Vision and Language Concepts through Optimal Transport Semantic Flow

通过最优传输语义流桥接视觉与语言概念

Chenyang Zhang, Anqi Dong, Guangming Zhu, Nuoye Xiong, Siyuan Wang, Lin Mei, Liang Zhang

机构 * School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院) KTH Royal Institute of Technology(瑞典皇家理工学院)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 提出最优传输流概念瓶颈模型(OTF-CBM),通过逆最优传输学习数据驱动的语义代价,并利用非平衡最优传输流匹配建模视觉块与文本概念间的语义转换,实现可解释的几何关系捕获,提升分类精度与概念忠实度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16159 2026-06-26 cs.CV cs.AI 版本更新 62%

Through the Looking Glass: A Dual Perspective on Weakly-Supervised Few-Shot Segmentation

透过镜中奇境:弱监督小样本分割的双重视角

Jiaqi Ma, Guo-Sen Xie, Fang Zhao, Zechao Li

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出同源异构网络,通过异构视觉聚合和转移模块增强双重视角互补性,结合异构CLIP文本信息,在弱监督小样本分割任务中以极少参数超越现有方法。

Comments Accepted by IEEE Transactions on Image Processing (TIP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08420 2026-06-26 cs.CV 新提交 57%

CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs

CheXanatomy: 面向胸部X光片的解剖感知视觉-语言建模

Sergios Gatidis, Curtis Langlotz, Christian Bluethgen

机构 * Stanford Center for Artificial Intelligence in Medicine and Imaging, Stanford University(斯坦福大学医学与影像人工智能中心) Department of Radiology, Stanford University(斯坦福大学放射学系)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

AI总结 提出CheXanatomy框架,通过自回归令牌空间监督将解剖知识融入预训练视觉-语言模型,实现解剖分割,在合成和真实X光片上性能媲美U-Net,并提升域迁移鲁棒性和样本效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06467 2026-06-26 cs.CV 版本更新 57%

GreenRFM: Learning a resource-efficient radiology vision-language foundation model via supervision-centric pre-training

GreenRFM:通过以监督为中心的预训练学习资源高效的放射学视觉-语言基础模型

Yingtai Li, Shuai Ming, Qiuli Wang, Mingyue Zhao, Hongchun Zhang, Yuhe Tian, Haoran Lai, Rongsheng Wang, Rui Zhou, Rundong Wang, Yujia Li, Zhiyang He, Xiaodong Tao, Wei Chen, Wei Wei, Shaohua Kevin Zhou

机构 * School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China (USTC)(生物医学工程学院,生命科学与医学系,中国科学技术大学) Center for Medical Imaging, Robotics, Analytic Computing & Learning (MIRACLE)(医学影像、机器人、分析计算与学习中心) Suzhou Institute for Advanced Research, USTC(苏州先进研究院,中国科学技术大学) Department of Radiology, The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine, USTC(放射科,中国科学技术大学第一附属医院,生命科学与医学系,中国科学技术大学) T Magnetic Resonance Translational Medicine Research Center, Department of Radiology, The First Affiliated Hospital (Southwest Hospital) of Army Medical University(7T磁共振转化医学研究中心,放射科,中国医学大学第一附属医院(西南医院))

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 提出GreenRFM框架,通过将噪声报告转化为结构化诊断信号,以监督为中心预训练,在有限资源下实现高效3D放射学视觉-语言表示,零样本CT-RATE AUC达84.8。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04238 2026-06-26 cs.CV 版本更新 57%

6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models

6根手指,1个肾脏:自然对抗性医学图像揭示视觉语言模型的关键弱点

Leon Mayer, Piotr Kalinowski, Caroline Ebersbach, Marcel Knopp, Tim Rädsch, Evangelia Christodoulou, Annika Reinke, Fiona R. Kolbinger, Lena Maier-Hein

机构 * German Cancer Research Center (DKFZ) Heidelberg, Division of Intelligent Medical Systems(德国癌症研究中心(DKFZ)海德堡,智能医学系统部门) Medical Faculty, Heidelberg University(海德堡大学医学院) Faculty of Mathematics and Computer Science, Heidelberg University(海德堡大学数学与计算机科学学院) HIDSS4Health - Helmholtz Information and Data Science School for Health, Karlsruhe/Heidelberg(HIDSS4Health - 哈勃-马克斯信息与数据科学健康学院,卡尔斯鲁厄/海德堡) Helmholtz Imaging, German Cancer Research Center (DKFZ)(哈勃-马克斯成像,德国癌症研究中心(DKFZ)) Engineering Faculty, Heidelberg University(海德堡大学工程学院) School of Computation, Information and Technology, TUM(技术大学(TUM)计算、信息与技术学院) Weldon School of Biomedical Engineering, Purdue University(普渡大学韦尔登生物医学工程学院) Department of Visceral, Thoracic and Vascular Surgery, University Hospital and Faculty of Medicine Carl Gustav Carus, TUD Dresden University of Technology(visceral、胸腔和血管外科部门,技术大学(TUD)德累斯顿大学医院和医学院) National Center for Tumor Diseases (NCT), NCT Heidelberg, a partnership between DKFZ and University Hospital Heidelberg(肿瘤疾病国家中心(NCT),海德堡NCT,DKFZ与海德堡大学医院之间的合作) Heidelberg University Hospital, Surgical Clinic, Surgical AI Research Group(海德堡大学医院,外科诊所,外科人工智能研究组) Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, UAE(Mohamed Bin Zayed人工智能大学(MBZUAI),阿布扎赫,阿拉伯联合酋长国)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 提出AdversarialAnatomyBench基准,测试25个视觉语言模型在罕见解剖变异上的表现,发现准确率从71%降至28%,且模型缩放和干预无法解决。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27146 2026-06-26 cs.RO 新提交 50%

PhysReflect-VLA: Physical Feasibility and Self-Reflective Regulation for Reliable Vision-Language-Action Policies

PhysReflect-VLA:面向可靠视觉-语言-动作策略的物理可行性与自反思调节

Jiayu Yang, Tao Yang, Weijun Li, Xiang Chang, Fei Chao, Changjing Shang, Qiang Shen

机构 * Xiamen University(厦门大学) Aberystwyth University(阿伯里斯特威斯大学)

专题命中 图文多模态 :multimodal(abstract)

AI总结 提出PhysReflect-VLA框架,通过物理可行性评估和结构化自反思增强VLA策略,在闭环控制中实现稳定多阶段操作,平均成功率提升5.4%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27144 2026-06-26 cs.RO 新提交 50%

PAMAE: Phase-Aware-MoE Action Experts Towards Reliable Flow-Matching Vision-Language-Action Policies

PAMAE: 面向可靠流匹配视觉-语言-动作策略的相位感知MoE动作专家

Jiayu Yang, Tao Yang, Xiang Chang, Fei Chao, Changjing Shang, Qiang Shen

机构 * Department of Artificial Intelligence, School of Informatics, Xiamen University(厦门大学信息学院人工智能系) Department of Computer Science, Aberystwyth University(阿伯里斯特威斯大学计算机科学系)

专题命中 图文多模态 :multimodal(abstract)

AI总结 针对多阶段机器人操作中VLA模型动作生成不可靠的问题,提出即插即用的相位感知混合专家动作模块PAMAE,通过稀疏专家混合和相位感知路由提升阶段一致性,在模拟任务中成功率提升9.2%。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 6 篇

2606.26107 2026-06-26 cs.CL cs.AI 新提交 81%

Low Resource Multimodal Translation of Nepali Spoken Words into Emotion-Conditioned Sign Language Avatars

低资源场景下尼泊尔口语词汇到情感条件手语虚拟人物的多模态翻译

Jatin Bhusal, Salma Tamang

机构 * Center for Human Mobility and Communications, Prateek Innovations(普拉蒂克创新公司人类移动与通信中心) Sunway International Business School, Birmingham City University(双威国际商学院,伯明翰城市大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 提出NEST-V1框架,通过共享声学编码器实现语音识别与情感分类,生成情感条件尼泊尔手语虚拟人物,在低资源场景下验证了技术可行性。

Comments 15 pages, 5 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25041 2026-06-26 cs.CV cs.AI cs.GR cs.SD 新提交 79%

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

Wan-Streamer v0.1: 端到端实时交互基础模型

Lianghua Huang, Zhi-Fan Wu, Wei Wang, Yupeng Shi, Mengyang Feng, Junjie He, Chen-Wei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chen Liang, Cheng Yu, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Yuzheng Wang, Zoubin Bi

机构 * Alibaba Group(阿里巴巴集团)

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.AI

AI总结 提出Wan-Streamer,一种原生流式、端到端的交互基础模型,通过单一Transformer联合建模语言、音频和视频,实现低延迟全双工音视频交互,无需外部模块。

Comments Website: https://wan-streamer.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03762 2026-06-26 eess.AS cs.LG 版本更新 70%

Conditional Flow Matching for Visually-Guided Acoustic Highlighting

基于条件流匹配的视觉引导声学增强

Hugo Malard, Gael Le Lan, Daniel Wong, David Lou Alon, Yi-Chiao Wu, Sanjeel Parekh

机构 * LTCI, Télécom Paris, Institut Polytechnique de Paris(LTCI,巴黎电信学院,巴黎理工学院) Meta

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 eess.AS

AI总结 提出条件流匹配生成框架解决视觉引导音频重混中的多对多映射问题,通过滚动损失和跨模态融合模块实现优于判别方法的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26793 2026-06-26 cs.CR cs.AI cs.LG 新提交 57%

MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG

MIRROR: 基于新颖性约束的记忆引导MCTS红队测试用于智能体RAG

Inderjeet Singh, Andrés Murillo, Motoyoshi Sekiya, Yuki Unno, Junichi Suga

机构 * Fujitsu Research of Europe, United Kingdom(欧洲富士通研究机构,英国) Fujitsu Limited, Japan(日本富士通有限公司)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

AI总结 提出MIRROR框架,通过记忆引导蒙特卡洛树搜索和显式新颖性约束,在多模态智能体RAG系统的四个攻击面上实现高攻击成功率,并降低跨表面方差。

Comments 6 pages, 2 figures. Accepted at the 2026 International Joint Conference on Neural Networks (IJCNN 2026), IEEE WCCI 2026; presented as an oral talk. Code and ART-SafeBench benchmark: https://github.com/FujitsuResearch/mirror

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08660 2026-06-26 cs.CL 版本更新 57%

A Systematic Survey of Semantic Role Labeling in the Era of Pretrained Language Models

预训练语言模型时代语义角色标注的系统综述

Huiyao Chen, Meishan Zhang, Jing Li, Lilja Øvrelid, Jan Hajič, Hao Fei, Min Zhang

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Shenzhen Loop Area Institute (SLAI)(深圳南山区研究院) University of Oslo(奥斯陆大学) Charles University(查尔斯大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

AI总结 提出四维分类法系统综述SRL研究,分析句法特征的作用条件,首次系统处理大语言模型时代的SRL,并扩展至多模态设置。

Comments 54 pages, 9 figures, 9 tables. Accepted at Artificial Intelligence Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06080 2026-06-26 cs.CV cs.CY cs.HC 版本更新 57%

AIDEN: Design and Pilot Study of an AI Assistant for the Visually Impaired

AIDEN:面向视障人士的AI助手设计与初步研究

Luis Marquez-Carpintero, Francisco Gomez-Donoso, Zuria Bauer, Bessie Dominguez-Dager, Alvaro Belmonte-Baeza, Mónica Pina-Navarro, Francisco Morillas-Espejo, Felix Escalona, Miguel Cazorla

机构 * Institute for Computer Research, University of Alicante(计算机研究所,阿利坎特大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

AI总结 提出AIDEN系统,结合YOLO实时目标检测、LLaVA场景描述与OCR,以及基于盖革计数器隐喻的连续触觉引导,避免听觉过载并保护隐私,实验表明用户满意度高。

Journal ref IEEE Access 14 (2026) 80406-80420

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 10 篇

2604.19193 2026-06-26 cs.CV 79%

How Far Are Video Models from True Multimodal Reasoning?

视频模型距离真正的多模态推理还有多远?

Xiaotian Zhang, Jianhui Wei, Yuan Wang, Jie Tan, Yichen Li, Yan Zhang, Ziyi Chen, Daoan Zhang, Dezhi YU, Wei Xu, Songtao Jiang, Zuozhu Liu

机构 * Zhejiang University(浙江大学) ByteDance(字节跳动)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出CLVG-Bench评估框架,通过上下文学习测试视频模型的零样本推理能力,揭示当前SOTA模型在逻辑推理和交互生成任务中表现不足,指出多模态推理和物理基础是关键瓶颈。

Journal ref ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26626 2026-06-26 cs.HC 新提交 78%

Reviving Reflection-in-Action: Instilling Designerly Thinking in AI-Supported Ideation through Multimodal Prompting

复兴反思行动:通过多模态提示在AI支持的构思中注入设计思维

Samangi Wadinambiarachchi, Jenny Waycott, Greg Wadley

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 针对当前AI创造力支持工具主要依赖文本提示的局限,提出SketchifAI原型,通过多模态输入(文本、草图、草图加标签)增强设计学生的表达意图、创造力支持和发散思维,发现草图模态提升流畅性但用户偏好文本,探讨如何通过草图反思保留设计技能。

Comments ACM Creativity and Cognition 2026 conference (C&C'26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27187 2026-06-26 cs.CV cs.CL 新提交 76%

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

HarmVideoBench:大型多模态模型中有害视频理解的基准测试

Jiajun Wu, Haoyu Kang, Yining Sun, Jiacheng Hou, Heng Zhang, Danyang Zhang, Zhenjun Zhao, Haochi Zhang, Leixin Sun, Eric Hanchen Jiang, Yushan Li, Ruiyu Li, Mengkai Huang, Yan Gao, Xu Zhang, Guancheng Wan

机构 * Central South University(中南大学) Tsinghua University(清华大学) South China Normal University(华南师范大学) ByteDance Inc(字节跳动公司) University of Zaragoza(阿拉维达大学) CosmosMind Wuhan University(武汉大学) University of California, Los Angeles(加州大学洛杉矶分校) Southeast University(东南大学) Tencent(腾讯公司) Nankai University(南开大学) Supermicro Computer Inc(Supermicro计算机公司) Huazhong University of Science and Technology(华中科技大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.CL

AI总结 提出HarmVideoBench,一个包含1379个视频和4137道选择题的多层次诊断基准,从三个维度评估模型对有害视频的深层理解,并引入BCR方法将宏平均准确率从61.7%提升至84.4%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01958 2026-06-26 cs.CV 版本更新 70%

MAVFusion: Efficient Infrared and Visible Video Fusion via Motion-Aware Sparse Interaction

MAVFusion: 基于运动感知稀疏交互的高效红外与可见光视频融合

Xilai Li, Weijun Jiang, Xiaosong Li, Yang Liu, Hongbin Wang, Tao Ye, Huafeng Li, Haishu Tan

机构 * Foshan University(佛山大学) Kunming University of Science and Technology(昆明理工大学) China University of Mining and Technology(中国矿业大学)

专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出MAVFusion端到端视频融合框架,通过运动感知稀疏交互机制,利用光流识别动态区域并分配跨模态注意力,对静态区域使用轻量弱交互模块,在保持时间一致性和细节的同时加速推理,达到14.16 FPS。

Comments Accepted at ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15347 2026-06-26 eess.IV cs.MM 版本更新 70%

Symmetric Entropy-Constrained Video Coding for Machines

面向机器的对称熵约束视频编码

Yuxiao Sun, Meiqin Liu, Chao Yao, Qi Tang, Jian Jin, Weisi Lin, Frederic Dufaux, Yao Zhao

专题命中 视频多模态 :MLLM(abstract,abstract_cn);分类 cs.MM

AI总结 提出SEC-VCM框架,通过双向熵约束机制对称对齐视频编解码器与视觉骨干网络,保留语义并丢弃无关信息,结合语义-像素双路径融合提升机器视觉任务性能,在多项任务上显著优于H.266/VVC。

Comments Accepted by IEEE Transactions on Image Processing. This is the author's accepted manuscript (AAM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26455 2026-06-26 cs.CV cs.AI cs.LG 新提交 62%

Active Adversarial Perturbation-driven Associative Memory Retrieval for RGB-Event Visual Object Tracking

主动对抗扰动驱动的关联记忆检索用于RGB-事件视觉目标跟踪

Xiao Wang, Xufeng Lou, Zikang Yan, Lan Chen, Sibao Chen, Yaowei Wang, Yonghong Tian, Jin Tang

机构 * School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院) School of Electronic and Information Engineering, Anhui University(安徽大学电子信息工程学院) Peng Cheng Laboratory(鹏城实验室) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) School of Computer Science, Peking University(北京大学计算机学院) School of Electronic and Computer Engineering, Shenzhen Graduate School, Peking University(北京大学深圳研究生院电子与计算机工程学院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 提出APRTrack框架,通过分层对抗扰动模拟模态退化与局部缺失,并设计足迹引导的通道校准Hopfield检索实现鲁棒的历史信息补偿,在多个数据集上验证了有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03371 2026-06-26 cs.CL 版本更新 57%

See, Infer, Intervene: Proactive World Modeling for Goal-Oriented Social Intelligence

观察、推断、干预:面向目标导向社交智能的主动世界建模

Honghui Zhang, Chenmeinian Guo, Yichen Yu, Guanyu Liu, Yujia Zhang, Yongming Qin, Chongguo Song, Mengyue Yang, Lei Yu, Tianyu Shi

机构 * Mita Technology(Mita技术公司) University of Bristol(布里斯托大学) University of Toronto(多伦多大学) McGill University(麦吉尔大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

AI总结 提出 See-Infer-Intervene (SII) 框架和主动意图世界模型 (PIWM),通过观察顾客行为、推断潜在意图并选择干预动作,实现零售场景中的主动辅助,在 GuidanceSalesBench 基准上达到 0.641 macro F1。

Comments 16 pages, 3 figures, 9 tables. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17683 2026-06-26 cs.LG cs.CV stat.ML 版本更新 57%

Probabilistic NDVI Forecasting from Sparse Satellite Time Series and Weather Covariates

基于稀疏卫星时间序列和天气协变量的概率NDVI预测

Irene Iele, Giulia Romoli, Daniele Molino, Elena Mulero Ayllón, Filippo Ruffini, Paolo Soda, Matteo Tortora

机构 * Department of Naval, Electrical, Electronics and Telecommunications Engineering, University of Genoa, Italy(海军、电子、电子与电信工程系,热那亚大学,意大利)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出一种概率预测框架,利用稀疏卫星数据和天气协变量进行田间NDVI预测,通过融合历史和未来数据实现多步分位数预测,实验表明优于传统和深度学习方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12062 2026-06-26 cs.CV 版本更新 57%

Learning Language-Driven Sequence-Level Modal-Invariant Representations for Video-Based Visible-Infrared Person Re-Identification

基于语言驱动的序列级模态不变表示学习用于视频可见光-红外行人重识别

Xiaomei Yang, Antai Liu, Xizhan Gao, Fa Zhu, Sijie Niu, Giancarlo Fortino

机构 * Shandong Key Laboratory of Ubiquitous Intelligent Computing, School of Information Science and Engineering, University of Jinan(山东省 Ubiquitous Intelligent Computing 重点实验室,济南大学信息科学与工程学院) College of Information Science and Technology & Artificial Intelligence, Nanjing Forestry University(信息科学与技术及人工智能学院,南京林业大学) Department of Informatics, Modeling, Electronics, and Systems, University of Calabria(信息学、建模、电子与系统系,卡利博大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

AI总结 提出LSMRL方法,通过CLIP的时空特征学习、语义扩散和跨模态交互模块,结合模态级损失,学习序列级模态不变表示,在VVI-ReID任务上超越现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06480 2026-06-26 cs.RO 版本更新 50%

History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation

基于历史条件的时空视觉令牌剪枝用于高效视觉-语言导航

Qitong Wang, Yijun Liang, Ming Li, Tianyi Zhou, Christopher Rasmussen

机构 * Department of Computer and Information Sciences at the University of Delaware(德克萨斯大学达勒姆分校计算机与信息科学系) University of Maryland’s Department of Computer Science(马里兰大学计算机科学系) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 视频多模态 :multimodal(abstract)

AI总结 提出一种无需训练的时空视觉令牌剪枝框架,通过空间令牌选择和时空压缩减少冗余计算,在保持导航精度的同时显著提升推理效率,并在真实机器人上验证了低延迟指令跟随导航。

Comments International Conference on Intelligent Robots and Systems (IROS) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 6 篇

2606.27010 2026-06-26 cs.IR cs.MM 新提交 83%

TriPAH: Imbalance-Aware Tri-Prompt Affinity Hashing for Cross-Modal Medical Retrieval

TriPAH:面向跨模态医学检索的不平衡感知三提示亲和哈希

Jiaming Bian, Songming Li, Yurui Song, Yunfei Chen, Yichao Cao, Jun Long

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.MM

AI总结 提出TriPAH框架,通过本体引导的患者级提示、轻量级提示-令牌混合器及不平衡感知多任务目标,解决跨模态医学哈希检索中的语义碎片、长尾标签和量化脆弱性问题,在三个数据集上超越现有方法。

Comments 10 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26458 2026-06-26 cs.AI 新提交 79%

MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation

MKG-RAG-Bench:多模态知识图谱增强生成中的检索基准

Xiaochen Wang, Bao Hoang, Han Liu, Ting Wang, Fenglong Ma

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Michigan State University(密歇根州立大学) Dalian University of Technology(大连理工大学) Stony Brook University(石溪大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 提出MKG-RAG-Bench基准,通过构建跨领域多模态知识图谱和问答数据集,系统评估多模态知识图谱增强生成中的检索性能,揭示检索质量对生成结果的决定性作用。

Comments Accepted by KDD'26

详情

展开后加载摘要…

URL PDF HTML 收藏