Less is More: Generating Grounded Navigation Instructions from Landmarks
专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL
Comments CVPR 2022 Camera-ready
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL
Comments CVPR 2022 Camera-ready
专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL
Journal ref EMNLP 2021
专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI
专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL
Comments Accepted at ACM Multimedia 2021. Code available at https://github.com/MILVLG/rosita
专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL
Comments Paper has been published in the AAAI2021 conference
专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL
Comments ECCV 2020, Code and pre-trained models are released: https://github.com/microsoft/Oscar
专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL
Comments ICPR 2020
LaTtE-Flow: 基于层间时间步专家流的Transformer
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; University of Maryland(马里兰大学) ; Nvidia(英伟达) ; Salesforce AI Research(Salesforce AI研究) ; Intuit AI Research(Intuit AI研究)
专题命中 图文多模态 :multimodal(abstract,comments);multimodal foundation model(abstract);分类 cs.CV
AI总结 提出LaTtE-Flow,一种基于预训练视觉语言模型的高效统一架构,通过层间时间步专家流和条件残差注意力机制,实现图像理解与生成,生成速度提升约6倍。
Comments Unified multimodal model, Flow-matching
当否定是在视觉-语言模型中一个几何问题
机构 * ETRO Department, Vrije Universiteit Brussel(布鲁塞尔自由大学ETRO系) ; imec ; Independent Researcher(独立研究员)
专题命中 图文多模态 :multimodal(abstract,comments);image-text(abstract);分类 cs.CV
AI总结 本文探讨了视觉-语言模型中否定理解的几何问题,提出基于多模态大语言模型的评估框架,并通过表示工程操控CLIP模型实现否定意识。
Comments Accepted to CVPR (Multimodal Algorithmic Reasoning Workshop) 2026
专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV;MLLM(comments)
Comments Accepted by ICCV'23 (Oral); Add evaluation on MLLM
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI;multimodal(comments)
Comments This paper is greatly modified and updated to be re-submitted to another conference. The new paper is under the name "Multimodal Adaptive Distillation for Leveraging Unimodal Encoders for Vision-Language Tasks", https://doi.org/10.48550/arXiv.2204.10496
默认以视觉为中心:解决面向盲人和低视力(BLV)用户的实时视觉语言模型(VLM)辅助中的隐含视觉假设问题
专题命中 图文多模态 :multimodal(title)
AI总结 针对现有面向BLV用户的实时VLM辅助工具存在的以视觉为中心的默认偏见问题,提出VIA-Agent模型,其在保持与Doubao相当成功率的同时,缩短了任务时间并减少了对话轮次,提升了用户信任度。
Comments Accepted to UIST 2026
打破幻觉:当积极与消极在多模态解码中相遇
机构 * School of Astronautics, Beihang University(北京航空航天大学航天学院) ; Longcat Interaction Team, Meituan(美团Longcat交互团队) ; Tianmushan Laboratory, Beihang University(北京航空航天大学天门山实验室)
专题命中 图文多模态 :multimodal(title)
AI总结 本文提出PND框架,通过在解码过程中引入正负对比路径,增强视觉真实性,无需重新训练即可在POPE、MME和CHAIR数据集上取得最佳性能。
Comments Accepted by CVPR 2026 (Conference on Computer Vision and Pattern Recognition). 11 pages, 5 figures. Code available at: https://github.com/JiangYubo4399/PND
生成向量搜索以提升多模态视觉-语言任务中的病理基础模型
专题命中 图文多模态 :multimodal(title)
AI总结 STHLM通过生成向量搜索方法提升多模态视觉-语言任务中病理基础模型的检索性能,实现10-30%的性能提升和10倍的维度压缩
Comments 13 pages main (54 total), 2 main figures (9 total)
专题命中 图文多模态 :multi-modal(title)
专题命中 图文多模态 :multimodal(title)
专题命中 图文多模态 :multimodal(title)
专题命中 图文多模态 :multimodal(title)
专题命中 图文多模态 :multimodal(title)
Comments to appear at CHI 2024
专题命中 图文多模态 :multimodal(title)
Comments 33 pages
专题命中 图文多模态 :image-text(title)
Comments ICPR2020
基于主导性的测试时适应以应对视觉-语言模型在模态特定偏移下的表现
机构 * Guangdong University of Technology(广东工业大学)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
AI总结 本文研究了视觉-语言模型在部署时视觉与文本分支不对称偏移的问题,提出MG-MTTA方法通过控制模态可靠性而非仅预测熵来提升性能。
Comments Accepted by ACM MM 2026
SLAP:基于部分最优传输的鱼类重识别选择性局部视觉-语言对齐
机构 * University of Verona(维罗纳大学) ; Institute of Marine Research(海洋研究所) ; University of Agder(阿格德大学)
专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV
AI总结 针对鱼类重识别中全局对齐易引入噪声的问题,提出基于POT的选择性局部视觉-语言对齐框架,在多数据集上实现优于CLIP类方法的性能,泛化性良好。
Comments This is the author version prior to incorporating the camera-ready comments. The final version will be included in the Proceedings of the European Conference on Computer Vision (ECCV) 2026
同一硬币的两面:面向视觉-语言模型跨任务攻击的协同演化搜索
机构 * College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) ; Center for Frontier AI Research, Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局前沿人工智能研究中心) ; School of Electrical Engineering and Automation, Fuzhou University(福州大学电气工程与自动化学院) ; ByteDance(字节跳动)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
AI总结 针对视觉-语言模型易受对抗扰动的问题,提出协同演化跨模态攻击框架,联合优化文本与视觉空间,在多任务上展现出强攻击性能与跨任务可迁移性。
Comments 15 pages, 7 figures, and 8 tables; includes supplementary material
用于3D理解的参数高效CLIP适配:通过统一分词实现
机构 * Fondazione Bruno Kessler(布鲁诺·科塞拉基金会) ; University of Trento(特伦托大学) ; University of Pisa(比萨大学) ; Beijing Forestry University(北京林业大学) ; Shanghai Jiao Tong University(上海交通大学) ; Hong Kong University of Science and Technology (GZ)(香港科技大学) ; Shandong University(山东大学) ; University of California, Merced(加州大学默塞德分校)
专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV
AI总结 本文提出参数高效框架UTok3D,通过学习尺度归一化3D分词器,实现冻结CLIP视觉主干在无标注情况下对不同尺度点云的复用,完成3D分割任务。
Comments 14 pages, tokenizer
自监督点云编码器在高效3D大语言模型中的效能研究
机构 * Beijing University of Posts and Telecommunications(北京邮电大学)
专题命中 图文多模态 :multi-modal(abstract);image-text(abstract);分类 cs.CV
AI总结 该研究探究低成本自监督点云编码器能否替代昂贵多模态编码器用于3D-LLM,经实验发现其与架构存在强交互,为高性价比3D-LLM设计提供指南。
Comments 14 pages, 3 figures. This work has been previously released as a preprint on ChinaXiv (No. ChinaXiv:202607.00167, DOI: 10.12074/202607.00167)
Precise Shield:通过神经层面指导解释和对齐VLLM安全
机构 * Nanjing University of Science and Technology(南京理工大学) ; National University of Singapore(新加坡国立大学) ; Beihang University(北京航空航天大学) ; Nanjing Forestry University(南京林业大学)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
AI总结 本文提出Precise Shield框架,通过对比有害与良性输入的激活模式识别安全神经元,并通过梯度掩码限制参数更新,提升VLLM安全性能并保持多语言多模态泛化能力。
通过目标感知数据对齐实现细粒度食品图像理解
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
AI总结 研究针对细粒度食品图像理解,提出以数据为中心的多模态对齐方法,先选视觉相关训练子集,再细化字幕,训练互补检索专家并融合决策,提升了检索性能,完整方法检索分数超纯VLM检索两倍且更高效。
当汇点帮助或伤害:面向大视觉-语言模型的注意力汇点统一框架
机构 * KAIST(韩国科学技术院) ; Chung-Ang University(中央大学) ; Technical University of Munich(慕尼黑工业大学)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
AI总结 本文研究了大视觉-语言模型中注意力汇点的影响,提出统一框架分析其作为冗余 artifacts 或全局先验的作用,并通过 Layer-wise Sink Gating 模块平衡全局推理与局部证据。
Comments Acknowledgments updated
VISTA-Bench: 视觉-语言模型是否真的能像纯文本一样理解可视化文本?
机构 * Dalian University of Technology(大连理工大学) ; Nanyang Technological University(南洋理工大学)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
AI总结 VISTA-Bench通过对比纯文本和可视化文本问题,揭示了视觉-语言模型在处理可视化文本时的模态差距,发现模型在语义相同的情况下表现显著下降。
Comments 32 pages, 16 figures