arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-05-11 至 2026-05-11 共收录 97 信号源:cs.CV, cs.AI, cs.LG

1. VLM训练与架构 20 篇

2605.07110 2026-05-11 cs.CL cs.SE 50%

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability

保障计算机使用代理:一种面向部署的架构-生命周期框架用于可靠性

Zejian Chen, Zhanyuan Liu, Chaozhuo Li, Mengxiang Han, Songyang Liu, Litian Zhang, Feng Gao, Yiming Hei, Xi Zhang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) China Academy of Information and Communications Technology, Artificial Intelligence Institute(信息与通信技术研究院,人工智能研究所)

专题命中 VLM训练与架构 :grounding(abstract)

AI总结 本文提出一种面向部署的架构-生命周期框架,用于提升计算机使用代理的可靠性,通过分析感知、决策和执行层,以及创建、部署、操作和维护阶段,解决能力形成、授权暴露、故障表现和控制放置之间的联系。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13372 2026-05-11 cs.CY 50%

Semantic Alignment Between Normative Theories of Ethics and the European Union Artificial Intelligence Act: A Transformer-Based Semantic Textual Similarity Analysis

伦理规范理论与欧盟人工智能法案之间的语义对齐:基于Transformer的语义文本相似性分析

Mehmet Murat Albayrakoglu, Mehmet Nafiz Aydin

专题命中 VLM训练与架构 :grounding(abstract)

AI总结 本文通过Transformer模型分析伦理理论与欧盟AI法案的语义对齐,发现义务伦理学在法案两部分中具有最高相似性。

Comments 18 pages, 5 tables, 3 figures; the concept of alignment introduced as an indication of influence

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他VLM 5 篇

2605.07512 2026-05-11 cs.CV 79%

Hierarchical Dual-Subspace Decoupling for Continual Learning in Vision-Language Models

层次双子空间解耦用于视觉-语言模型中的持续学习

Mengxin Qin, Xiang Zhang, Kun Wei, Xu Yang, Cheng Deng

机构 * School of Electronic Engineering, Xidian University(西电电子工程学院)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

AI总结 本文提出HDSD框架,通过分解参数空间和结构化参数分解,减少子空间干扰和参数漂移,提升视觉-语言模型持续学习性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07494 2026-05-11 cs.CV 79%

DIMoE-Adapters: Dynamic Expert Evolution for Continual Learning in Vision-Language Models

DIMoE-Adapters:动态专家进化用于视觉语言模型的持续学习

Mengxin Qin, Xiang Zhang, Xi Wang, Kun Wei, Xu Yang, Cheng Deng

机构 * School of Electronic Engineering, Xidian University(西安电子科技大学电子工程学院)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

AI总结 本文提出DIMoE-Adapters框架,通过动态专家进化方法平衡持续学习中的稳定性与可塑性,解决多领域任务增量学习中的领域迁移问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06115 2026-05-11 cs.AI 77%

CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs

CrossCult-KIBench:一种用于多模态大语言模型跨文化知识插入的基准测试

Zhen Zeng, Leijiang Gu, Feng Li, Jing Yu, Zenglin Shi

机构 * Hefei University of Technology(合肥工业大学) Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE, Minzu University of China(民族语言智能分析与安全治理国家重点实验室,中央民族大学) School of Information Engineering, Minzu University of China(中央民族大学信息工程学院)

专题命中 其他VLM :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.AI

AI总结 本文提出CrossCult-KIBench基准测试,用于评估跨文化知识插入的有效性及非目标文化的影响,通过9800个图像场景测试不同语言文化的适应性,提出MCKI方法并揭示了当前模型在文化适应与行为保持间的平衡难题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05340 2026-05-11 cs.CR cs.AI 70%

How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study

VLMs在物理世界中离隐私意识还有多远?一项实证研究

Junran Wang, Xinjie Shen, Zehao Jin, Pan Li

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 其他VLM :vision-language model(abstract,abstract_cn);分类 cs.AI

AI总结 本文通过实证研究揭示VLMs在物理环境中隐私意识的不足,提出ImmersedPrivacy框架评估模型在复杂场景中的隐私感知能力,发现现有模型在感知和隐私冲突处理上存在显著缺陷。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07478 2026-05-11 cs.CV 57%

AudioFace: Language-Assisted Speech-Driven Facial Animation with Multimodal Language Models

AudioFace: 基于语言的语音驱动面部动画与多模态语言模型

Kai Zheng, Zejian Kang, Rui Mao, Hongyuan Zou, Yuanchen Fei, Xuanyang Xu, Xiangru Huang

机构 * Westlake University(西湖大学) Zhejiang University(浙江大学) Tiangong University(天工大学) Hunan University(湖南大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出AudioFace框架,通过语言和发音信息指导语音驱动的面部动作生成,提升语音与面部运动的对应精度。

详情

展开后加载摘要…

URL PDF HTML 收藏