arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-05-27 至 2026-05-27 共收录 25 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 25 篇

2605.26294 2026-05-27 cs.CV 89%

CNNs, Transformers, Hybrid, and Vision Language Models for Skin Cancer Detection

用于皮肤癌检测的CNN、Transformer、混合模型和视觉语言模型

Durjoy Dey, Yuhong Yan, Hassan Hajjdiab

机构 * Department of Computer Science and Software Engineering, Concordia University, Montreal, Canada(计算机科学与软件工程系,康科迪亚大学,加拿大蒙特利尔) Ebovir Biotechnologie Inc., Montreal, Canada(Ebovir生物技术公司,加拿大蒙特利尔)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision language model(title,abstract);分类 cs.CV

AI总结 本文在PAD-UFES-20数据集上统一评估了12种深度学习模型(包括CNN、ViT、混合卷积Transformer和视觉语言模型),结果表明混合模型和基于SigLIP的VLM在排名性能和临床相关操作点之间取得了最佳平衡。

Comments 13 pages, 3 figures, accepted at ICPRAI 2026, The Fifth International Conference on Pattern Recognition and Artificial Intelligence. To appear in Lecture Notes in Computer Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26689 2026-05-27 cs.CV cs.CL 86%

PinPoint: Prompting with Informative Interior Points

PinPoint: 通过信息性内部点进行提示

Pouya Sadeghi, Shawn He, Pedro Pablo Guerrero Vela, C. Thomas, Alex Wong, Sirisha Rambhatla

机构 * University of Waterloo(滑铁卢大学) Critical ML Apple(苹果公司)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 针对指代图像分割中VLM与SAM结合时因提示模糊导致的性能差距,提出无需训练的确定性点选择器PinPoint,通过融合视觉线索选择稳定、信息丰富的内部点,在无训练下达到监督和强化学习方法的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14799 2026-05-27 cs.CV cs.CR cs.SI 84%

Can Visual Mamba Improve AI-Generated Image Detection? An In-Depth Investigation

视觉Mamba能否提升AI生成图像检测?一项深入研究

Mamadou Keita, Wassim Hamidouche, Hessen Bougueffa Eutamene, Abdelmalik Taleb-Ahmed, Xianxun Zhu, Abdenour Hadid

机构 * Laboratory of IEMN, CNRS, Centrale Lille, UMR 8520, Univ. Polytechnique Hauts-de-France(伊姆纳实验室,国家科学研究中心,里尔中央理工大学,UMR 8520,法国高等技术大学) Khalifa University(卡利法大学) School of Communication and Information Engineering, Shanghai University(上海大学通信与信息工程学院) Sorbonne Center for Artificial Intelligence, Sorbonne University Abu Dhabi(索邦人工智能中心,索邦大学阿布扎克分校)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本研究系统评估了Vision Mamba模型在AI生成图像检测中的性能,与CNN、ViT和VLM检测器进行对比,分析了准确性、效率和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26661 2026-05-27 cs.CV cs.AI 84%

Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language Models

在预训练视觉语言模型的后验分布外检测中尊重模态差距

Yuanwei Hu, Bo Peng, Yadan Luo, Zhen Fang, Ling Chen, Jie Lu

机构 * The University of Queensland(昆士兰大学) University of Technology Sydney(悉尼科技大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 针对预训练视觉语言模型在后验分布外检测中文本原型与视觉原型存在模态差距的问题,提出在线伪监督框架直接在视觉特征空间学习类原型,实现新最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01253 2026-05-27 cs.CV 81%

ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection

ODOV:开放域开放词汇目标检测基准

Yupeng Zhang, Ruize Han, Fangnan Zhou, Wei Feng, Liang Wan

机构 * College of Intelligence and Computing, Tianjin University(天津大学智能计算学院) Key Research Center for Surface Monitoring and Analysis of Relics, State Administration of Cultural Heritage(文物表面监测与分析国家重点研究中心) Faculty of Computer Science and Artificial Intelligence, Shenzhen University of Advanced Technology(深圳先进技术大学计算机科学与人工智能学院)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);分类 cs.CV

AI总结 针对真实场景中域偏移和类别偏移同时发生的问题,提出开放域开放词汇目标检测任务,构建OD-LVIS基准数据集,并设计基于VLM的基线方法,通过域无关类别提示和域投影嫁接模块提升检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26621 2026-05-27 cs.CV cs.AI 81%

MedVol-R1: Reward-Driven Evidence Grounding for Volumetric Reasoning Segmentation

MedVol-R1:基于奖励驱动的证据基础用于体积推理分割

Zichun Wang, Hairong Shi, Bingzheng Wei, Yan Xu, Zihua Wang

机构 * School of Biological Science and Medical Engineering, Beihang University, Beijing, China(生物科学与医学工程学院,北京航空航天大学) Center for Information and Computer Science, School of Science for Open and Environmental Systems, Graduate School of Science and Technology, Keio University, Kanagawa, Japan(信息与计算机科学中心,开放与环境系统科学学院,科技研究生学校,东京大学,神奈川,日本) Bytedance Inc., China(字节跳动公司,中国) Tsinghua University, Beijing, China(清华大学,北京,中国)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

AI总结 提出MedVol-R1框架,通过强化学习将临床推理解耦为可验证的2D证据锚点,再传播为3D掩膜,实现体积推理分割,在多个基准上达到最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26460 2026-05-27 cs.CV cs.AI 81%

AnchorDiff: Training-Free Concept Grounding for MM-DiTs via Anchor-Based Graph Propagation

AnchorDiff: 基于锚点图传播的无训练概念定位用于多模态扩散Transformer

Jian Zhang, Zhijun Zhang

机构 * School of Automation Science and Engineering(自动化科学与工程学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

AI总结 提出AnchorDiff方法,通过锚点选择和混合图传播解耦语义定位与结构细化,解决多模态扩散Transformer中视觉混淆概念间的概念泄漏问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26441 2026-05-27 cs.CV cs.AI 81%

Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective

从博弈视角重新思考弱监督视频时间定位

Xiang Fang, Zeyu Xiong, Wanlong Fang, Xiaoye Qu, Chen Chen, Jianfeng Dong, Keke Tang, Pan Zhou, Yu Cheng, Daizong Liu

机构 * Hubei Key Laboratory of Distributed System Security(湖北分布式系统安全重点实验室) Hubei Engineering Research Center on Big Data Security(大数据安全工程研究中心) School of Cyber Science and Engineering(网络安全科学与工程学院) Huazhong University of Science and Technology(华中科技大学) University of Central Florida(佛罗里达中央大学) Zhejiang Gongshang University(浙江工商大学) Guangzhou University(广州大学) The Chinese University of Hong Kong(香港中文大学) Peking University(北京大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

AI总结 本文从博弈论视角出发,通过多元合作博弈建模帧与词的不确定对应关系,实现多级跨模态交互,从而在弱监督下提升视频时间定位的准确性。

Comments Published in ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27168 2026-05-27 cs.CL cs.AI cs.CY 79%

Grounding Text Embeddings in Stakeholder Associations

将文本嵌入与利益相关者关联对齐

Jonathan Rystrøm, Sofie Burgos-Thorsen, Zihao Fu, Johan Irving Søltoft, Kenneth C. Enevoldsen, Chris Russell

机构 * University of Oxford(牛津大学) Institute for Wicked Problems(复杂问题研究所) The Chinese University of Hong Kong(香港中文大学) Danish Technical University(丹麦技术大学) Aarhus University(奥胡斯大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 提出利益相关者对齐练习方法,通过评估嵌入模型与人类专家的语义距离一致性,发现神经文本嵌入在丹麦政策案例中可靠性显著低于专家(差距19-26个百分点),且该差距在美国联邦AI用例中复现(16个百分点)。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26421 2026-05-27 cs.CV 79%

HydraPrompt: An Adaptive and Asymmetric Framework of Vision-Language Models for Synthetic Image Detection

HydraPrompt: 面向合成图像检测的视觉语言模型自适应非对称框架

Senyuan Shi, Hao Tan, Zichang Tan, Shuhan Feng, Ajian Liu, Sergio Escalera, Jun Wan

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) School of Advanced Interdisciplinary Sciences (SAIS), University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院) Shenzhen Institute of Advanced Technology (SIAT), Chinese Academy of Sciences(中国科学院深圳先进技术研究所) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) University of Barcelona(巴塞罗那大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

AI总结 提出一种非对称提示框架HydraPrompt,通过动态调整类别中心对齐细粒度图像线索,结合条件监督对比学习,实现合成图像检测的SOTA性能。

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26383 2026-05-27 cs.CV 70%

Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion

基于多阶段SAM3特征融合的零样本物体重识别在自我中心厨房视频中的应用

Dmytro Klepachevskyi, Alexander Wong, Sirisha Rambhatla, Yuhao Chen

机构 * University of Waterloo(滑铁卢大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract_cn);分类 cs.CV

AI总结 针对自我中心厨房视频中物体重识别的挑战,提出一种基于SAM3分割的多阶段零样本方法,通过融合SAM3、DINOv2和CLIP特征并引入掩码形状IoU和k-倒数重排序,将mAP从45.3%提升至52.8%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22274 2026-05-27 cs.CV 70%

CAGE-SGG: Counterfactual Active Graph Evidence for Open-Vocabulary Scene Graph Generation

CAGE-SGG:用于开放词汇场景图生成的反事实主动图证据

Suiyang Guang, Chenyu Liu, Ruohan Zhang, Siyuan Chen

机构 * Institute of Intelligent Vision and Embodied Cognition(智能视觉与具身认知研究院)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 提出基于反事实关系验证的开放词汇场景图生成框架,通过分解谓词为软证据基并使用反事实验证器确保关系有视觉证据支持,从而提升可靠性、可解释性和泛化能力。

Comments This manuscript has been withdrawn by the authors because we found a methodological flaw in the formulation and evaluation of the proposed approach. The issue affects the reliability of the experimental results and the conclusions drawn from them. Therefore, the authors consider the current version unsuitable for citation or further use

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27203 2026-05-27 cs.CV cs.AI 62%

Generative Animations: A Multi-Model Pipeline for Prompt-Driven Motion Synthesis

生成式动画:面向提示驱动运动合成的多模型流水线

Mannat Khurana, Sanyam Jain, Rishav Agarwal

机构 * Canva Adobe

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

AI总结 提出一种结合大语言模型和分割模型的流水线,将自然语言提示自动转换为符合场景几何、深度遮挡和3D透视变换的动画运动路径。

Comments 5 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26576 2026-05-27 cs.CV cs.LG 62%

TrackRef3D: Multi-View Consistent Track-then-Label for Open-World Referring Segmentation in 3D Gaussian Splatting

TrackRef3D: 面向开放世界3D高斯泼溅分割的多视角一致跟踪-标注方法

Yuyang Tan, Renhe Zhang, Hang Zhang, Ao Li, Xin Tan

机构 * East China Normal University, Shanghai, China(华东师范大学,上海,中国) Shanghai AI Laboratory(上海人工智能实验室) University of Electronic Science and Technology of China, Chengdu, China(电子科技大学,成都,中国)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

AI总结 提出TrackRef3D全自动流水线,通过多视角一致跟踪-标注范式解耦目标发现与语义定位,无需人工标注实现开放世界3D高斯泼溅分割。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20505 2026-05-27 cs.AI cs.CL cs.LG 62%

LiPUP-MA: A Residential Experience-centric Multi-Agent Framework for Living-in-the-loop Participatory Urban Planning

LiPUP-MA:一种以居住体验为中心的循环参与式城市多智能体规划框架

Hang Ni, Yuzhi Wang, Yizhi Song, Hao Liu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 提出LiPUP-MA多智能体框架,通过模拟居住生活与体验驱动的计划修订循环,利用基于图的经验库和空间约束技能增强规划器,解决参与式城市规划中经验落地与反馈空间化问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27245 2026-05-27 cs.LG 57%

Symbolic Regression via Latent Iterative Refinement

通过潜在迭代细化的符号回归

Xieting Chu, Sriram Vishwanath, Vijay Ganesh

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 提出潜在方程嵌入(LEE)框架,通过迭代推断在功能基础化的潜在空间中缩小符号回归的推断差距,生成更简单且准确的表达式。

Comments Preprint. 21 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27116 2026-05-27 cs.CV 57%

COVD: Continual Open-Vocabulary Object Detection with Novel Concept Injection

COVD: 通过新概念注入的持续开放词汇目标检测

Yupeng Zhang, Ruize Han, Yuzhong Feng, Zixin Ren, Yuntong Tian, Liang Wan

机构 * Tianjin University(天津大学) Shenzhen University of Advanced Technology(深圳大学)

专题命中 视觉定位与Grounding :VLM(abstract_cn);分类 cs.CV

AI总结 提出持续开放词汇目标检测新任务COVD,通过冻结视觉编码器并仅更新文本分支参数注入新概念,实现无需额外参数的高效持续学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27009 2026-05-27 cs.LG 57%

SCENT: Aligning Mass Spectra with Molecular Structure for Olfactory Perception

SCENT: 将质谱与分子结构对齐用于嗅觉感知

Ziqi Zhang, Eunyeong Jin, Miguel Vasco, Farzaneh Taleb, Nona Rajabi, Alexandra Gutmann, Jonathan Williams, Antônio H. Ribeiro, Danica Kragic

机构 * Dept. of Intelligent Systems, KTH Royal Institute of Technology(智能系统系,皇家理工学院) Atmospheric Chemistry Dept., Max Planck Institute for Chemistry(大气化学部,马克斯·普朗克研究所) Dept. of Information Technology, Uppsala University(信息科技系,乌普萨拉大学) Science for Life Laboratory (SciLifeLab), Uppsala(生命科学实验室(SciLifeLab),乌普萨拉)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 提出SCENT多模态对比学习框架,通过将电子电离质谱表示与预训练化学结构嵌入对齐,在无需分子结构的情况下实现与结构模型相当的嗅觉预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26926 2026-05-27 cs.AI 57%

From Norms to Indicators (N2I-RAG): An Agentic Retrieval-Augmented Generation Framework for Legal Indicator Computation

从规范到指标 (N2I-RAG): 一种用于法律指标计算的智能检索增强生成框架

Youssef Al Mouatamid, Marie Bonnin, Jihad Zahir

机构 * LISI Laboratory(LISI实验室) Cadi Ayyad University(卡迪·阿亚德大学) Univ Brest(布列塔尼大学) IRD, Univ Brest, CNRS, Ifremer, LEMAR(IRD、布列塔尼大学、CNRS、Ifremer、LEMAR)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 提出N2I-RAG框架,通过自适应检索、基于LLM的智能体和验证机制,实现从法律文本到指标的透明、可追溯的自动计算,在法国海洋环境法语料库上优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26861 2026-05-27 cs.CV 57%

REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization

REVERSE: 强化证据验证与搜索的智能体图像地理定位

Yong Li, Furong Jia, Dacheng Yin, Kang Rong, Fengyun Rao, Jing Lyu, Fan Zhang

机构 * Peking University(北京大学) The Hong Kong University of Science and Technology(香港科技大学) WeChat Vision, Tencent Inc(腾讯公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 提出REVERSE框架,通过多轮智能体推理强化证据搜索与验证的交互,在图像地理定位任务中优于强检索增强基线,以4B模型媲美更大模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19186 2026-05-27 cs.AI 57%

Discoverable Agent Knowledge -- A Formal Framework for Agentic KG Affordances (Extended Version)

可发现的智能体知识——智能体知识图谱能力的形式化框架(扩展版)

Terry R. Payne, Valentina Tamma, Enrico Daga

机构 * School of Computer Science and Informatics, University of Liverpool, UK(利物浦大学计算机科学与信息学学院) Open University(开放大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出一个四维形式化框架(语义表达性、智能体可发现性、任务相对基础性和认知信任范围),并从中推导出智能体能力概况(AAP),作为VoID和DCAT之上的语义层,支持智能体在规划时进行原则性的知识图谱选择、组合和故障诊断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27045 2026-05-27 cs.CL 50%

ExTax: Explainable Disinformation Detection via Persuasion, Emotion, and Narrative Role Taxonomies

ExTax:基于说服、情感和叙事角色分类学的可解释虚假信息检测

Shang Luo, Yingguang Yang, Zhenchen Sun, Yang Liu, Bin Chong, Jingru Chen, Yancheng Chen, Jiayu Liang, Kefu Xu, Hao Peng, Philip S. Yu

机构 * Peking University(北京大学) University of Science and Technology of China(中国科学技术大学) North China University of Science and Technology(华北理工大学) Tsinghua University(清华大学) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) University of Chinese Academy of Sciences(中国科学院大学) Soochow University(苏州大学) Beihang University(北航) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 提出ExTax框架,统一说服修辞、情感操纵和叙事角色为17维分类空间,通过熵驱动动态标签平滑和多头注意力融合分类与上下文特征,实现可解释的虚假信息检测,在跨域基准上达到0.8456 Macro F1。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26604 2026-05-27 cs.GT cs.DC cs.NI econ.TH 50%

Credibility Trilemma in Polymatroidal Service Markets

多拟阵服务市场中的可信度三难困境

Lauri Lovén, Sujit Gujar, Kalle Timperi, Hassan Mehmood, Praveen Kumar Donta, Sasu Tarkoma, Schahram Dustdar

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文研究多拟阵服务市场中市场运营者的策略行为,证明在非模多拟阵上不存在同时满足收益最优、激励兼容和可信的静态密封投标机制,并引入不可信成本度量该困境的福利损失。

Comments 75 pages, 3 figures. Prepared for submission to the ACM Transactions on Economics and Computation (TEAC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26405 2026-05-27 cs.CL 50%

Towards Just-in-Time Adaptive Feedback: Enhancing Student Learning via Knowledge-Grounded LLM

面向即时自适应反馈:通过知识增强的大语言模型提升学生学习

Younghun Lee, Amir Bralin, Nobel Sanjay Rebello, Dan Goldwasser

机构 * Department of Computer Science(计算机科学系) Department of Physics and Astronomy(物理与天文学系) College of Education(教育学院)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 提出一个框架,利用领域专家知识增强大语言模型,在真实教学场景中提供即时自适应反馈,并在大规模大学课程中提升学生成绩超过80%。

Comments 8 pages, Accepted to 21st Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06580 2026-05-27 cs.CL 50%

Stylistic Evolution and LLM Neutrality in Singlish Language

新加坡英语中的文体演变与LLM中立性

Linus Tze En Foo, Weihan Angela Ng, Wenkai Li, Lynnette Hui Xian Ng

机构 * Independent Researcher(独立研究者) ETH Zürich(苏黎世联邦理工学院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 通过分析十年间非正式数字信息的文体变化,研究大型语言模型(LLM)能否生成时间中立的输出,发现文体可分离性随时间距离增加,且LLM在真实性和时间中立性之间存在结构性权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏