arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-03-18 至 2026-03-18 共收录 20 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 20 篇

2512.12633 2026-03-18 cs.CV cs.AI 90%

DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model

DiG:通过差异 grounding 提升多模态大语言模型的细粒度感知

Zhou Tao, Shida Wang, Yongxiang Hua, Haoyu Cao, Linli Xu

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI

AI总结 本文提出DiG框架,通过学习相似图像对的差异识别提升多模态大语言模型的细粒度感知能力,实验表明其在多个视觉感知基准上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16664 2026-03-18 cs.CV cs.AI 84%

Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation

Kestrel: 为降低LVLM幻觉而引入自反思

Jiawei Mao, Hardy Chen, Haoqin Tu, Yuhan Wang, Letian Zhang, Zeyu Zheng, Huaxiu Yao, Zirui Wang, Cihang Xie, Yuyin Zhou

机构 * UC Santa Cruz(加州大学圣克ruz分校) UC Berkeley(加州大学伯克利分校) UNC-Chapel Hill(北卡罗来纳大学教堂山分校) Apple(苹果公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 Kestrel提出一种无需训练的框架,通过显式视觉 grounding 与证据验证自反思机制减少LVLM幻觉,实验显示在POPE和MME-Hallucination基准上性能提升,同时提供透明的验证轨迹。

Comments 16 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16558 2026-03-18 cs.CV cs.MM 83%

Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models

基于分割的注意力熵:在大视觉-语言模型中检测和缓解对象幻觉

Jiale Song, Jiaxin Luo, Xue-song Tang, Kuangrong Hao, Mingbo Zhao

机构 * School of Information and Intelligent Science, Donghua University, Shanghai, 201620, China(信息与智能科学学院,东华大学,上海,201620,中国)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV

AI总结 本文提出基于分割的注意力熵(SAE),通过语义分割量化视觉注意力不确定性,设计可靠性评分和注意力调整方法,有效缓解大视觉-语言模型中的对象幻觉问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13353 2026-03-18 cs.CV cs.AI 81%

VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding

VideoITG: 多模态视频理解中的指导性时间定位

Shihao Wang, Guo Chen, De-an Huang, Zhiqi Li, Minghan Li, Guilin Liu, Jose M. Alvarez, Lei Zhang, Zhiding Yu

机构 * The Hong Kong Polytechnic Univ(香港理工大学) Nanjing Univ(南京大学) NVIDIA(NVIDIA公司) Harvard Univ(哈佛大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

AI总结 VideoITG通过设计VidThinker流水线,自动生成指令条件化描述,检索相关视频片段并选择关键帧,提升多模态视频理解任务的性能。

Comments Accepted by CVPR

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15154 2026-03-18 eess.IV cs.CV 74%

Vision-Language Model Based Multi-Expert Fusion for CT Image Classification

基于视觉-语言模型的多专家融合用于CT图像分类

Jianfa Bai, Kejin Lu, Runtian Yuan, Qingqiu Li, Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng

机构 * College of Computer Science and Artificial Intelligence, Shanghai Key Laboratory of Intelligent Information Processing, Fudan University(复旦大学计算机科学与人工智能学院,上海智能信息处理重点实验室) University of Oxford(牛津大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

AI总结 本文提出一种三阶段源感知多专家框架,通过构建肺部感知3D专家、开发MedSigLIP基专家和训练源分类器,提升多源CT图像中新冠检测的鲁棒性,实验结果显示在不同阶段模型在宏F1、ACC和AUC指标上均取得优异成绩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16538 2026-03-18 cs.CV cs.LG 73%

Streetscape Analysis with Generative AI (SAGAI): Vision-Language Assessment and Mapping of Urban Scenes

利用生成式AI进行街景分析(SAGAI):基于视觉-语言评估和城市场景制图

Joan Perez, Giovanni Fusco

机构 * Urban Geo Analytics, France(法国城市地理分析)

专题命中 视觉定位与Grounding :vision-language model(abstract);LLaVA(abstract);分类 cs.CV、cs.LG

AI总结 本文提出SAGAI,一种利用开放数据和视觉-语言模型分析街景的模块化流程,通过定制提示生成空间指标,实现城市场景的自动化制图与评估,展示了其在城市研究中的广泛应用潜力。

Comments 25 pages, 6 figures in main paper, 6 figures in appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13798 2026-03-18 cs.CV cs.AI cs.LG 67%

CFM: Language-aligned Concept Foundation Model for Vision

CFM:面向视觉的语言对齐概念基础模型

Kai Wittenmayer, Sukrut Rao, Amin Parchami-Araghi, Bernt Schiele, Jonas Fischer

机构 * Max Planck Institute for Informatics(马克斯·普朗克信息研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 CFM提出一种语言对齐的概念基础模型,提供细粒度可解释的概念,提升视觉任务的解释能力,实现分类、分割和描述生成的高性能表现。

Comments 53 pages, 29 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07923 2026-03-18 cs.CV cs.AI 62%

Exploring the Underwater World Segmentation without Extra Training

探索无需额外训练的水下世界分割

Bingyu Li, Tao Huo, Da Zhang, Zhiyuan Zhao, Junyu Gao, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI),中国电信,中国) University of Science and Technology of China, China(中国科学技术大学,中国) Northwestern Polytechnical University, China(西北工业大学,中国)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出AquaOV255水下分割数据集及Earth2Ocean框架,通过几何引导视觉掩码生成和类别-视觉语义对齐模块,在无需额外训练的情况下实现高效的水下开放词汇分割。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02869 2026-03-18 cs.CY cs.AI cs.CV 62%

Representing Beauty: Towards a Participatory but Objective Latent Aesthetics

表现美:迈向一种参与但客观的潜在美学

Alexander Michael Rusnak

机构 * Digital Humanities Laboratory(数字人文实验室) École Polytechnique Fédérale de Lausanne(瑞士联邦理工学院洛桑分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

AI总结 本文探讨神经网络如何表征美,通过跨模型表征收敛研究,发现美能产生更相似的表示,而不美的图像则不能。研究提出美的现实基础,并强调人类感知与创作在塑造深度学习模型潜在空间中的作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06993 2026-03-18 cs.AI cs.CV 62%

IMAIA: Interactive Maps AI Assistant for Travel Planning and Geo-Spatial Intelligence

IMAIA:交互式地图AI助手用于旅行规划和地理空间智能

Jieren Deng, Zhizhang Hu, Ziyan He, Aleksandar Cvetkovic, Pak Kiu Chung, Dragomir Yankov, Chiqun Zhang

机构 * Microsoft(微软) Amazon(亚马逊)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

AI总结 IMAIA通过自然语言交互整合矢量地图和卫星影像,结合地理空间信息提升地图相关问答和摄像头到地点的关联能力,实现空间感知的对话式地图系统。

Comments Accepted to The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15699 2026-03-18 cs.PF cs.AI cs.SE 57%

This Is Taking Too Long -- Investigating Time as a Proxy for Energy Consumption of LLMs

这需要太长时间 -- 调查时间作为LLMs能耗的代理

Lars Krupp, Daniel Geißler, Francisco M. Calatrava-Nicolas, Vishal Banwari, Paul Lukowicz, Jakob Karolus

机构 * Centre for Applied Autonomous Sensor Systems (AASS)(应用自主传感器系统中心)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 研究通过推理时间估算API基于LLMs的能耗,验证了时间测量与实际能耗之间的关联,为用户理解LLMs能耗提供方法。

Comments This work was accepted at PerCom 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16207 2026-03-18 cs.AI 57%

Proactive Rejection and Grounded Execution: A Dual-Stage Intent Analysis Paradigm for Safe and Efficient AIoT Smart Homes

主动拒绝与 grounded 执行:一种双阶段意图分析范式用于安全高效的 AIoT 智能家居

Xinxin Jin, Zhengwei Ni, Zhengguo Sheng, Victor C. M. Leung

机构 * School of Information and Electronic Engineering (Sussex Artificial Intelligence Institute), Zhejiang Gongshang University(信息与电子工程学院(Sussex人工智能研究院),浙江工商大学) Sussex Artificial Intelligent Institute, Zhejiang Gongshang University(Sussex人工智能研究院,浙江工商大学) Department of Engineering and Design, University of Sussex(工程与设计系, Sussex大学) Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(人工智能研究院,深圳MSU-BIT大学) College of Computer Science and Software Engineering, Shenzhen University(计算机科学与软件工程学院,深圳大学) Department of Electrical and Computer Engineering, The University of British Columbia(电气与计算机工程系,不列颠哥伦比亚大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出双阶段意图感知框架,通过语义防火墙和确定性验证器,解决LLM在IoT中的可靠性与交互效率问题,提升家居任务执行精度与用户干扰最小化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15957 2026-03-18 cs.LG 57%

GASP: Guided Asymmetric Self-Play For Coding LLMs

GASP:指导性的非对称自我对弈用于编码大语言模型

Swadesh Jana, Cansu Sancaktar, Tomáš Daniš, Georg Martius, Antonio Orvieto, Pavel Kolev

机构 * University of Tübingen(图宾根大学) Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) ELLIS Institute Tübingen(图宾根ELLIS研究所) Tübingen AI Center(图宾根人工智能中心)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 GASP通过引入真实数据目标问题,改进了非对称自我对弈方法,提升了LiveCodeBench上的pass@20指标,并解决了其他方法无法达到的难题。

Comments Accepted at ICLR 2026 Workshop on AI with Recursive Self-Improvement (RSI 2026) as Spotlight, and ICLR 2026 Workshop on Lifelong Agents (LLA 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00512 2026-03-18 cs.CV 57%

Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding

基于语义边界的小波帧选择用于长视频理解

Wang Chen, Yuhui Zeng, Yongdong Luo, Tianyu Xie, Luojun Lin, Jiayi Ji, Yan Zhang, Xiawu Zheng

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学) College of Computer and Data Science, Fuzhou University(福州大学计算机与数据科学学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出WFS-SB方法,通过小波变换检测语义边界,提升长视频中大视觉语言模型的性能,实验显示在多个基准上均优于现有方法。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19570 2026-03-18 cs.CV 57%

VALD: Multi-Stage Vision Attack Detection for Efficient LVLM Defense

VALD:为高效LVLM防御的多阶段视觉攻击检测

Nadav Kadvil, Malak Fares, Ayellet Tal

机构 * Technion – Israel Institute of Technology(技术学院 – 以色列理工学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出VALD方法,通过图像变换与智能数据整合快速过滤清洁输入,结合文本嵌入空间和LLM解决攻击问题,实现高效准确的LVLM防御。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19476 2026-03-18 cs.LG 57%

LogicXGNN: Grounded Logical Rules for Explaining Graph Neural Networks

LogicXGNN: 基于逻辑规则的图神经网络解释方法

Chuqin Geng, Ziyu Zhao, Zhaoyue Wang, Haolin Ye, Yuhe Jiang, Xujie Si

机构 * University of Toronto(多伦多大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 LogicXGNN通过构建逻辑规则提升图神经网络解释的可信度和可解释性,提出数据驱动的 fidelity 指标并显著提升解释质量。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16253 2026-03-18 cs.CV 57%

Open-Vocabulary Octree-Graph for 3D Scene Understanding

开放词汇八叉树图用于3D场景理解

Zhigang Wang, Yifei Su, Chenhui Li, Dong Wang, Yan Huang, Bin Zhao, Xuelong Li

机构 * Northwestern Polytechnical University(西北工业大学) Shanghai AI Laboratory(上海人工智能实验室) University of Chinese Academy of Sciences(中国科学院大学) CASIA TeleAI

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出Octree-Graph,通过CGSM和IFA算法获取3D实例及语义特征,构建适应性八叉树结构以高效表示场景,实验表明其在多种任务中具有广泛适用性和有效性。

Comments Accepted by ICCV25. 11 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15712 2026-03-18 cond-mat.mtrl-sci cs.AI 57%

LLM-Driven Discovery of High-Entropy Catalysts via Retrieval-Augmented Generation

通过检索增强生成发现高效熵催化剂

AI Scientists, Xinyi Lin, Danqing Yin, Ying Guo

机构 * School of Biomedical Sciences, Li Ka Shing Faculty of Medicine, The University of Hong Kong(香港大学生物医学科学学院,利卡申医学学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文展示如何利用大语言模型加速催化剂发现,通过检索增强生成框架生成250多个候选催化剂,其中82%具有热力学稳定性,同时满足多目标约束,最佳催化剂在性能和成本上均有显著提升。

Journal ref Open Conference of AI Agents for Science 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15677 2026-03-18 cs.CL 50%

MedArena: Comparing LLMs for Medicine-in-the-Wild Clinician Preferences

MedArena: 比较医疗领域LLM的临床医生偏好

Eric Wu, Kevin Wu, Jason Hom, Paul H. Yi, Angela Zhang, Alejandro Lozano, Jeff Nirschl, Jeff Tangney, Kevin Byram, Braydon Dymm, Narender Annapureddy, Eric Topol, David Ouyang, James Zou

机构 * Department of Electrical Engineering, Stanford University(斯坦福大学电气工程系) Department of Biomedical Data Science, Stanford University(斯坦福大学生物医学数据科学系) Division of Hospital Medicine, Department of Medicine, Stanford School of Medicine(斯坦福医学院医学部住院医学科) Department of Radiology, St. Jude Children's Research Hospital(圣 Jude 儿童研究医院放射科) University of California, San Francisco(旧金山大学) Department of Pathology and Laboratory Medicine, University of Wisconsin School of Medicine and Public Health(威斯康星大学医学与公共卫生学院病理学与实验室医学系) Doximity, San Francisco, CA, USA(Doximity公司) Department of Medicine, Division of Rheumatology and Immunology, Vanderbilt University Medical Center(范德比尔特大学医学中心医学系风湿病与免疫学科) Department of Neurology, Charleston Area Medical Center(查尔斯顿医疗中心神经科) Department of Translational Medicine, Scripps Research Translational Institute(斯克里普斯研究转化研究所转化医学系) Kaiser Permanente Division of Research(凯撒医疗集团研究部)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 MedArena通过真实临床问题和医生偏好比较LLM,揭示临床实用性与基准性能的差异,强调可读性和临床细节的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13713 2026-03-18 cs.CL 50%

On Theoretically-Driven LLM Agents for Multi-Dimensional Discourse Analysis

关于理论驱动的LLM代理在多维话语分析中的应用

Maciej Uberna, Michał Wawer, Jarosław A. Chudziak, Marcin Koszowy

机构 * Laboratory of The New Ethos, Warsaw University of Technology, Poland(新伦理实验室,华沙理工大学,波兰) Faculty of Electronics and Information Technology, Warsaw University of Technology, Poland(电子与信息技术学院,华沙理工大学,波兰)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出一种多代理框架,通过引入显式理论知识提升话语分析中改写策略的识别能力,实验表明理论增强的LLM代理在检测强化和泛化等任务上表现显著优于基线模型,宏F1分数提升近30%。

Comments 8 pages, 4 figures, 3 tables. This is the accepted version of the paper presented at the 18th International Conference on Agents and Artificial Intelligence (ICAART 2026), Marbella, Spain

Journal ref Proceedings of the 18th International Conference on Agents and Artificial Intelligence (ICAART 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏