arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7314 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7314 篇

2507.07939 2025-07-23 cs.CL 85%

SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware Alignment

Guoxin Zang, Xue Li, Donglin Di, Lanshun Nie, Dechen Zhan, Yang Song, Lei Fan

机构 * Harbin Institute of Technology(哈尔滨工业大学) University of New South Wales(新南威尔士大学)

专题命中 视觉定位与Grounding :visual language model(title);vision-language model(abstract);VLM(abstract);visual reasoning(abstract)

Comments Accepted by ACMMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14538 2025-04-02 eess.IV cs.AI cs.CV cs.LG 85%

Vision-Language Models for Acute Tuberculosis Diagnosis: A Multimodal Approach Combining Imaging and Clinical Data

Ananya Ganapthy, Praveen Shastry, Naveen Kumarasami, Anandakumar D, Keerthana R, Mounigasri M, Varshinipriya M, Kishore Prasath Venkatesh, Bargava Subramanian, Kalyan Sivasailam

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08722 2025-03-18 cs.CV cs.AI cs.LG 85%

A Recipe for Improving Remote Sensing VLM Zero Shot Generalization

Aviad Barzilai, Yotam Gigi, Amr Helmy, Vered Silverman, Yehonathan Refael, Bolous Jaber, Tomer Shekel, George Leifman, Genady Beryozkin

专题命中 视觉定位与Grounding :VLM(title,abstract);visual language model(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06912 2025-03-04 cs.CV cs.AI cs.LG 85%

Compositional Entailment Learning for Hyperbolic Vision-Language Models

Avik Pal, Max van Spengler, Guido Maria D'Amely di Melendugno, Alessandro Flaborea, Fabio Galasso, Pascal Mettes

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted as oral paper at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04559 2024-10-28 cs.CL cs.AI cs.CV cs.LG 85%

Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition

Aditya K Surikuchi, Raquel Fernández, Sandro Pezzelle

专题命中 视觉定位与Grounding :grounding(title,abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG

Comments In proceedings of EMNLP 2024 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19696 2024-05-01 cs.CV cs.AI cs.CL cs.LG 85%

Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners

Chun Feng, Joy Hsu, Weiyu Liu, Jiajun Wu

专题命中 视觉定位与Grounding :grounding(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI、cs.LG

Comments CVPR 2024. The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10093 2023-11-08 cs.CV cs.AI cs.CL cs.LG 85%

Investigating the Role of Attribute Context in Vision-Language Models for Object Recognition and Detection

Kyle Buettner, Adriana Kovashka

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted at Winter Conference on Applications of Computer Vision (WACV), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.06403 2022-07-14 cs.CV cs.AI cs.CL cs.GR cs.LG 85%

3D Concept Grounding on Neural Fields

Yining Hong, Yilun Du, Chunru Lin, Joshua B. Tenenbaum, Chuang Gan

专题命中 视觉定位与Grounding :grounding(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Project page: http://3d-cg.csail.mit.edu

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15925 2025-11-19 cs.CV cs.AI 84%

MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Classification and Detection

Andrea Moglia, Elia Clement Nastasio, Luca Mainardi, Pietro Cerveri

机构 * Department of Electronics, Information, and Bioengineering(电子、信息与生物工程系) Polytechnic University of Milan(米兰理工学院) Department of Industrial, and Information Engineering(工业与信息工程系) University of Pavia(帕维亚大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Journal ref Moglia, A., Nastasio, E.C., Mainardi, L. et al. MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Observation and Localization in CT Images. J Healthc Inform Res (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19294 2025-10-01 cs.CV cs.AI cs.CL 84%

Object Detection with Multimodal Large Vision-Language Models: An In-depth Review

Ranjan Sapkota, Manoj Karkee

专题命中 视觉定位与Grounding :vision-language model(title,abstract);vision language model(abstract);分类 cs.CV、cs.AI

Comments First Peer Reviewed Review Paper for Object Detection with Vision-Language Models (VLMs)

Journal ref Information Fusion, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09480 2025-04-15 cs.CV cs.AI 84%

Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation

Yongchao Feng, Yajie Liu, Shuai Yang, Wenrui Cai, Jinqing Zhang, Qiqi Zhan, Ziyue Huang, Hongxi Yan, Qiao Wan, Chenguang Liu, Junzhe Wang, Jiahui Lv, Ziqi Liu, Tengyuan Shi, Qingjie Liu, Yunhong Wang

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments A Review and Evaluation about Vision-Language Model for Object Detection and Segmentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01525 2024-07-18 cs.CV cs.AI cs.CL 84%

ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities

Chenming Zhu, Tai Wang, Wenwei Zhang, Kai Chen, Xihui Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted by ECCV 2024. A comprehensive and hierarchical 3D reasoning grounding benchmark in the era of foundation models. Project page: https://zcmax.github.io/projects/ScanReason

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.05704 2024-04-24 cs.CV cs.AI cs.CL 84%

Visual Grounding Methods for VQA are Working for the Wrong Reasons!

Robik Shrestha, Kushal Kafle, Christopher Kanan

专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI

Comments Published in ACL 2020 under the title "A negative case analysis of visual grounding methods for VQA"

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.15462 2024-01-09 cs.CV cs.CL cs.LG 84%

Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations

Ziyan Yang, Kushal Kafle, Franck Dernoncourt, Vicente Ordonez

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.LG

Comments CVPR 2023. Fix ReferIt results. Code: https://github.com/uvavision/AMC-grounding Project Webpage: https://vislang.ai/amc

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25467 2026-08-12 cs.CV 版本更新 84%

GridVAD: Open-Set Video Anomaly Detection via Spatial Reasoning over Stratified Frame Grids

GridVAD: 通过分层帧网格上的空间推理实现开放集视频异常检测

Mohamed Eltahir, Ahmed O. Ibrahim, Obada Siralkhatim, Tabarak Abdallah, Sondos Mohamed

机构 * King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学) Independent Researcher(独立研究员) National Center for Research (NCR)(国家研究中心)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);grounding(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV

AI总结 GridVAD通过分层帧网格的空间推理生成开放集候选描述,结合自一致性巩固和空间跟踪模块实现像素级异常掩码,其在UCSD Ped2上达到77.59的Pixel-AUROC,优于其他方法。

Comments Accepted at the Large-scale Video Object Segmentation (LVOS) Workshop in conjunction with ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09147 2026-08-11 cs.CV 新提交 84%

RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection

RefineAny3D:作为语义对齐的深度细化用于单目3D检测

Zhihao Zhang, Gengwei Zhang, Tianlong Chen, Xiaoming Liu

机构 * Michigan State University(密歇根州立大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 RefineAny3D将单目3D检测的深度细化转化为视觉对齐问题,通过VLM实现无需数值预测的深度修正,在多类检测工具上均有性能提升且可泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12715 2026-08-05 cs.CV 版本更新 84%

VLC Fusion: Vision-Language Conditioned Sensor Fusion for Robust Object Detection

VLC Fusion:面向鲁棒目标检测的视觉-语言条件传感器融合

Aditya Taparia, Noel Ngu, Mario Leiva, Joshua Shay Kricheli, John Corcoran, Nathaniel D. Bastian, Gerardo Simari, Paulo Shakarian, Ransalu Senanayake

机构 * Arizona State University(亚利桑那州立大学) Department of Computer Science and Engineering, Universidad Nacional del Sur and Institute for Computer Science and Engineering(计算机科学与工程系,国家南方大学和计算机科学与工程研究所) U.S. Department of Defense(美国国防部) United States Military Academy(美国军事学院) Syracuse University(雪城大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本文提出VLC Fusion视觉-语言条件传感器融合框架,利用VLM捕捉环境线索动态调整模态权重,在多传感器融合目标检测任务中,于自动驾驶与军事目标数据集上较传统方法实现更优性能。

Comments 27 pages, 20 figures, Accepted for presentation at ECML PKDD 2026, Shortlisted for the Best Research Track Student Paper Award

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01113 2026-08-04 cs.CV 新提交 84%

CoT-Edit: Let CoT Guide Instruction Video Editing

CoT-Edit:让思维链(CoT)指导指令视频编辑

Sen Liang, Fengbin Guan, Youliang Zhang, Xin Li, Zhibo Chen

机构 * University of Science and Technology of China(中国科学技术大学) Zhongguancun Academy(中关村学院) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出CoT-Edit的plan--guide--edit框架,以CoT增强的MLLM为规划器生成空间先验,结合扩散编辑器实现高保真指令视频编辑,性能优于多个基准方法

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00232 2026-08-04 cs.CV 新提交 84%

Real-Time Visual Obstruction Detection in Surgical Augmented Reality

手术增强现实中的实时视觉遮挡检测

Shih-Chin Yang, Yanming Xiu, Hanting Ye, Qi Chen, Elias Rotondo, Maria Gorlatova

机构 * Duke University(杜克大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 针对手术AR虚拟内容遮挡手术器械的问题,提出结合VLM与分割推理的延迟感知流水线,构建伪AR基准,实现87.43%准确率、479 ms延迟,较云端基线降延迟62.90%。

Comments ISMAR 2026 Mecidal Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14675 2026-07-24 cs.RO cs.AI 版本更新 84%

An Intelligent-Cloud Edge Multimodal Interaction System for Robots

一种用于机器人的智能云边缘多模态交互系统

Zihan Guo, Xiaoqi Li

机构 * Hainan University(海南大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.AI

AI总结 针对复杂环境下资源受限机器人的交互问题,提出云边缘多模态交互框架,集成增强YOLO手势检测器与LLM、VLM智能体,改进手势检测方法,经实验验证该系统在手势检测精度、任务成功率及用户满意度方面表现良好,证明了方法的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25763 2026-06-25 cs.CV 新提交 84%

ShutterMuse: Capture-Time Photography Guidance with MLLMs

ShutterMuse: 基于多模态大语言模型的拍摄时刻摄影指导

Jiayu Li, Yixiao Fang, Tianyu Hu, Wei Cheng, Ping Huang, Zheheng Fan, Gang Yu, Xingjun Ma

机构 * Fudan University(复旦大学) StepFun

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 提出CaptureGuide-Bench基准测试,评估MLLM在摄影师构图与主体姿态推荐方面的能力,并构建ShutterMuse统一模型,通过监督与强化微调实现最佳综合性能。

Comments Project Page:https://lijayutnt.github.io/ShutterMuse

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18609 2026-06-18 cs.CV 新提交 84%

Hallucination Detection and Correction in Medical VLMs via Counter-Evidence Verification

基于反事实证据验证的医学视觉语言模型幻觉检测与纠正

Nan Zhou, Ke Zou, Meng Liu, Linchao He, Jiaqi Zhu, Yi Zhang, Hu Chen, Huazhu Fu

机构 * College of Computer Science, Sichuan University(四川大学计算机科学学院) Yong Loo Lin School of Medicine, National University of Singapore(新加坡国立大学杨潞龄医学院) Key Laboratory of Data Protection and Intelligent Management, Ministry of Education, Sichuan University(四川大学数据保护与智能管理教育部重点实验室) National Key Laboratory of Autonomous Intelligent Unmanned Systems, Beijing Institute of Technology(北京理工大学自主智能无人系统国家重点实验室) Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局高性能计算研究所)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract_cn);vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 提出CoEV框架,通过文本与视觉证据的双向验证检测并纠正医学VLM幻觉,无需重新训练,在四个数据集上显著提升检测和纠正性能。

Comments MICCAI 2026 Accept. Submission Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17433 2026-06-17 cs.CV 新提交 84%

LADBench: A Benchmark for Logical Fault Detection in Images

LADBench: 图像中逻辑故障检测的基准

Sahasra Kondapalli, Lara Radovanovic, Aadi Palnitkar, Mingyang Mao, Xiaomin Lin

机构 * University of South Florida(南佛罗里达大学)

专题命中 视觉定位与Grounding :VLM(summary_cn);vision language model(abstract);visual question answering(abstract);grounding(abstract)

AI总结 提出LAD-Bench基准,包含1000多张合成图像的四域逻辑异常,通过分层提示协议评估模型,揭示现有VLM在隐式逻辑故障检测上的不足。

Comments Accepted to the IEEE International Conference on Development and Learning (ICDL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10528 2026-06-05 cs.CV 84%

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs

BareBones: 视觉语言模型中零样本几何理解的基准测试

Aaditya Baranwal, Vishal Yadav, Abhishek Rajora

机构 * University of Central Florida(佛罗里达大学中央分校) University of Calgary(卡尔加里大学)

专题命中 视觉定位与Grounding :LLaVA(abstract,abstract_cn);vision-language model(abstract);VLM(abstract_cn);grounding(abstract)

AI总结 提出BareBones基准,通过去除RGB纹理仅保留轮廓,测试26个视觉语言模型在零样本几何形状理解上的表现,发现模型存在严重的纹理偏置悬崖。

Comments Accepted at CVPR (13th FGVC Workshop) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14799 2026-05-27 cs.CV cs.CR cs.SI 84%

Can Visual Mamba Improve AI-Generated Image Detection? An In-Depth Investigation

视觉Mamba能否提升AI生成图像检测?一项深入研究

Mamadou Keita, Wassim Hamidouche, Hessen Bougueffa Eutamene, Abdelmalik Taleb-Ahmed, Xianxun Zhu, Abdenour Hadid

机构 * Laboratory of IEMN, CNRS, Centrale Lille, UMR 8520, Univ. Polytechnique Hauts-de-France(伊姆纳实验室,国家科学研究中心,里尔中央理工大学,UMR 8520,法国高等技术大学) Khalifa University(卡利法大学) School of Communication and Information Engineering, Shanghai University(上海大学通信与信息工程学院) Sorbonne Center for Artificial Intelligence, Sorbonne University Abu Dhabi(索邦人工智能中心,索邦大学阿布扎克分校)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本研究系统评估了Vision Mamba模型在AI生成图像检测中的性能,与CNN、ViT和VLM检测器进行对比,分析了准确性、效率和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23304 2026-05-25 cs.CV 84%

General Hazard Detection

通用危险检测

Stephanie Ng, CP Lim, SueJen Looi, Hendrik Zurlinden, David Nguyen, Lei Wei, Saeid Nahavandi, Hailing Zhou

机构 * Swinburne University of Technology(斯winburne大学) National Transport Research Organisation(国家交通运输研究组织) Google Cloud(谷歌云) Deakin University(德金大学)

专题命中 视觉定位与Grounding :LLaVA(abstract,abstract_cn);vision-language model(abstract);VLM(abstract_cn);visual reasoning(abstract)

AI总结 针对现有危险检测系统在抽象安全概念上的局限性,提出基于规则合规评估的CompliVision数据集和结合LLaVA视觉推理与人在回路反馈的通用危险检测框架。

Comments 20 pages, 7 figures and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15325 2026-05-18 cs.CV 84%

COPRA: Conditional Parameter Adaptation with Reinforcement Learning for Video Anomaly Detection

COPRA:基于强化学习的条件参数适应用于视频异常检测

Darryl Cherian Jacob, Xinyu Liu, Kai Wang, Pan He

机构 * Auburn University(奥本大学) Tencent Hunyuan(腾讯文元)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 COPRA通过生成输入特定的参数更新,动态适应冻结的VLM,提升视频异常检测的适应性和泛化能力,同时拓展到多选视频问答和密集标注等任务。

Comments Manuscript currently under review for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05008 2026-05-15 cs.CV 84%

Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation

多模态因果驱动表示学习用于通用化医学图像分割

Xusheng Liang, Lihua Zhou, Nianxin Li, Miao Xu, Ziyang Song, Dong Yi, Jinlin Wu, Jiawei Ma, Hongbin Liu, Zhen Lei, Jiebo Luo

机构 * City University of Hong Kong(香港城市大学) Shenzhen Loop Area Institute(深圳河套学院) CAIR, HKISI, Chinese Academy of Sciences(中国科学院计算智能研究所) UESTC(电子科技大学) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本文提出MCDRL框架,结合因果推理与VLM解决医学图像分割的领域泛化问题,通过文本提示识别病变区域并消除领域特定影响,提升分割准确性与泛化能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12237 2026-05-13 cs.CV 84%

UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs

UHR-Micro: 诊断和缓解地球观测VLMs中的分辨率错觉

Shuo Ni, Tong Wang, Jing Zhang, He Chen, Haonan Guo, Ning Zhang, Bo Du

机构 * National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing(国家空间智能信息处理科技重点实验室) Beijing Institute of Technology(北京理工大学) School of Computer Science(计算机学院) Wuhan University(武汉大学) Zhongguancun Academy(中关村学院) State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing(测绘遥感信息工程国家重点实验室) Hong Kong Polytechnic University(香港理工大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract_cn);vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 UHR-Micro通过11253条指令和1212张超高清图像,评估VLM在高分辨率地球观测图像中的空间极限表现,揭示高分辨率输入下空间定位和证据解析的失败,并提出MAP代理以证据为中心提升微尺度感知。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09598 2026-05-13 cs.CV 84%

SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

SoccerLens: 超越准确性的 grounded 球赛视频理解

Ismael Elsharkawi, Ahmed Sait, Silvio Giancola, Bernard Ghanem, Hossam Sharara, Abdelrahman Eldesokey

机构 * Department of Computer Science and Engineering, The American University in Cairo(美国亚历山大大学计算机科学与工程系) Image And Visual Understanding Lab (IVUL), KAUST(卡塔尔大学图像与视觉理解实验室)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本文提出 SoccerLens 基准测试,评估球赛视频理解中视觉 grounding 的有效性,发现现有模型在准确率高但 grounding 表现差,揭示了时空复杂领域中 grounded 评估的重要性。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏