arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7360 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7360 篇

2507.19847 2025-07-30 cs.CV 79%

Knowledge Regularized Negative Feature Tuning of Vision-Language Models for Out-of-Distribution Detection

Wenjie Zhu, Yabin Zhang, Xin Jin, Wenjun Zeng, Lei Zhang

机构 * Hong Kong Polytechnic University(香港理工大学) Eastern Institute of Technology(东部技术研究所) Stanford University(斯坦福大学) Ningbo Institute of Digital Twin, Eastern Institute of Technology(宁波数字孪生研究所,东部技术研究所) Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-罗克琴堡研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒研究实验室)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20188 2025-07-29 cs.CV 79%

SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection

Mohammed-En-Nadhir Zighem, Abdenour Hadid

机构 * Sorbonne Center for Artificial Intelligence, Sorbonne University Abu Dhabi, UAE(索邦人工智能中心,阿布扎比分校,阿联酋)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11261 2025-07-29 cs.CV 79%

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition

Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He

机构 * South China University of Technology(华南理工大学) Guangdong Engineering Center for Large Model and GenAI Technology(广东省大模型与生成式人工智能技术工程中心) State Key Laboratory of Subtropical Building and Urban Science(亚热带建筑科学国家重点实验室) Ministry of Education Key Laboratory of Big Data and Intelligent Robot(教育部大数据与智能机器人重点实验室) Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information(广东省计算智能与网络信息重点实验室) Singapore Management University(新加坡国立大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19599 2025-07-29 cs.CV 79%

Object-centric Video Question Answering with Visual Grounding and Referring

Haochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai, Xiaolong Jiang, Yao Hu, Weidi Xie, Stratis Gavves

机构 * University of Amsterdam(阿姆斯特丹大学) SAI, Shanghai Jiao Tong University(上海交通大学SAI研究所) Xiaohongshu Inc(小红书公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01092 2025-07-23 cs.CV cs.RO eess.IV 79%

One-Shot Affordance Grounding of Deformable Objects in Egocentric Organizing Scenes

Wanjun Jia, Fan Yang, Mengfei Duan, Xianchi Chen, Yinxi Wang, Yiming Jiang, Wenrui Chen, Kailun Yang, Zhiyong Li

机构 * School of Artificial Intelligence and Robotics, Hunan University, China(人工智能与机器人学院,湖南大学) National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, China(机器人视觉感知与控制技术国家工程研究中心,湖南大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to IROS 2025. Source code and benchmark dataset will be publicly available at https://github.com/Dikay1/OS-AGDO

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14426 2025-07-22 cs.CV 79%

CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding

Zhou Chen, Joe Lin, Sathyanarayanan N. Aakur

机构 * Auburn University(亚伯拉罕大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to NeSy 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01558 2025-07-21 cs.HC cs.AI 79%

Visual Grounding Methods for Efficient Interaction with Desktop Graphical User Interfaces

El Hassane Ettifouri, Jessica López Espejel, Laura Minkova, Tassnim Dardouri, Walid Dahhane

机构 * Research and Innovation Lab, Novelis(创新研究实验室,Novelis)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Preprint submitted to Engineering Applications of Artificial Intelligence journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12232 2025-07-17 cs.CV 79%

MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM

Tao Chen, Jingyi Zhang, Decheng Liu, Chunlei Peng

机构 * Tao Chen, Jingyi Zhang, Decheng Liu, Chunlei Peng(作者)

专题命中 视觉定位与Grounding :VLM(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12123 2025-07-17 cs.CV 79%

Open-Vocabulary Indoor Object Grounding with 3D Hierarchical Scene Graph

Sergey Linok, Gleb Naumov

机构 * Moscow Institute of Physics and Technology(莫斯科物理技术学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 13 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07744 2025-07-11 cs.CV 79%

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding

David Pujol-Perich, Sergio Escalera, Albert Clapés

机构 * Universitat de Barcelona(巴塞罗那大学) Computer Vision Center(计算机视觉中心)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06719 2025-07-10 cs.CV cs.RO 79%

A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding

Zhenyang Liu, Sixiao Zheng, Siyu Chen, Cairong Zhao, Longfei Liang, Xiangyang Xue, Yanwei Fu

机构 * Fudan University, Shanghai Innovation Institute(复旦大学) Zhejiang University(浙江大学) Tongji University(同济大学) NeuHelium Co., Ltd(NeuHelium公司) Fudan University(复旦大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05424 2025-07-09 cs.CL cs.AI 79%

"Lost-in-the-Later": Framework for Quantifying Contextual Grounding in Large Language Models

Yufei Tao, Adam Hiatt, Rahul Seetharaman, Ameeta Agrawal

机构 * Portland State University(波特兰州立大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04741 2025-07-08 cs.CV 79%

Vision-Language Models Can't See the Obvious

Yasser Dahou, Ngoc Dung Huynh, Phuc H. Le-Khac, Wamiq Reyaz Para, Ankit Singh, Sanath Narayan

机构 * Technology Innovation Institute(技术创新研究所)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14607 2025-07-01 cs.CV 79%

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

Tianming Liang, Kun-Yu Lin, Chaolei Tan, Jianguo Zhang, Wei-Shi Zheng, Jian-Fang Hu

机构 * Sun Yat-sen University(中山大学) Southern University of Science and Technology(南方科技大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2025. Project page: \url{https://isee-laboratory.github.io/ReferDINO}

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22817 2025-07-01 cs.CV 79%

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding

Xingyilang Yin, Jiale Wang, Xi Yang, Mutian Xu, Xu Gu, Nannan Wang

机构 * Xidian University(西安电子科技大学) SSE, CUHKSZ(华南理工大学深圳校区)

专题命中 视觉定位与Grounding :VLM(title);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20384 2025-06-26 cs.AI 79%

Paladin-mini: A Compact and Efficient Grounding Model Excelling in Real-World Scenarios

Dror Ivry, Oran Nahum

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18476 2025-06-24 cs.CV 79%

Context Consistency Learning via Sentence Removal for Semi-Supervised Video Paragraph Grounding

Yaokun Zhong, Siyu Jiang, Jian Zhu, Jian-Fang Hu

机构 * School of Computer Science and Engineering, Sun Yat-sen University, China(中山大学计算机科学与工程学院) Guangdong Province Key Laboratory of Information Security Technology, China(广东省信息安全技术重点实验室) Key Laboratory of Machine Intelligence and Advanced Computing, Ministry of Education, China(教育部机器智能与高级计算重点实验室) Guangdong University of Foreign Studies(广东外语外贸大学) Guangdong University of Technology(广东工业大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16493 2025-06-23 cs.RO cs.AI 79%

Grounding Language Models with Semantic Digital Twins for Robotic Planning

Mehreen Naeem, Andrew Melnik, Michael Beetz

机构 * Institute for Artificial Intelligence, University of Bremen(人工智能研究所,不莱梅大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14495 2025-06-18 cs.CV 79%

I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs

Yu Qi, Lipeng Gu, Honghua Chen, Liangliang Nan, Mingqiang Wei

机构 * School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(计算机科学与技术学院,南京航空航天大学) Urban Data Science Section, Delft University of Technology(都市数据科学部门,代尔夫特理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14238 2025-06-18 cs.CV 79%

Unified Representation Space for 3D Visual Grounding

Yinuo Zheng, Lipeng Gu, Honghua Chen, Liangliang Nan, Mingqiang Wei

机构 * School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(计算机科学与技术学院,南京航空航天大学) Urban Data Science Section, Delft University of Technology(数据科学部,代尔夫特理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07896 2025-06-10 cs.AI cs.CL 79%

Evaluating Large Language Models on the Frame and Symbol Grounding Problems: A Zero-shot Benchmark

Shoko Oka

机构 * Independent Researcher(独立研究者)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 52 pages, Additional resources available on GitHub repository

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07671 2025-06-10 cs.CL cs.AI 79%

GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation

Ionut-Teodor Sorodoc, Leonardo F. R. Ribeiro, Rexhina Blloshmi, Christopher Davis, Adrià de Gispert

机构 * Amazon AGI(亚马逊人工智能研究院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments ACL 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20389 2025-06-10 cs.CV 79%

From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs

Ang Cao, Sergio Arnaud, Oleksandr Maksymets, Jianing Yang, Ayush Jain, Sriram Yenamandra, Ada Martin, Vincent-Pierre Berges, Paul McVay, Ruslan Partsey, Aravind Rajeswaran, Franziska Meier, Justin Johnson, Jeong Joon Park, Alexander Sax

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Project page: https://liftgs.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14853 2025-06-10 cs.CV 79%

Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection

Peipeng Yu, Jianwei Fei, Hui Gao, Xuan Feng, Zhihua Xia, Chip Hong Chang

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05890 2025-06-09 cs.CV 79%

Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation

Yiheng Li, Yang Yang, Zichang Tan, Huan Liu, Weihua Chen, Xu Zhou, Zhen Lei

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) CAIR, HKISI, Chinese Academy of Sciences(中国科学院计算机辅助研究部) School of Computer Science and Engineering, the Faculty of Innovation Engineering, M.U.S.T(慕斯科技大学计算机科学与工程学院) Sangfor Technologies Inc.(Sangfor技术有限公司) Beijing Jiaotong University(北京交通大学) Alibaba Group(阿里巴巴集团)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17433 2025-06-05 cs.AI 79%

MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models

Zhengyi Zhao, Shubo Zhang, Yuxi Zhang, Yanxi Zhao, Yifan Zhang, Zezhong Wang, Huimin Wang, Yutian Zhao, Bin Liang, Yefeng Zheng, Binyang Li, Kam-Fai Wong, Xian Wu

机构 * The Chinese University of Hong Kong(香港中文大学) University of International Relations(国际关系大学) Jarvis Research Center, Tencent YouTu Lab(腾讯YouTu实验室) Westlake University(西湖大学)

专题命中 视觉定位与Grounding :vision-language model(title);vision language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24103 2025-06-02 cs.CV 79%

Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic Priors

Peiran Xu, Yadong Mu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17143 2025-05-28 cs.AI 79%

ASP-based Multi-shot Reasoning via DLV2 with Incremental Grounding

Francesco Calimeri, Giovambattista Ianni, Francesco Pacenza, Simona Perri, Jessica Zangari

机构 * University of Calabria(卡布里亚大学) Gruppo Nazionale Calcolo Scientifico-Istituto Nazionale di Alta Matematica(国家科学计算集团-国家高级数学研究所) DLVSystem Srl(DLVSystem公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Under consideration in Theory and Practice of Logic Programming (TPLP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19242 2025-05-27 cs.CV 79%

Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model

Alaa Dalaq, Muzammil Behzad

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22436 2025-05-27 cs.CV 79%

NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving

Fuhao Li, Huan Jin, Bin Gao, Liaoyuan Fan, Lihui Jiang, Long Zeng

机构 * Tsinghua University(清华大学) Huawei Noah’s Ark Lab(华为诺亚实验室) The University of Hong Kong(香港大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏