arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7360 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7360 篇

2509.05154 2025-09-08 eess.IV cs.CV 79%

VLSM-Ensemble: Ensembling CLIP-based Vision-Language Models for Enhanced Medical Image Segmentation

Julia Dietlmeier, Oluwabukola Grace Adegboro, Vayangi Ganepola, Claudia Mazo, Noel E. O'Connor

机构 * Insight Research Ireland Centre for Data Analytics, DCU, Dublin, Ireland(爱尔兰洞察研究爱尔兰数据分析中心,都柏林大学,都柏林) Research Ireland Centre for Research Training in Machine Learning, DCU, Dublin, Ireland(爱尔兰研究爱尔兰机器学习研究培训中心,都柏林大学,都柏林) School of Computing, Dublin City University, Dublin, Ireland(计算学院,都柏林城市大学,都柏林)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Medical Imaging with Deep Learning (MIDL 2025) short paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04676 2025-09-08 cs.AI cs.HC 79%

An Approach to Grounding AI Model Evaluations in Human-derived Criteria

Sasha Mitts

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 4 figures, 6 pages, presented at CHI 2025 Workshop on Human-AI Interaction for Augmented Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02807 2025-09-04 cs.CV 79%

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?

Mennatullah Siam

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Work under review in NeurIPS 2025 with the title "Are we using Motion in Referring Segmentation? A Motion-Centric Evaluation"

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19493 2025-09-04 cs.CR cs.CV 79%

Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents

Zhixin Lin, Jungang Li, Shidong Pan, Yibo Shi, Yue Yao, Dongliang Xu

专题命中 视觉定位与Grounding :MLLM(title);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01197 2025-09-04 cs.CV cs.RO 79%

A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding

Zhan Shi, Song Wang, Junbo Chen, Jianke Zhu

机构 * College of Software Technology, Zhejiang University(浙江大学软件技术学院) College of Computer Science, Zhejiang University(浙江大学计算机科学学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12553 2025-09-04 cs.CV 79%

ViDDAR: Vision Language Model-Based Task-Detrimental Content Detection for Augmented Reality

Yanming Xiu, Tim Scargill, Maria Gorlatova

机构 * Duke University(杜克大学)

专题命中 视觉定位与Grounding :vision language model(title,abstract);分类 cs.CV

Comments The paper has been accepted to the 2025 IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR), and selected for publication in the 2025 IEEE Transactions on Visualization and Computer Graphics (TVCG) special issue

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01401 2025-08-29 cs.CV 79%

Language-to-Space Programming for Training-Free 3D Visual Grounding

Boyu Mi, Hanqing Wang, Tai Wang, Yilun Chen, Jiangmiao Pang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19165 2025-08-27 cs.CV 79%

Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual Grounding

Yuzhen Li, Min Liu, Yuan Bian, Xueping Wang, Zhaoyang Li, Gen Li, Yaonan Wang

机构 * School of Artificial Intelligence and Robotics, Hunan University(人工智能与机器人学院,湖南大学) College of Information Science and Engineering, Hunan Normal University(信息科学与工程学院,湖南师范大学) School of Informatics, University of Edinburgh(信息学院,爱丁堡大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18372 2025-08-27 cs.CV 79%

OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding

Hieu Nguyen, Phuc-Tan Nguyen, Thien-Phuc Tran, Minh-Quang Nguyen, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le

机构 * University of Science(科学大学) University of Dayton(戴维森大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04958 2025-08-26 cs.CV cs.MM 79%

Boosting Temporal Sentence Grounding via Causal Inference

Kefan Tang, Lihuo He, Jisheng Dang, Xinbo Gao

机构 * School of Electronic Engineering, Xidian University Xi'an China School of Information Science \& Engineering, Lanzhou University Lanzhou China Xidian University Lanzhou University

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11903 2025-08-19 cs.CV 79%

OVG-HQ: Online Video Grounding with Hybrid-modal Queries

Runhao Zeng, Jiaqi Mao, Minghao Lai, Minh Hieu Phan, Yanjie Dong, Wei Wang, Qi Chen, Xiping Hu

机构 * Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(人工智能研究院,深圳MSU-BIT大学) University of Adelaide(阿德莱德大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11350 2025-08-18 cs.CV 79%

HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model

Zhenhao Zhang, Hanqing Wang, Xiangyu Zeng, Ziyu Cheng, Jiaxin Liu, Haoyu Yan, Zhirui Liu, Kaiyang Ji, Tianxiang Gui, Ke Hu, Kangyi Chen, Yahao Fan, Mokai Pan

专题命中 视觉定位与Grounding :multimodal large language model(title);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10976 2025-08-18 cs.AI 79%

Grounding Rule-Based Argumentation Using Datalog

Martin Diller, Sarah Alice Gaggl, Philipp Hanisch, Giuseppina Monterosso, Fritz Rauschenbach

机构 * Logic Programming and Argumentation Group, TU Dresden, Germany(图灵编程与论证组,德累斯顿理工大学,德国) Knowledge-Based Systems Group, TU Dresden, Germany(知识系统组,德累斯顿理工大学,德国) DIMES - University of Calabria, Italy(迪梅斯-卡拉布里亚大学,意大利)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04058 2025-08-18 cs.CV 79%

LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding

Feng Xiao, Hongbin Xu, Guocan Zhao, Wenxiong Kang

机构 * School of Automation Science and Engineering, South China University of Technology(自动化科学与工程学院,华南理工大学) School of Future Technology, South China University of Technology(未来技术学院,华南理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00372 2025-08-15 cs.CV 79%

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning

Zhixi Cai, Fucai Ke, Simindokht Jahangard, Maria Garcia de la Banda, Reza Haffari, Peter J. Stuckey, Hamid Rezatofighi

机构 * Monash University(莫纳什大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08633 2025-08-13 cs.AI cs.LO 79%

Diminution: On Reducing the Size of Grounding ASP Programs

HuanYu Yang, Fengming Zhu, YangFan Wu, Jianmin Ji

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14594 2025-08-12 cs.CV 79%

Solving Zero-Shot 3D Visual Grounding as Constraint Satisfaction Problems

Qihao Yuan, Kailai Li, Jiaming Zhang

机构 * Bernoulli Institute for Mathematics, Computer Science and Artificial Intelligence, University of Groningen(格罗宁根大学伯努利学院) Computer Vision for Human-Computer Interaction Lab (cv:hci), Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院人机交互计算机视觉实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05684 2025-08-11 cs.CR cs.LG 79%

MM-FusionNet: Context-Aware Dynamic Fusion for Multi-modal Fake News Detection with Large Vision-Language Models

Junhao He, Tianyu Liu, Jingyuan Zhao, Benjamin Turner

机构 * Huaiyin Institute of Technology(淮阴职业技术学院) Universidad Autónoma de Asunción(阿斯unción自治大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04546 2025-08-07 cs.CV 79%

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding

Minghang Zheng, Yuxin Peng, Benyuan Sun, Yi Yang, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王宣计算机技术研究所) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) Central Media Technology Institute, Huawei(华为中央媒体技术研究所)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04299 2025-08-07 cs.CV 79%

Length Matters: Length-Aware Transformer for Temporal Sentence Grounding

Yifan Wang, Ziyi Liu, Xiaolong Sun, Jiawei Wang, Hongmin Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03654 2025-08-06 cs.CL cs.CV 79%

Can Large Vision-Language Models Understand Multimodal Sarcasm?

Xinyu Wang, Yue Zhang, Liqiang Jing

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 视觉定位与Grounding :vision-language model(title);visual language model(abstract);分类 cs.CV

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02479 2025-08-05 cs.CV 79%

Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding

Xinquan Yu, Wei Lu, Xiangyang Luo

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01699 2025-08-05 cs.CV 79%

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding

Zuhao Yang, Yingchen Yu, Yunqing Zhao, Shijian Lu, Song Bai

机构 * Nanyang Technological University(南洋理工大学) ByteDance Inc.(字节跳动公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01402 2025-08-05 cs.CV 79%

ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models

Chuangchuang Tan, Jinglu Wang, Xiang Ming, Renshuai Tao, Yunchao Wei, Yao Zhao, Yan Lu

机构 * Beijing Jiaotong University(北京交通大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04351 2025-08-01 cs.CV 79%

LidaRefer: Context-aware Outdoor 3D Visual Grounding for Autonomous Driving

Yeong-Seung Baek, Heung-Seon Oh

机构 * School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19947 2025-07-31 cs.RO cs.CL cs.IT cs.LG cs.SY eess.SY math.IT 79%

Spatial Language Likelihood Grounding Network for Bayesian Fusion of Human-Robot Observations

Supawich Sitdhipol, Waritwong Sukprasongdee, Ekapol Chuangsuwanich, Rina Tse

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

Comments Accepted to the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC); Supplementary video: https://cu-asl.github.io/fp-lgn/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04047 2025-07-31 cs.CV 79%

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

Ziyu Zhu, Xilin Wang, Yixuan Li, Zhuofan Zhang, Xiaojian Ma, Yixin Chen, Baoxiong Jia, Wei Liang, Qian Yu, Zhidong Deng, Siyuan Huang, Qing Li

机构 * Tsinghua University(清华大学) Beijing Institute of Technology(北京理工大学) Beihang University(北航) State Key Laboratory of General Artificial Intelligence, BIGAI, China(国家一般人工智能重点实验室, BIGAI, 中国)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Embodied AI; 3D Vision Language Understanding; ICCV 2025 Highlight; https://mtu3d.github.io; Spatial intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21507 2025-07-30 cs.CV cs.MM 79%

VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding

Shibo Gao, Peipei Yang, Yangyang Liu, Yi Chen, Han Zhu, Xuyao Zhang, Linlin Huang

机构 * Beijing Jiaotong University(北京交通大学) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 21 pages, 19 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21450 2025-07-30 cs.CV cs.RO 79%

Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation

Bolei Chen, Jiaxu Kang, Yifei Wang, Ping Zhong, Qi Wu, Jianxin Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Submitted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21080 2025-07-30 cs.CL cs.AI 79%

Which symbol grounding problem should we try to solve?

Vincent C. Müller

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Journal ref (2015) Journal of Experimental and Theoretical Artificial Intelligence, 27 (1, ed. D. Jones & A. Beavers), 73-78

详情

展开后加载摘要…

URL PDF HTML 收藏