arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7360 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7360 篇

2512.12246 2025-12-16 cs.CV 79%

Moment and Highlight Detection via MLLM Frame Segmentation

通过MLLM帧分割进行时刻和亮点检测

I Putu Andika Bagas Jiwanta, Ayu Purwarianti

机构 * School of Electrical Engineering and Informatics(电气工程与信息学院) Institut Teknologi Bandung(万隆技术大学)

专题命中 视觉定位与Grounding :MLLM(title,abstract);分类 cs.CV

AI总结 本文提出通过MLLM帧分割实现视频时刻和亮点检测,利用分割损失与因果LM损失结合,提升检测精度与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09117 2025-12-11 cs.AI 79%

A Categorical Analysis of Large Language Models and Why LLMs Circumvent the Symbol Grounding Problem

大型语言模型的范畴分析及为何LLMs绕过了符号 grounding 问题

Luciano Floridi, Yiyang Jia, Fernando Tohmé

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本文通过范畴分析,探讨了LLMs如何绕过符号 grounding 问题,而非解决它。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20409 2025-12-11 cs.LO cs.AI 79%

A Unified Formal Theory on the Logical Limits of Symbol Grounding

符号接地逻辑界限的统一正式理论

Zhangchi Liu

机构 * Zhangchi Liu(独立研究者)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本文提出符号接地问题的统一理论,通过形式证明展示其逻辑界限,指出外部动态过程与具身交互的必要性。

Comments 13 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00518 2025-12-10 cs.CV cs.CL 79%

Fine-grained Spatiotemporal Grounding on Egocentric Videos

细粒度眼动视频时空定位

Shuo Liang, Yiwu Zhong, Zi-Yuan Hu, Yeyao Tao, Liwei Wang

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出EgoMask基准和EgoMask-Train数据集,针对眼动视频细粒度时空定位的挑战,通过微调提升模型性能。

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05996 2025-12-09 cs.CV cs.CY cs.RO eess.IV 79%

FishDetector-R1: Unified MLLM-Based Framework with Reinforcement Fine-Tuning for Weakly Supervised Fish Detection, Segmentation, and Counting

FishDetector-R1: 基于统一MLLM框架的弱监督鱼类检测、分割与计数强化微调方法

Yi Liu, Jingyu Song, Vedanth Kallakuri, Katherine A. Skinner

机构 * University of Michigan(密歇根大学)

专题命中 视觉定位与Grounding :MLLM(title,abstract);分类 cs.CV

AI总结 FishDetector-R1通过统一MLLM框架和强化学习微调,实现了弱监督下的鱼类检测、分割与计数的高效准确提升。

Comments 18 pages, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05482 2025-12-08 cs.CV 79%

Concept-based Explainable Data Mining with VLM for 3D Detection

基于概念的可解释数据挖掘与VLM用于3D检测

Mai Tsujimoto

机构 * The University of Tokyo(东京大学)

专题命中 视觉定位与Grounding :VLM(title);vision-language model(abstract);分类 cs.CV

AI总结 本文提出基于概念的可解释数据挖掘方法,利用VLMs识别稀有物体以提升3D检测性能,减少标注负担并提高模型效果。

Comments 28 pages including appendix. Code: https://github.com/mm1129/concept_based_rare_detector_2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04231 2025-12-05 cs.RO cs.AI 79%

CRAFT-E: A Neuro-Symbolic Framework for Embodied Affordance Grounding

CRAFT-E:一种用于具身赋能力量接地的神经符号框架

Zhou Chen, Joe Lin, Carson Bulgin, Sathyanarayanan N. Aakur

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 CRAFT-E通过结合神经符号方法和具身感知,提供了一种可解释且可定制的框架,用于辅助机器人系统中基于赋能力量的物体选择。

Comments 20 pages. 3 figures, 4 tables. Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10194 2025-12-02 cs.CV 79%

B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding

B2N3D: 从二元关系到N元关系的3D物体接地的渐进式学习

Feng Xiao, Hongbin Xu, Hai Ci, Wenxiong Kang

机构 * School of Automation Science and Engineering, South China University of Technology(自动化科学与工程学院,华南理工大学) ByteDance Seed(字节跳动种子) Show Lab, National University of Singapore(Show Lab,新加坡国立大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 B2N3D通过引入N元关系学习提升3D物体接地的准确性,利用分组监督损失和混合注意力机制实现更精确的多模态关系建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23151 2025-12-01 cs.CV 79%

Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding

学习拒绝:用于视频时间定位中处理难相关查询的拒绝感知强化微调

Jin-Seop Lee, SungJoon Lee, SeongJun Jung, Boyang Li, Jee-Hyong Lee

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出RA-RFT方法,通过拒绝感知强化微调提升视频时间定位中对难相关查询的处理能力。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23867 2025-11-25 cs.CV 79%

Sim-DETR: Unlock DETR for Temporal Sentence Grounding

Sim-DETR: 解锁用于时间句子定位的DETR

Jiajin Tang, Zhengxuan Wei, Yuchen Zhu, Cheng Shi, Guanbin Li, Liang Lin, Sibei Yang

机构 * ShanghaiTech University(上海科技大学) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 Sim-DETR通过改进DETR的解码器层,解决了时间句子定位中的查询冲突问题,提升了模型性能。

Comments This work is accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18883 2025-11-24 cs.CV 79%

Universal Video Temporal Grounding with Generative Multi-modal Large Language Models

通用视频时间定位与生成多模态大语言模型

Zeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang, Yanfeng Wang, Weidi Xie

机构 * SAI, Shanghai Jiao Tong University(上海交通大学SAI实验室) ByteDance Seed(字节跳动种子)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出UniTime模型,利用生成多模态大语言模型实现通用视频时间定位,有效处理多类型视频并提升VideoQA任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15720 2025-11-21 cs.AI 79%

Automated Hazard Detection in Construction Sites Using Large Language and Vision-Language Models

利用大语言和视觉-语言模型进行施工工地自动危险检测

Islem Sahraoui

机构 * University of Houston Cullen College of Engineering Department of Civil and Environmental Engineering(德克萨斯大学休斯顿分校库伦工程学院土木与环境工程系)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.AI

AI总结 本研究利用大语言和视觉-语言模型,通过分析文本和图像数据,提高施工工地的安全隐患检测效率。

Comments Master thesis, University of Houton

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10241 2025-11-21 cs.CV 79%

TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding

TubeRMC: 基于互约束的管状重建用于弱监督空间-时间视频定位

Jinxuan Li, Yi Zhang, Jian-Fang Hu, Chaolei Tan, Tianming Liang, Beihao Xia

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 TubeRMC通过引入互约束的管状重建方法,在弱监督条件下提升空间-时间视频定位的准确性和一致性。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15379 2025-11-20 cs.CV 79%

Zero-Shot Open-Vocabulary Human Motion Grounding with Test-Time Training

零样本开放词汇人体运动接地与测试时训练

Yunjiao Zhou, Xinyan Chen, Junlang Qian, Lihua Xie, Jianfei Yang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出ZOMG框架,通过语言语义分区和软掩码优化,在无需标注或微调的情况下实现零样本开放词汇的人体运动接地,提升了运动理解的性能和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13924 2025-11-19 cs.CV 79%

Start Small, Think Big: Curriculum-based Relative Policy Optimization for Visual Grounding

Qingyang Yan, Guangyao Chen, Yixiong Zou

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments AAAI 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11438 2025-11-17 cs.CV 79%

VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models

Mingjie Xu, Jinpeng Chen, Yuzhi Zhao, Jason Chun Lok Li, Yue Qiu, Zekang Du, Mengyang Wu, Pingping Zhang, Kun Li, Hongzheng Yang, Wenao Ma, Jiaheng Wei, Qinbin Li, Kangcheng Liu, Wenqiang Lei

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

Comments This is the extended version of the paper accepted at AAAI 2026, which includes all technical appendices and additional experimental details

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10705 2025-11-17 cs.AI cs.CL 79%

Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents

Yuan Zhao, Hualei Zhu, Tingyu Jiang, Shen Li, Xiaohang Xu, Hao Henry Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10837 2025-11-13 cs.AI 79%

Exploring the Paradigm Shift from Grounding to Skolemization for Complex Query Answering on Knowledge Graphs

Yuyin Lu, Hegang Chen, Shanrui Xie, Yanghui Rao, Haoran Xie, Fu Lee Wang, Qing Li

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) School of Data Science, Lingnan University(岭南大学数据科学学院) School of Science and Technology, Hong Kong Metropolitan University(香港理工大学科技学院) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06908 2025-11-11 cs.CV cs.MM 79%

Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding

Yuzhen Li, Min Liu, Zhaoyang Li, Yuan Bian, Xueping Wang, Erbo Zhai, Yaonan Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21844 2025-11-11 cs.CV 79%

Test-Time Adaptation of Vision-Language Models for Open-Vocabulary Semantic Segmentation

Mehrdad Noori, David Osowiechi, Gustavo Adolfo Vargas Hakim, Ali Bahri, Moslem Yazdanpanah, Sahar Dastani, Farzad Beizaee, Ismail Ben Ayed, Christian Desrosiers

机构 * LIVIA, ÉTS Montréal, Canada International Laboratory on Learning Systems (ILLS)(LIVIA,蒙特利尔ÉTS,加拿大国际学习系统实验室)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02044 2025-11-11 cs.CV 79%

A multi-modal vision-language model for generalizable annotation-free pathology localization

Hao Yang, Hong-Yu Zhou, Jiarun Liu, Weijian Huang, Cheng Li, Zhihuan Li, Yuanxu Gao, Qiegen Liu, Yong Liang, Qi Yang, Song Wu, Tao Tan, Hairong Zheng, Kang Zhang, Shanshan Wang

机构 * Paul C. Lauterbur Research Center for Biomedical Imaging(生物医学成像研究中心) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) Pengcheng Laboratory(鹏城实验室) University of Chinese Academy of Sciences(中国科学院大学) Chinese Medicine Guangdong Laboratory(广东中医药实验室) Beijing Chaoyang Hospital, Capital Medical University(首都医科大学北京朝阳医院)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22672 2025-10-29 cs.CV cs.CL cs.RO 79%

Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views

Anna Deichler, Jonas Beskow

机构 * KTH Royal Institute of Technology(皇家理工学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 10 pages, 6 figures, 2 tables. Accepted to the NeurIPS 2025 Workshop on SPACE in Vision, Language, and Embodied AI (SpaVLE). Dataset: https://huggingface.co/datasets/annadeichler/KTH-ARIA-referential

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15256 2025-10-29 cs.CV 79%

Normal and Abnormal Pathology Knowledge-Augmented Vision-Language Model for Anomaly Detection in Pathology Images

Jinsol Song, Jiamu Wang, Anh Tien Nguyen, Keunho Byeon, Sangjeong Ahn, Sung Hak Lee, Jin Tae Kwak

机构 * Korea University(韩国大学) The Catholic University of Korea(韩国天主大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2025. Code is available at: https://github.com/QuIIL/ICCV2025_Ano-NAViLa

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03201 2025-10-28 cs.CV 79%

AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding

Yidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao, Nicu Sebe

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) University of Trento(特伦托大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08216 2025-10-28 cs.AI 79%

Grounding Methods for Neural-Symbolic AI

Rodrigo Castellano Ontiveros, Francesco Giannini, Marco Gori, Giuseppe Marra, Michelangelo Diligenti

机构 * University of Siena(锡耶纳大学) Scuola Normale Superiore(正规大学) KU Leuven(卢森堡大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Journal ref Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence (IJCAI-25), pp. 4806-4814, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21076 2025-10-28 cs.CV 79%

DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding

Weihao Xuan, Junjue Wang, Heli Qi, Zihang Chen, Zhuo Zheng, Yanfei Zhong, Junshi Xia, Naoto Yokoya

机构 * The University of Tokyo(东京大学) RIKEN AIP(理化学研究所AIP) Waseda University(早稻田大学) Wuhan University(武汉大学) Stanford University(斯坦福大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21083 2025-10-27 cs.CV 79%

Knowledge-Driven Vision-Language Model for Plexus Detection in Hirschsprung's Disease

Youssef Megahed, Atallah Madi, Dina El Demellawy, Adrian D. C. Chan

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted into the ICAAI 2025 - The 9th International Conference on Advances in Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19678 2025-10-24 cs.CL cs.CV 79%

Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs

Hao Fang, Changle Zhou, Jiawei Kong, Kuofeng Gao, Bin Chen, Shu-Tao Xia

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视觉定位与Grounding :grounding(title);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18582 2025-10-23 cs.CV 79%

The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers

Daiqing Qi, Handong Zhao, Jing Shi, Simon Jenni, Yifei Fan, Franck Dernoncourt, Scott Cohen, Sheng Li

机构 * University of Virginia(弗吉尼亚大学) Adobe(Adobe公司)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

Journal ref CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12718 2025-10-23 cs.CV cs.MM 79%

ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding

Zhenxing Zhang, Yaxiong Wang, Lechao Cheng, Zhun Zhong, Dan Guo, Meng Wang

机构 * School of Computer Science and Information Engineering, Hefei University of Technology, China(计算机科学与信息工程学院,合肥工业大学,中国) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China(人工智能研究所,合肥综合性国家科学中心,中国)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏