arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7314 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7314 篇

2601.09449 2026-01-15 cs.CV 83%

PrivLEX: Detecting legal concepts in images through Vision-Language Models

PrivLEX:通过视觉-语言模型检测图像中的法律概念

Darya Baranouskaya, Andrea Cavallaro

机构 * EPFL(苏黎世联邦理工学院) Idiap Research Institute(Idiap研究机构)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV

AI总结 PrivLEX通过视觉-语言模型实现图像中法律概念的检测,无需显式标签即可进行可解释的隐私分类。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07761 2026-01-13 cs.CV 83%

Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding

视频证据推理:通过显式证据接地实现高效的视频理解

Yanxiang Huang, Guohua Gao, Zhaoyang Wei, Jianyuan Ni

机构 * Department of Applied Mathematics, The Hong Kong Polytechnic University(应用数学系,香港理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本文提出CoE框架,通过显式证据接地模块和强化学习优化的锚定协议,提升视频推理的准确性和可靠性。

Comments 6 pages

Journal ref ICME 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00730 2026-01-05 cs.CV 83%

Grading Handwritten Engineering Exams with Multimodal Large Language Models

用多模态大语言模型评分手写工程考试

Janez Perš, Jon Muhovič, Andrej Košir, Boštjan Murovec

机构 * University of Ljubljana, Faculty of Electrical Engineering(卢布尔雅那大学电子工程学院)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV

AI总结 本研究利用多模态大语言模型对斯洛文尼亚语手写工程考试进行自动评分,通过多阶段设计实现高可靠性,达到与人工评分约8分的均方差,验证了结构化提示和参考定位的重要性。

Comments 10 pages, 5 figures, 2 tables. Supplementary material available at https://lmi.fe.uni-lj.si/en/janez-pers-2/supplementary-material/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23219 2025-12-30 cs.CV 83%

MM-UAVBench: How Well Do Multimodal Large Language Models See, Think, and Plan in Low-Altitude UAV Scenarios?

MM-UAVBench: 多模态大语言模型在低空无人机场景中的感知、推理与规划能力如何?

Shiqi Dai, Zizhi Ma, Zhicong Luo, Xuesong Yang, Yibin Huang, Wanyue Zhang, Chi Chen, Zonghao Guo, Wang Xu, Yufei Sun, Maosong Sun

机构 * Tsinghua University(清华大学) Nankai University(南开大学) Northwest Polytechnical University(西北工业大学) Chinese Academy of Sciences(中国科学院) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 MM-UAVBench通过系统评估多模态大语言模型在低空无人机场景中的感知、推理与规划能力,揭示其在复杂任务中的性能瓶颈与改进方向。

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17601 2025-12-24 cs.CV 83%

HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly Detection

HeadHunt-VAD: 在MLLM中寻找鲁棒的异常敏感头部以实现无调优视频异常检测

Zhaolin Cai, Fan Li, Ziwei Zheng, Haixia Bi, Lijun He

专题命中 视觉定位与Grounding :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 HeadHunt-VAD通过直接在冻结的MLLM中寻找鲁棒的异常敏感头部,实现高效的无调优视频异常检测。

Comments AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13891 2025-12-24 cs.CV 83%

Weakly Supervised Ephemeral Gully Detection In Remote Sensing Images Using Vision Language Models

基于视觉语言模型的弱监督瞬时沟壑遥感图像检测

Seyed Mohamad Ali Tousi, Ramy Farag, John A. Lory, G. N. DeSouza

机构 * Vision Guided and Intelligent Robotics Laboratory (ViGIR)(视觉引导与智能机器人实验室) EECS Dept.(电子工程与计算机科学系) Division of Plant Science and Technology(植物科学与技术系)

专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract);分类 cs.CV

AI总结 本文提出基于视觉语言模型的弱监督瞬时沟壑检测方法,通过半监督学习和噪声感知损失函数提升遥感图像检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12715 2025-12-23 cs.CV cs.RO 83%

AsyMoE: Leveraging Modal Asymmetry for Enhanced Expert Specialization in Large Vision-Language Models

AsyMoE:利用模态不对称性提升大视觉-语言模型专家专业化

Heng Zhang, Haichuan Hu, Yaomin Shen, Weihao Yu, Yilei Yuan, Haochen You, Guo Cheng, Zijian Zhang, Lubin Gan, Huihui Wei, Hao Zhang, Jin Huang

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV

AI总结 AsyMoE通过三个专门专家组解决视觉-语言模态不对称问题,提升大模型专家专业化性能。

Comments This submission has been withdrawn by the authors due to a fundamental error in the methodology that affects the validity of the main results

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15949 2025-12-19 cs.CV 83%

The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs

感知观测站:多模态大语言模型的鲁棒性与基础性表征

Tejas Anvekar, Fenil Bardoliya, Pavan K. Turaga, Chitta Baral, Vivek Gupta

机构 * Arizona State University(亚利桑那州立大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 感知观测站通过系统性扰动和真实数据集,评估多模态大语言模型在视觉基础性和关系结构上的鲁棒性,揭示其在扰动下的表现,为模型分析提供系统基础。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11215 2025-12-15 cs.CV 83%

SmokeBench: Evaluating Multimodal Large Language Models for Wildfire Smoke Detection

SmokeBench: 评估多模态大语言模型用于野火烟雾检测

Tianye Qi, Weihao Li, Nick Barnes

机构 * Australian National University(澳大利亚国立大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV

AI总结 SmokeBench评估多模态大语言模型在野火烟雾检测中的性能,发现模型在烟雾定位方面存在显著局限,尤其在早期阶段。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02715 2025-12-03 cs.CV 83%

GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding

GeoViS: 基于地理奖励的视觉搜索用于遥感视觉定位

Peirong Zhang, Yidan Zhang, Luxiao Xu, Jinliang Lin, Zonghao Guo, Fengxiang Wang, Xue Yang, Kaiwen Wei, Lei Wang

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所) University of Chinese Academy of Sciences(中国科学院大学) Tsinghua University(清华大学) National University of Defense Technology(国防科技大学) Shanghai Jiao Tong University(上海交通大学) Chongqing University(重庆大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 GeoViS通过地理奖励机制,实现了遥感影像中的细粒度视觉定位,通过逐步搜索和推理提升小目标检测精度与跨领域泛化能力。

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00936 2025-12-02 cs.CV 83%

SceneProp: Combining Neural Network and Markov Random Field for Scene-Graph Grounding

SceneProp:结合神经网络和马尔可夫随机场进行场景图接地

Keita Otani, Tatsuya Harada

机构 * The University of Tokyo(东京大学) RIKEN AIP(日本科学技术研究所(RIKEN)先进研究所(AIP))

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

AI总结 SceneProp通过将场景图接地重新公式化为马尔可夫随机场的MAP推断问题,有效提升了复杂查询下的视觉接地性能。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00395 2025-12-02 cs.CV 83%

Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction

更优、更强、更快:在基于多模态大语言模型的分割中解决三重困境

Jiazhen Liu, Mingkuan Feng, Long Chen

机构 * The Hong Kong University of Science and Technology(香港科技大学)

专题命中 视觉定位与Grounding :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 STAMP通过同时预测文本和分割掩码,解决了多模态大语言模型在保持对话能力、高分割性能和快速推理间的三重困境。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00305 2025-12-02 cs.AI 83%

ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning

ChartPoint: 通过 grounding 反射引导 MLLMs 进行图表推理

Zhengzhuo Xu, SiNan Du, Yiyan Qi, SiwenLu, Chengjin Xu, Chun Yuan, Jian Guo

机构 * Tsinghua University(清华大学) International Digital Economy Academy(国际数字经济学院) Beihang University(北航) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.AI

AI总结 ChartPoint通过引入 grounding 反射机制,提升 MLLMs 在图表推理中的表现,开发了两个指令微调模型,在多个图表基准上取得显著提升。

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02607 2025-11-27 cs.CV cs.CL 83%

UniChange: Unifying Change Detection with Multimodal Large Language Model

UniChange: 通过多模态大语言模型统一变化检测

Xu Zhang, Danyang Li, Xiaohang Dong, Tianhao Wu, Hualong Yu, Jianye Wang, Qicheng Li, Xiang Li

机构 * TMCC, Computer Science, Nankai University(天津大学计算机学院,南开大学) VCIP, Computer Science, Nankai University(南开大学计算机科学系) NKIARI, Futian, Shenzhen(深圳福田国家工程研究中心) CMEE, Sichuan Agricultural University(四川农业大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 UniChange通过多模态大语言模型统一变化检测任务,引入特殊标记和文本提示,实现BCD和SCD的统一,并在多个基准测试中取得最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12140 2025-11-18 cs.CL cs.CV 83%

Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding

Pinxue Guo, Chongruo Wu, Xinyu Zhou, Lingyi Hong, Zhaoyu Chen, Jinglun Li, Kaixun Jiang, Sen-ching Samson Cheung, Wei Zhang, Wenqiang Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19110 2025-11-14 cs.CV 83%

LISA: A Layer-wise Integration and Suppression Approach for Hallucination Mitigation in Multimodal Large Language Models

Zhihui Guo, Xin Man, Hui Xu, Jie Shao, Zhiguo Jiang, Xianchao Zhang, Heng Tao Shen

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08173 2025-11-12 cs.CV 83%

VLMDiff: Leveraging Vision-Language Models for Multi-Class Anomaly Detection with Diffusion

Samet Hicsonmez, Abd El Rahman Shabayek, Djamila Aouada

机构 * University of Luxembourg(卢森堡大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06201 2025-11-11 cs.CV cs.HC 83%

Scene-Aware Urban Design: A Human-AI Recommendation Framework Using Co-Occurrence Embeddings and Vision-Language Models

Rodrigo Gallardo, Oz Fishman, Alexander Htet Kyaw

机构 * Department of Architecture(建筑系) Department of EECS(电子工程与计算机科学系) Massachusetts Institute of Technology(麻省理工学院)

专题命中 视觉定位与Grounding :vision-language model(title);vision language model(abstract);grounding(abstract);分类 cs.CV

Comments Accepted to NEURIPS 2025 Creative AI Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26781 2025-11-04 cs.CV 83%

ChartAB: A Benchmark for Chart Grounding & Dense Alignment

Aniruddh Bansal, Davit Soselia, Dang Nguyen, Tianyi Zhou

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27196 2025-11-03 cs.CL cs.AI 83%

MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models

Zixin Chen, Hongzhan Lin, Kaixin Li, Ziyang Luo, Yayue Deng, Jing Ma

机构 * Hong Kong Baptist University(香港 Baptist 大学) Beijing University of Posts and Telecommunications(北京邮电大学) National University of Singapore(新加坡国立大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23397 2025-10-28 cs.CV 83%

VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations

Lu Dong, Haiyu Zhang, Han Lin, Ziang Yan, Xiangyu Zeng, Hongjie Zhang, Yifei Huang, Yi Wang, Zhen-Hua Ling, Limin Wang, Yali Wang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Beihang University(北京航空航天大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23010 2025-10-24 cs.CV 83%

SeG-SR: Integrating Semantic Knowledge into Remote Sensing Image Super-Resolution via Vision-Language Model

Bowen Chen, Keyan Chen, Mohan Yang, Zhengxia Zou, Zhenwei Shi

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17034 2025-10-21 cs.CV 83%

Where, Not What: Compelling Video LLMs to Learn Geometric Causality for 3D-Grounding

Yutong Zhong

机构 * New York University(纽约大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16036 2025-10-21 cs.CV 83%

IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection

Zewen Li, Zitong Yu, Qilang Ye, Weicheng Xie, Wei Zhuo, Linlin Shen

机构 * School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) School of Computing and Information Technology, Great Bay University(大亚湾大学计算机与信息科技学院) College of Computer Science, Nankai University(南开大学计算机学院) School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室) National Engineering Laboratory of Big Data System Computing Technology, Shenzhen University(大数据系统计算技术国家工程实验室)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Instrumentation and Measurement (TIM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15866 2025-10-20 cs.CV cs.NE 83%

BiomedXPro: Prompt Optimization for Explainable Diagnosis with Biomedical Vision Language Models

Kaushitha Silva, Mansitha Eashwara, Sanduni Ubayasiri, Ruwan Tennakoon, Damayanthi Herath

机构 * University of Peradeniya(珀德尼亚大学) RMIT University(皇家墨尔本理工大学)

专题命中 视觉定位与Grounding :vision language model(title);vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments 10 Pages + 15 Supplementary Material Pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26165 2025-10-16 cs.CV 83%

Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models

Yuansen Liu, Haiming Tang, Jinlong Peng, Jiangning Zhang, Xiaozhong Ji, Qingdong He, Wenbin Wu, Donghao Luo, Zhenye Gan, Junwei Zhu, Yunhang Shen, Chaoyou Fu, Chengjie Wang, Xiaobin Hu, Shuicheng Yan

机构 * Yuan-Hou(元侯)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11852 2025-10-15 cs.LG 83%

Evaluating Open-Source Vision-Language Models for Multimodal Sarcasm Detection

Saroj Basnet, Shafkat Farabi, Tharindu Ranasinghe, Diptesh Kanoji, Marcos Zampieri

机构 * George Mason University(乔治·马歇尔大学) Lancaster University(兰卡斯特大学) University of Surrey(萨里大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);LLaVA(abstract);分类 cs.LG

Comments Accepted to ICDMW 2025 Workshop on Multimodal AI (MMAI). Full workshop info: https://icdmw25mmai.github.io/

Journal ref Proc. IEEE International Conference on Data Mining Workshops (ICDMW 2025), Workshop on Multimodal AI (MMAI 2025), Los Angeles, USA, December 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08307 2025-10-15 cs.CV cs.RO 83%

DSM: Constructing a Diverse Semantic Map for 3D Visual Grounding

Qinghongbing Xie, Zijian Liang, Fuhao Li, Long Zeng

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China(清华大学深圳国际研究生院,清华大学,深圳,中国)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract);分类 cs.CV

Comments 8 pages, 6 figures, Project Page: https://binicey.github.io/DSM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06427 2025-10-14 cs.CV 83%

When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection

Rabin Dulal, Lihong Zheng, Muhammad Ashad Kabir

机构 * School of Computing, Mathematics and Engineering(计算数学与工程学院) Charles Sturt University(查尔斯·斯特劳特大学) Gulbali Institute for Agriculture, Water and Environment(古尔巴利农业、水与环境研究所) Food Agility CRC Ltd(食品敏捷性CRC有限公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

Journal ref Australasian Joint Conference on Artificial Intelligence 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11368 2025-10-07 cs.CV 83%

From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation

Jingkun Chen, Haoran Duan, Xiao Zhang, Boyan Gao, Vicente Grau, Jungong Han

机构 * Department of Engineering Science, University of Oxford(工程科学系,牛津大学) Department of Automation, Tsinghua University(自动化系,清华大学) School of Information Science and Technology, Northwest University(信息科学与技术学院,西北大学)

专题命中 视觉定位与Grounding :vision language model(title);vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏