arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7387 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7387 篇

2508.17667 2025-08-26 cs.CV cs.AI 62%

Hierarchical Vision-Language Learning for Medical Out-of-Distribution Detection

Runhe Lai, Xinhua Lu, Kanghao Chen, Qichao Chen, Wei-Shi Zheng, Ruixuan Wang

机构 * Peng Cheng Laboratory(鹏城实验室) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Nottingham Malaysia(诺丁汉大学(马来西亚)) Key Laboratory of Machine Intelligence and Advanced Computing, MOE(机器智能与高级计算重点实验室)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 2 figures, Accepted by MICCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16643 2025-08-26 cs.LG cs.AI 62%

From Classical Probabilistic Latent Variable Models to Modern Generative AI: A Unified Perspective

Tianhua Chen

机构 * School of Computing and Engineering University of Huddersfield(计算与工程学院赫德斯菲尔德大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments This is a substantially improved and expanded version of an earlier manuscript hosted on SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5244929

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16157 2025-08-25 cs.CV cs.AI 62%

Beyond Human-prompting: Adaptive Prompt Tuning with Semantic Alignment for Anomaly Detection

Pi-Wei Chen, Jerry Chun-Wei Lin, Wei-Han Chen, Jia Ji, Zih-Ching Chen, Feng-Hao Yeh, Chao-Chun Chen

机构 * Silesian University of Technology(西里西亚技术大学) National Cheng Kung University(国立成功大学) Nvidia AI Technology Center(Nvidia AI技术中心)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03847 2025-08-22 cs.LG cs.AI 62%

KEA Explain: Explanations of Hallucinations using Graph Kernel Analysis

Reilly Haskins, Benjamin Adams

机构 * Department of Computer Science and Software Engineering, University of Canterbury(计算机科学与软件工程系,坎特伯雷大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04107 2025-08-20 cs.CV cs.AI 62%

Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder

Jingchao Wang, Zhijian Wu, Dingjiang Huang, Yefeng Zheng, Hong Wang

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13223 2025-08-20 cs.CV cs.AI 62%

MIRAGE: Towards AI-Generated Image Detection in the Wild

Cheng Xia, Manxi Lin, Jiexiang Tan, Xiaoxiong Du, Yang Qiu, Junjun Zheng, Xiangheng Kong, Yuning Jiang, Bo Zheng

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10935 2025-08-19 cs.CV cs.LG cs.RO 62%

HQ-OV3D: A High Box Quality Open-World 3D Detection Framework based on Diffision Model

Qi Liu, Yabei Li, Hongsong Wang, Lei He

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11256 2025-08-18 cs.CV cs.AI 62%

Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception

Junjie Wang, Keyu Chen, Yulin Li, Bin Chen, Hengshuang Zhao, Xiaojuan Qi, Zhuotao Tian

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments arXiv admin note: text overlap with arXiv:2505.04410

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10771 2025-08-15 cs.CV cs.AI 62%

AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences

Jieyu Li, Xin Zhang, Joey Tianyi Zhou

机构 * National University of Singapore(新加坡国立大学) Agency for Science, Technology and Research(科技研究局)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Proceedings of the 33rd ACM International Conference on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10556 2025-08-15 cs.CV cs.AI 62%

Retrieval-Augmented Prompt for OOD Detection

Ruisong Han, Zongbo Han, Jiahao Zhang, Mingyue Cheng, Changqing Zhang

机构 * College of Intelligence and Computing, Tianjin University, Tianjin, China(智能与计算学院,天津大学,天津,中国) State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10397 2025-08-15 cs.CV cs.AI 62%

PQ-DAF: Pose-driven Quality-controlled Data Augmentation for Data-scarce Driver Distraction Detection

Haibin Sun, Xinghui Song

机构 * College of Computer Science and Engineering, Shandong University of Science and Technology(计算机科学与工程学院,山东科技大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00873 2025-08-15 cs.LG cs.CV 62%

Unifying Self-Supervised Clustering and Energy-Based Models

Emanuele Sansone, Robin Manhaeve

机构 * KU Leuven(卢森堡大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

Comments Changes from previous version: change introductions and added acknowledgments. Integral version of workshop paper arXiv:2309.15420. Improved GEDI version (from two stages to single stage training) arxiv:2212.13425 - ACCEPTED TO TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09715 2025-08-14 cs.CV cs.LG 62%

NEURAL: Attention-Guided Pruning for Unified Multimodal Resource-Constrained Clinical Evaluation

Devvrat Joshi, Islem Rekik

机构 * BASIRA Lab, Imperial-X (I-X) and Department of Computing, Imperial College London(BASIRA实验室、Imperial-X(I-X)及帝国理工学院计算机系)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14153 2025-08-14 cs.CV cs.AI cs.CL 62%

Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions

Lucas Möller, Pascal Tilli, Ngoc Thang Vu, Sebastian Padó

机构 * University of Stuttgart(斯图加特大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted at Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11747 2025-08-13 cs.CV cs.AI 62%

OE3DIS: Open-Ended 3D Point Cloud Instance Segmentation

Phuc D. A. Nguyen, Minh Luu, Anh Tran, Cuong Pham, Khoi Nguyen

机构 * Movian AI Posts & Telecommunications Inst. of Tech(电信技术研究所)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments Accepted at ICCVW'25 - OpenSUN3D: 5th Workshop on Open-World 3D Scene Understanding with Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24277 2025-08-11 cs.LG cs.AI 62%

Evaluating and Designing Sparse Autoencoders by Approximating Quasi-Orthogonality

Sewoong Lee, Adam Davies, Marc E. Canby, Julia Hockenmaier

机构 * Siebel School of Computing and Data Science University of Illinois Urbana-Champaign(计算与数据科学学院 耶鲁大学伊利诺伊分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01932 2025-08-05 cs.CV cs.AI 62%

Proactive Disentangled Modeling of Trigger-Object Pairings for Backdoor Defense

Kyle Stein, Andrew A. Mahyari, Guillermo Francia, Eman El-Sheikh

机构 * Department of Intelligent Systems and Robotics, University of West Florida(智能系统与机器人系,西佛罗里达大学) Florida Institute For Human and Machine Cognition (IHMC)(佛罗里达人类与机器认知研究所) Center for Cybersecurity, University of West Florida(网络安全中心,西佛罗里达大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Journal ref Computers, Materials & Continua, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01338 2025-08-05 cs.CV cs.AI 62%

Weakly-Supervised Image Forgery Localization via Vision-Language Collaborative Reasoning Framework

Ziqi Sheng, Junyan Wu, Wei Lu, Jiantao Zhou

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04154 2025-08-05 cs.CV cs.AI 62%

CA-W3D: Leveraging Context-Aware Knowledge for Weakly Supervised Monocular 3D Detection

Chupeng Liu, Runkai Zhao, Weidong Cai

机构 * School of Computer Science, Faculty of Engineering, The University of Sydney(计算机科学学院,工程学院,悉尼大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06566 2025-08-04 cs.CL cs.AI cs.LG 62%

Natural Language Interaction with a Household Electricity Knowledge-based Digital Twin

Carolina Fortuna, Vid Hanžel, Blaž Bertalanič

机构 * 1 Department of Communication Systems, Jo z ef Stefan Institute, Slovenia Jozef Stefan Institute, Ljubljana, Slovenia

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments Accepted at IEEE SmartGridComm'24

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22802 2025-07-31 cs.CV cs.AI 62%

Advancing Fetal Ultrasound Image Quality Assessment in Low-Resource Settings

Dongli He, Hu Wang, Mohammad Yaqub

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫莫德·宾·扎耶德人工智能大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted to the MICCAI 2025 MIRASOL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04858 2025-07-31 cs.AI cs.LG 62%

Don't Lag, RAG: Training-Free Adversarial Detection Using RAG

Roie Kazoom, Raz Lapid, Moshe Sipper, Ofer Hadar

机构 * Electrical and Computer Engineering, Ben Gurion University, Beer Sheba 84105, Israel(电子与计算机工程系,本· Gurion 大学) Computer Science, Ben Gurion University, Beer Sheba 84105, Israel(计算机科学系,本· Gurion 大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI、cs.LG

Comments Accepted at VecDB @ ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21353 2025-07-30 cs.CV cs.LG 62%

Group Relative Augmentation for Data Efficient Action Detection

Deep Anil Patel, Iain Melvin, Zachary Izzo, Martin Renqiang Min

机构 * NEC Laboratories America(NEC美洲实验室)

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20913 2025-07-29 cs.CV cs.AI 62%

HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection

Jialei Cui, Jianwei Du, Yanzhe Li, Lei Gao, Hui Jiang, Chenfu Bao

机构 * Baidu Inc.(百度公司) Southeast University(东南大学) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21226 2025-07-29 cs.CV cs.AI 62%

MemeBLIP2: A novel lightweight multimodal system to detect harmful memes

Jiaqi Liu, Ran Tong, Aowei Shen, Shuzheng Li, Changlin Yang, Lisha Xu

机构 * Mathematics and Statistics Department, University of Texas at Dallas(德克萨斯大学达拉斯分校数学与统计学系)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 11 pages, 3 figures. Accepted at the First Workshop on Multimodal Knowledge and Language Modeling (MKLM), IJCAI-25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05211 2025-07-28 cs.CV cs.AI 62%

All in One: Visual-Description-Guided Unified Point Cloud Segmentation

Zongyan Han, Mohamed El Amine Boudjoghra, Jiahua Dong, Jinhong Wang, Rao Muhammad Anwer

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫兹哈德大学人工智能大学) Technical University of Munich(慕尼黑技术大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17467 2025-07-24 cs.CV cs.AI 62%

Probing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls

Elena Pitta, Tom Kouwenhoven, Tessa Verhoef

机构 * Leiden Institute of Advanced Computer Science (LIACS), Leiden University, The Netherlands(莱顿先进计算机科学研究所(LIACS)、莱顿大学、荷兰)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments LUHME: 2nd Workshop on Language Understanding in the Human-Machine Era

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16877 2025-07-24 cs.CV cs.AI cs.CL 62%

ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension

Yizhi Hu, Zezhao Tian, Xingqun Qi, Chen Su, Bingkun Yang, Junhui Yin, Muyi Sun, Man Zhang, Zhenan Sun

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16213 2025-07-23 cs.CV cs.AI 62%

Advancing Visual Large Language Model for Multi-granular Versatile Perception

Wentao Xiang, Haoxian Tan, Cong Wei, Yujie Zhong, Dengjie Li, Yujiu Yang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Meituan Inc.(美团公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments To appear in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15094 2025-07-22 cs.CV cs.AI 62%

BleedOrigin: Dynamic Bleeding Source Localization in Endoscopic Submucosal Dissection via Dual-Stage Detection and Tracking

Mengya Xu, Rulin Zhou, An Wang, Chaoyang Lyu, Zhen Li, Ning Zhong, Hongliang Ren

机构 * Department of Electronic Engineering, The Chinese University of Hong Kong(电子工程系,香港中文大学) The Chinese University of Hong Kong Shenzhen Research Institute(香港中文大学深圳研究院) Department of Gastroenterology, Qilu Hospital of Shandong University(胃肠病科,山东大学齐鲁医院) Department of Mechanical Engineering, The University of Hong Kong(机械工程系,香港大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments 27 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏