arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7387 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7387 篇

2508.12605 2025-08-19 cs.CV 70%

ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images

Wenjie Liao, Jieyu Yuan, Yifang Xu, Chunle Guo, Zilong Zhang, Jihong Li, Jiachen Fu, Haotian Fan, Tao Li, Junhui Cui, Chongyi Li

机构 * Media Evaluation Lab, ByteDance Inc.(字节跳动公司媒体评估实验室)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12349 2025-08-19 cs.CV 70%

EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos

Junyi Ma, Erhang Zhang, Yin-Dong Zheng, Yuchen Xie, Yixuan Zhou, Hesheng Wang

机构 * IRMV Lab, the Department of Automation, Shanghai Jiao Tong University(IRMV实验室,自动化系,上海交通大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments Extended journal version of arXiv:2506.03662

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09999 2025-08-15 cs.CL cs.LG 70%

XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs

Yuzhuo Xiao, Zeyu Han, Yuhan Wang, Huaizu Jiang

机构 * Guizhou University(贵州大学) Northeastern University(东北大学) UC Santa Cruz(加州大学圣克ruz分校)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.LG

Comments For associated code and dataset, see https://github.com/neu-vi/XFacta

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19694 2025-08-12 cs.CV 70%

UltraAD: Fine-Grained Ultrasound Anomaly Classification via Few-Shot CLIP Adaptation

Yue Zhou, Yuan Bi, Wenjuan Tong, Wei Wang, Nassir Navab, Zhongliang Jiang

机构 * Computer Aided Medical Procedures (CAMP)(计算机辅助医疗程序) TU Munich, Germany(慕尼黑工业大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) The First Affiliated Hospital of Sun Yat-Sen University(中山大学附属第一医院)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01663 2025-08-12 cs.CV 70%

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement

Xuan Yu, Dayan Guan, Yanfeng Gu

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Code is available at https://github.com/xavier-yu114/Zoom-Refine

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05323 2025-08-08 cs.CV 70%

Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting

Frank Ruis, Gertjan Burghouts, Hugo Kuijf

机构 * TNO(荷兰技术院) Intelligent Imaging(智能成像)

专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03967 2025-08-07 cs.CV cs.CR cs.IR 70%

RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification

Mamadou Keita, Wassim Hamidouche, Hessen Bougueffa Eutamene, Abdelmalik Taleb-Ahmed, Abdenour Hadid

机构 * Laboratory of IEMN, Univ. Polytechnique Hauts-de-France(IEMN实验室,法国高等技术法国大学) KU 6G Research Center, Khalifa University(KU 6G研究中心,哈利法大学) Sorbonne Center for Artificial Intelligence, Sorbonne University(人工智能研究中心,索邦大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02951 2025-08-06 cs.AI 70%

MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine

Mahtab Bigverdi, Wisdom Ikezogwo, Kevin Zhang, Hyewon Jeong, Mingyu Lu, Sungjae Cho, Linda Shapiro, Ranjay Krishna

机构 * University of Washington(华盛顿大学) Massachusetts Institute of Technology(麻省理工学院) Seoul National University(首尔国立大学)

专题命中 视觉定位与Grounding :LLaVA(abstract);grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18937 2025-08-05 cs.CV cs.CL 70%

Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description

Mahmoud Ahmed, Junjie Fei, Jian Ding, Eslam Mohamed Bakr, Mohamed Elhoseiny

机构 * King Abdullah University of Science and Technology(卡斯特大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21619 2025-07-30 cs.CV 70%

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO

Wei Guan, Jun Lan, Jian Cao, Hao Tan, Huijia Zhu, Weiqiang Wang

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20025 2025-07-29 cs.CV 70%

Region-based Cluster Discrimination for Visual Representation Learning

Yin Xie, Kaicheng Yang, Xiang An, Kun Wu, Yongle Zhao, Weimo Deng, Zimin Ran, Yumeng Wang, Ziyong Feng, Roy Miles, Ismail Elezi, Jiankang Deng

机构 * DeepGlint University of Technology Sydney(悉尼科技大学) Huawei London Research Center(华为伦敦研究中心) Imperial College London(伦敦帝国理工学院)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Accepted as a highlight paper at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10568 2025-07-28 cs.CV 70%

AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark

Aruna Gauba, Irene Pi, Yunze Man, Ziqi Pang, Vikram S. Adve, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Rice University(Rice大学) Carnegie Mellon University(卡内基梅隆大学) AIFARMS Center for Digital Agriculture at UIUC(伊利诺伊大学厄巴纳-香槟分校数字农业中心)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments Project Website: https://agmmu.github.io/ Huggingface: https://huggingface.co/datasets/AgMMU/AgMMU_v1/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07783 2025-07-28 cs.CV cs.CL 70%

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

Zhaokai Wang, Xizhou Zhu, Xue Yang, Gen Luo, Hao Li, Changyao Tian, Wenhan Dou, Junqi Ge, Lewei Lu, Yu Qiao, Jifeng Dai

机构 * Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Sensetime

专题命中 视觉定位与Grounding :LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV

Journal ref TPAMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18300 2025-07-25 cs.CV 70%

LMM-Det: Make Large Multimodal Models Excel in Object Detection

Jincheng Li, Chunyu Xie, Ji Ao, Dawei Leng, Yuhui Yin

机构 * AI Research(360人工智能研究院)

专题命中 视觉定位与Grounding :visual question answering(abstract);grounding(abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17436 2025-07-24 cs.CV 70%

Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection

Yehao Lu, Minghe Weng, Zekang Xiao, Rui Jiang, Wei Su, Guangcong Zheng, Ping Lu, Xi Li

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Polytechnic Institute, Zhejiang University(浙江大学 polytechnic 院) ZTE(ZTE 公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14449 2025-07-22 cs.CV 70%

IRGPT: Understanding Real-world Infrared Image with Bi-cross-modal Curriculum on Large-scale Benchmark

Zhe Cao, Jin Zhang, Ruiheng Zhang

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments 11 pages, 7 figures. This paper is accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12887 2025-07-18 eess.IV cs.CV 70%

RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions

Junzhi Ning, Cheng Tang, Kaijing Zhou, Diping Song, Lihao Liu, Ming Hu, Wei Li, Huihui Xu, Yanzhou Su, Tianbin Li, Jiyao Liu, Jin Ye, Sheng Zhang, Yuanfeng Ji, Junjun He

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) Shanghai Institute of Laser Technology(上海激光技术研究所) Eye Hospital, Wenzhou Medical University(温州医学院眼医院) Monash University(墨尔本大学) Shanghai Jiao Tong University(上海交通大学) Fuzhou University(福州大学) Fudan University(复旦大学) Imperial College London(伦敦帝国理工学院) Stanford University(斯坦福大学)

专题命中 视觉定位与Grounding :VLM(abstract);visual language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12386 2025-07-15 cs.CV 70%

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Yi Wang, Xinhao Li, Ziang Yan, Yinan He, Jiashuo Yu, Xiangyu Zeng, Chenting Wang, Changlian Ma, Haian Huang, Jianfei Gao, Min Dou, Kai Chen, Wenhai Wang, Yu Qiao, Yali Wang, Limin Wang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Nanjing University(南京大学) Shanghai Innovation Institute(上海创新研究院) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07340 2025-07-14 cs.CV 70%

Entity Re-identification in Visual Storytelling via Contrastive Reinforcement Learning

Daniel A. P. Oliveira, David Martins de Matos

机构 * INESC-ID(INESC-ID研究所) Instituto Superior Técnico, Universidade de Lisboa(里斯本大学技术学院,里斯本大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02900 2025-07-11 cs.CV 70%

MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine

Yunfei Xie, Ce Zhou, Lang Gao, Juncheng Wu, Xianhang Li, Hong-Yu Zhou, Sheng Liu, Lei Xing, James Zou, Cihang Xie, Yuyin Zhou

机构 * Huazhong University of Science and Technology(华中科技大学) UC Santa Cruz(加州大学圣克ruz分校) Harvard University(哈佛大学) Stanford University(斯坦福大学)

专题命中 视觉定位与Grounding :LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV

Comments The dataset is publicly available at https://yunfeixie233.github.io/MedTrinity-25M/. Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04152 2025-07-08 cs.CV 70%

LVLM-Composer's Explicit Planning for Image Generation

Spencer Ramsey, Jeffrey Lee, Amina Grant

机构 * Northern Caribbean University(北加勒比大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04151 2025-07-08 cs.CV 70%

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation

Fernando Gabriela Garcia, Spencer Burns, Ryan Shaw, Hunter Young

机构 * Autonomous University of Nuevo León(新莱昂自治大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23309 2025-07-02 eess.IV cs.CV 70%

SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting

Yiming Huang, Long Bai, Beilei Cui, Kun Yuan, Guankun Wang, Mobarak I. Hoque, Nicolas Padoy, Nassir Navab, Hongliang Ren

机构 * The Chinese University of Hong Kong, Hong Kong SAR, China(香港中文大学) Shenzhen Research Institute, CUHK, Shenzhen, China(深圳研究学院) Technical University of Munich, Munich, Germany(慕尼黑技术大学) University of Strasbourg & IHU Strasbourg, Strasbourg, France(斯特拉斯堡大学及斯特拉斯堡IHU) University College London, London, United Kingdom(伦敦大学学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments MICCAI 2025. Project Page: MICCAI-2025-SurgTPGS/" target="_blank" rel="noopener">https://lastbasket.github.io/MICCAI-2025-SurgTPGS/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15180 2025-06-19 cs.CV 70%

ReSeDis: A Dataset for Referring-based Object Search across Large-Scale Image Collections

Ziling Huang, Yidan Zhang, Shin'ichi Satoh

机构 * National Institute of Informatics(日本国立信息学研究所) The University of Tokyo(东京大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14471 2025-06-18 cs.CV 70%

Dense360: Dense Understanding from Omnidirectional Panoramas

Yikang Zhou, Tao Zhang, Dizhe Zhang, Shunping Ji, Xiangtai Li, Lu Qi

机构 * Wuhan University(武汉大学) Insta360 Peking University(北京大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10756 2025-06-13 cs.RO cs.AI 70%

Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding

Yuhang Zhang, Haosheng Yu, Jiaping Xiao, Mir Feroskhan

机构 * School of Mechanical and Aerospace Engineering, Nanyang Technological University(机械与航空航天工程学院,南洋理工大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10474 2025-06-13 cs.CV 70%

LLMs Are Not Yet Ready for Deepfake Image Detection

Shahroz Tariq, David Nguyen, M. A. P. Chamikara, Tingmin Wu, Alsharif Abuadbba, Kristen Moore

机构 * CSIRO’s Data61(澳大利亚CSIRO数据61研究所)

专题命中 视觉定位与Grounding :vision-language model(abstract);visual reasoning(abstract);分类 cs.CV

Comments 6 pages, 3 figures, and 2 tables. paper is under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13856 2025-06-13 cs.RO cs.CV 70%

Simultaneous Localization and Affordance Prediction of Tasks from Egocentric Video

Zachary Chavis, Hyun Soo Park, Stephen J. Guy

机构 * Department of Computer Science and Engineering (CS&E), University of Minnesota(计算机科学与工程系(CS&E),明尼苏达大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07850 2025-06-10 cs.CV 70%

SAM2Auto: Auto Annotation Using FLASH

Arash Rocky, Q. M. Jonathan Wu

机构 * Department of Electrical and Computer Engineering, University of Windsor(电气与计算机工程系,温莎大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01908 2025-06-03 cs.CV 70%

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Hongyu Li, Songhao Han, Yue Liao, Junfeng Luo, Jialin Gao, Shuicheng Yan, Si Liu

机构 * BUAA(北京航空航天大学) NUS(国立大学新加坡) Meituan(美团)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏