arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7370 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7370 篇

2509.21356 2025-09-29 cs.CV cs.AI 73%

Phrase-grounded Fact-checking for Automatically Generated Chest X-ray Reports

Razi Mahmood, Diego Machado-Reyes, Joy Wu, Parisa Kaviani, Ken C. L. Wong, Niharika D'Souza, Mannudeep Kalra, Ge Wang, Pingkun Yan, Tanveer Syeda-Mahmood

机构 * Rensselaer Polytechnic Institute, NY, USA(罗文学院) IBM Research, Almaden, CA, USA(IBM研究院) Stanford University, CA, USA(斯坦福大学) Massachusetts General Hospital (MGH), Boston, USA(麻省总医院)

专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments In proceedings MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19875 2025-09-25 cs.CV cs.AI 73%

Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection

Yunqing Hu, Zheming Yang, Chang Zhao, Wen Ji

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of AI for Industries(工业人工智能研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13794 2025-09-22 cs.CV cs.AI 73%

LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation

Yang Zhou, Shiyu Zhao, Yuxiao Chen, Zhenting Wang, Can Jin, Dimitris N. Metaxas

机构 * Rutgers University(新泽西罗格斯大学)

专题命中 视觉定位与Grounding :grounding(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13642 2025-09-18 cs.LG cs.CV 73%

LLM-I: LLMs are Naturally Interleaved Multimodal Creators

Zirun Guo, Feng Zhang, Kai Jia, Tao Jin

机构 * Zhejiang University(浙江大学)

专题命中 视觉定位与Grounding :grounding(abstract);MLLM(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13234 2025-09-17 cs.AI cs.CV cs.HC 73%

Simulating Clinical AI Assistance using Multimodal LLMs: A Case Study in Diabetic Retinopathy

Nadim Barakat, William Lotter

机构 * Dana-Farber Cancer Institute & Tufts University School of Medicine(达纳-法伯癌症研究所及塔夫茨大学医学院) Dana-Farber Cancer Institute Brigham and Women’s Hospital & Harvard Medical School(达纳-法伯癌症研究所布里特妇女医院及哈佛医学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03961 2025-09-05 cs.CV cs.AI 73%

Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection

Yijun Zhou, Yikui Zhai, Zilu Ying, Tingfeng Xian, Wenlve Zhou, Zhiheng Zhou, Xiaolin Tian, Xudong Jia, Hongsheng Zhang, C. L. Philip Chen

机构 * College of Electronics and Information Engineering, Wuyi University(威怡大学电子与信息工程学院) School of Electronic and Information Engineering and the Key Laboratory of Big Data and Intelligent Robot, Ministry of Education, South China University of Technology(电子与信息工程学院和大数据与智能机器人重点实验室,华南理工大学) State Key Laboratory of Lunar and Planetary Sciences, Macau University of Science and Technology(澳门大学地球和行星科学国家重点实验室) College of Engineering and Computer Science, California State University, Northridge(工程与计算机科学学院,加州大学北岭分校) Department of Geography, The University of Hong Kong(地理系,香港大学) Faculty of Computer Science and Engineering, S(计算机科学与工程学院,S)

专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00284 2025-09-03 cs.CV cs.AI 73%

Generative AI for Industrial Contour Detection: A Language-Guided Vision System

Liang Gong, Tommy, Wang, Sara Chaker, Yanchen Dong, Fouad Bousetouane, Brenden Morton, Mark Mendez

机构 * The University of Chicago(芝加哥大学) FabTrack

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12263 2025-09-01 cs.CV cs.AI 73%

Region-Level Context-Aware Multimodal Understanding

Hongliang Wei, Xianqi Zhang, Xingtao Wang, Xiaopeng Fan, Debin Zhao

机构 * Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究院长) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18132 2025-08-26 cs.IR cs.AI cs.LG 73%

Test-Time Scaling Strategies for Generative Retrieval in Multimodal Conversational Recommendations

Hung-Chun Hsu, Yuan-Ching Kuo, Chao-Han Huck Yang, Szu-Wei Fu, Hanrong Ye, Hongxu Yin, Yu-Chiang Frank Wang, Ming-Feng Tsai, Chuan-Ju Wang

机构 * Research Center for Information Technology Innovation, Academia Sinica(资讯科技创新研究所以) NVIDIA(NVIDIA公司) Department of Computer Science, National Chengchi University(国立政治大学计算机科学系)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09333 2025-08-12 cs.CV cs.AI 73%

Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

Yufei Zhan, Shurong Zheng, Yousong Zhu, Hongyin Zhao, Fan Yang, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室) Wuhan AI Research, Wuhan, China(武汉人工智能研究所)

专题命中 视觉定位与Grounding :vision language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025. Codes and datasets are released at https://github.com/jefferyZhan/Griffon

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16623 2025-07-23 cs.CV cs.LG 73%

Automatic Fine-grained Segmentation-assisted Report Generation

Frederic Jonske, Constantin Seibold, Osman Alperen Koras, Fin Bahnsen, Marie Bauer, Amin Dada, Hamza Kalisch, Anton Schily, Jens Kleesiek

机构 * Institute for AI in Medicine, University Medicine Essen(人工智能医学研究所,埃森大学医学中心)

专题命中 视觉定位与Grounding :LLaVA(abstract);grounding(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00151 2025-07-11 cs.CV cs.AI 73%

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Ahmad Mohammadshirazi, Pinaki Prasad Guha Neogi, Ser-Nam Lim, Rajiv Ramnath

机构 * Department of Computer Science(计算机科学系) Engineering, Ohio State University, Ohio, US(工程系,俄亥俄州立大学,俄亥俄,美国) Department of Computer Science, University of Central Florida, Florida, US(计算机科学系,中央佛罗里达大学,佛罗里达,美国)

专题命中 视觉定位与Grounding :visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21892 2025-06-30 cs.CV cs.AI 73%

SODA: Out-of-Distribution Detection in Domain-Shifted Point Clouds via Neighborhood Propagation

Adam Goodge, Xun Xu, Bryan Hooi, Wee Siong Ng, Jingyi Liao, Yongyi Su, Xulei Yang

机构 * Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR), Singapore(信息通信研究所,科技研究局(A*STAR),新加坡) School of Computing, National University of Singapore, Singapore(计算学院,新加坡国立大学,新加坡)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06184 2025-06-30 cs.CV cs.AI cs.CE cs.HC cs.MA 73%

PEACE: Empowering Geologic Map Holistic Understanding with MLLMs

Yangyu Huang, Tianyi Gao, Haoran Xu, Qihao Zhao, Yang Song, Zhipeng Gui, Tengchao Lv, Hao Chen, Lei Cui, Scarlett Li, Furu Wei

机构 * Microsoft Research(微软研究院) Chinese Academy of Geological Sciences(中国地质科学研究院) Wuhan University(武汉大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09174 2025-06-24 cs.CV cs.AI 73%

DART: An Automated End-to-End Object Detection Pipeline with Data Diversification, Open-Vocabulary Bounding Box Annotation, Pseudo-Label Review, and Model Training

Chen Xin, Andreas Hartel, Enkelejda Kasneci

专题命中 视觉定位与Grounding :InternVL(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments Corrected minor typos; no changes to results or conclusions

Journal ref Expert Systems with Applications 258 (2024): 125124

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07848 2025-06-10 cs.CV cs.AI 73%

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

Teng Hu, Zhentao Yu, Zhengguang Zhou, Jiangning Zhang, Yuan Zhou, Qinglin Lu, Ran Yi

机构 * Shanghai Jiao Tong University(上海交通大学) Tencent Hunyuan(腾讯文英) Zhejiang University(浙江大学)

专题命中 视觉定位与Grounding :grounding(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16149 2025-05-23 cs.CV cs.AI cs.CL 73%

When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification

Zirui Pang, Haosheng Tan, Yuhan Pu, Zhijie Deng, Zhouan Shen, Keyu Hu, Jiaheng Wei

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Glasgow(格拉斯哥大学) Boston University(波士顿大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 视觉定位与Grounding :vision-language model(abstract);LLaVA(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17045 2025-05-23 cs.CV cs.AI 73%

GeoBiked: A Dataset with Geometric Features and Automated Labeling Techniques to Enable Deep Generative Models in Engineering Design

Phillip Mueller, Sebastian Mueller, Lars Mikelsons

机构 * BMW Group(宝马集团) University of Augsburg(奥格斯堡大学)

专题命中 视觉定位与Grounding :VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13399 2025-04-21 cs.CV cs.AI 73%

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety

Shashank Shriram, Srinivasa Perisetla, Aryan Keskar, Harsha Krishnaswamy, Tonko Emil Westerhof Bossen, Andreas Møgelmose, Ross Greer

机构 * Machine Intelligence, Interaction, and Imagination (Mi 3 ) Laboratory(机器智能、交互与想象(Mi 3)实验室) University of California, Merced(加州大学默塞德分校) Aalborg Universitet(奥胡斯大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.17422 2025-04-08 cs.CV cs.AI cs.RO 73%

Open-Vocabulary Action Localization with Iterative Visual Prompting

Naoki Wake, Atsushi Kanehira, Kazuhiro Sasabuchi, Jun Takamatsu, Katsushi Ikeuchi

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 5 figures, 6 tables. Published in IEEE Access. Last updated on April 7th, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00954 2025-04-02 cs.CV cs.AI 73%

IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval

Bangwei Liu, Yicheng Bao, Shaohui Lin, Xuhong Wang, Xin Tan, Yingchun Wang, Yuan Xie, Chaochao Lu

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14908 2025-03-20 cs.GR cs.AI cs.CV 73%

POSTA: A Go-to Framework for Customized Artistic Poster Generation

Haoyu Chen, Xiaojie Xu, Wenbo Li, Jingjing Ren, Tian Ye, Songhua Liu, Ying-Cong Chen, Lei Zhu, Xinchao Wang

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08549 2025-03-03 cs.CV cs.AI 73%

The Devil is in Temporal Token: High Quality Video Reasoning Segmentation

Sitong Gong, Yunzhi Zhuge, Lu Zhang, Zongxin Yang, Pingping Zhang, Huchuan Lu

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Journal ref CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20277 2025-02-28 cs.CV cs.AI 73%

Explainable, Multi-modal Wound Infection Classification from Images Augmented with Generated Captions

Palawat Busaranuvong, Emmanuel Agu, Reza Saadati Fard, Deepak Kumar, Shefalika Gautam, Bengisu Tulu, Diane Strong

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02746 2025-02-20 cs.CV cs.LG 73%

Contrastive Localized Language-Image Pre-Training

Hong-You Chen, Zhengfeng Lai, Haotian Zhang, Xinze Wang, Marcin Eichner, Keen You, Meng Cao, Bowen Zhang, Yinfei Yang, Zhe Gan

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.LG

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07905 2025-02-13 cs.CV cs.LG 73%

DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities

Chashi Mahiul Islam, Samuel Jacob Chacko, Preston Horne, Xiuwen Liu

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.LG

Comments 19 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15346 2025-02-11 cs.CV cs.AI 73%

YOLO-RD: Introducing Relevant and Compact Explicit Knowledge to YOLO by Retriever-Dictionary

Hao-Tang Tsui, Chien-Yao Wang, Hong-Yuan Mark Liao

专题命中 视觉定位与Grounding :VLM(abstract);visual language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14680 2024-12-30 cs.CV cs.AI cs.RO 73%

A Light-Weight Framework for Open-Set Object Detection with Decoupled Feature Alignment in Joint Space

Yonghao He, Hu Su, Haiyong Yu, Cong Yang, Wei Sui, Cong Wang, Song Liu

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17981 2024-12-20 cs.CV cs.AI 73%

From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information

Qirui Jiao, Daoyuan Chen, Yilun Huang, Yaliang Li, Ying Shen

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 32 pages, 22 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02262 2024-12-04 cs.CV cs.LG 73%

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation

Sepand Dyanatkar, Angran Li, Alexander Dungate

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.LG

Comments Accepted to Climate Change AI Workshop at NeurIPS 2024. 9 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏