arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7387 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7387 篇

2501.16981 2025-03-07 cs.CV 70%

Modulating CNN Features with Pre-Trained ViT Representations for Open-Vocabulary Object Detection

Xiangyu Gao, Yu Dai, Benliu Qiu, Lanxiao Wang, Heqian Qiu, Hongliang Li

专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01794 2025-03-04 cs.CV 70%

OFF-CLIP: Improving Normal Detection Confidence in Radiology CLIP with Simple Off-Diagonal Term Auto-Adjustment

Junhyun Park, Chanyu Moon, Donghwan Lee, Kyungsu Kim, Minho Hwang

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments 10 pages, 3 figures, and 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01785 2025-03-04 cs.CV 70%

Visual-RFT: Visual Reinforcement Fine-Tuning

Ziyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, Jiaqi Wang

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments project page: https://github.com/Liuziyu77/Visual-RFT

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20698 2025-03-03 cs.CV 70%

Towards General Visual-Linguistic Face Forgery Detection(V2)

Ke Sun, Shen Chen, Taiping Yao, Ziyin Zhou, Jiayi Ji, Xiaoshuai Sun, Chia-Wen Lin, Rongrong Ji

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments 8 pages, 5 figures, Accpet by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.04796 2025-02-27 cs.CV 70%

Local-Prompt: Extensible Local Prompts for Few-Shot Out-of-Distribution Detection

Fanhu Zeng, Zhen Cheng, Fei Zhu, Hongxin Wei, Xu-Yao Zhang

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments Accepted by The Thirteenth International Conference on Learning Representations (ICLR 2025). Code is available at https://github.com/AuroraZengfh/Local-Prompt

Journal ref The Thirteenth International Conference on Learning Representations (ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09855 2025-02-18 cs.CV 70%

Text4Seg: Reimagining Image Segmentation as Text Generation

Mengcheng Lan, Chaofeng Chen, Yue Zhou, Jiaxing Xu, Yiping Ke, Xinjiang Wang, Litong Feng, Wayne Zhang

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments ICLR 2025. Project page: https://mc-lan.github.io/Text4Seg/

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15795 2025-02-17 cs.CV 70%

AdaCLIP: Adapting CLIP with Hybrid Learnable Prompts for Zero-Shot Anomaly Detection

Yunkang Cao, Jiangning Zhang, Luca Frittoli, Yuqi Cheng, Weiming Shen, Giacomo Boracchi

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments Accepted by ECCV 2024

Journal ref European Conference on Computer Vision, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10417 2025-01-27 cs.AI 70%

M^2ConceptBase: A Fine-Grained Aligned Concept-Centric Multimodal Knowledge Base

Zhiwei Zha, Jiaan Wang, Zhixu Li, Xiangru Zhu, Wei Song, Yanghua Xiao

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.AI

Comments Accepted by CIKM2024. The code and data can be found at https://github.com/AwellmanZha/M2ConceptBase

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09218 2025-01-17 q-bio.QM cs.AI 70%

Interpretable Droplet Digital PCR Assay for Trustworthy Molecular Diagnostics

Yuanyuan Wei, Yucheng Wu, Fuyang Qu, Yao Mu, Yi-Ping Ho, Ho-Pui Ho, Wu Yuan, Mingkun Xu

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14505 2025-01-16 cs.CV 70%

T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

Kaiyue Sun, Kaiyi Huang, Xian Liu, Yue Wu, Zihan Xu, Zhenguo Li, Xihui Liu

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Project page: https://t2v-compbench-2025.github.io/ Code: https://github.com/KaiyueSun98/T2V-CompBench/tree/V2

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07554 2025-01-14 cs.CV cs.CL 70%

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing

Varun Biyyala, Bharat Chanderprakash Kathuria, Jialu Li, Youshan Zhang

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments WACV workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14750 2025-01-14 cs.CV cs.CL 70%

FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension

Junzhuo Liu, Xuzheng Yang, Weiwei Li, Peng Wang

专题命中 视觉定位与Grounding :visual reasoning(abstract);grounding(abstract);分类 cs.CV

Comments 18 pages, EMNLP 2024 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03786 2025-01-08 cs.CV 70%

KAnoCLIP: Zero-Shot Anomaly Detection through Knowledge-Driven Prompt Learning and Enhanced Cross-Modal Integration

Chengyuan Li, Suyang Zhou, Jieping Kong, Lei Qi, Hui Xue

专题命中 视觉定位与Grounding :vision-language model(abstract);visual question answering(abstract);分类 cs.CV

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16869 2024-12-24 cs.CV 70%

CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models

Yeyuan Wang, Dehong Gao, Bin Li, Rujiao Long, Lei Yi, Xiaoyan Cai, Libin Yang, Jinxia Zhang, Shanqing Yu, Qi Xuan

专题命中 视觉定位与Grounding :grounding(abstract);MLLM(abstract);分类 cs.CV

Comments 5 pages, Accepted by ICASSP2025, full paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02291 2024-12-24 cs.CV 70%

Focusing on what to decode and what to train: SOV Decoding with Specific Target Guided DeNoising and Vision Language Advisor

Junwen Chen, Yingcheng Wang, Keiji Yanai

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14145 2024-12-19 cs.CV 70%

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation

Jianyu Zhang, Li Zhang, Shijian Li

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13947 2024-12-19 cs.CV 70%

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition

Ethan Baron, Idan Tankel, Peter Tu, Guy Ben-Yosef

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04788 2024-12-03 cs.CV 70%

SemiCD-VL: Visual-Language Model Guidance Makes Better Semi-supervised Change Detector

Kaiyu Li, Xiangyong Cao, Yupeng Deng, Jiayi Song, Junmin Liu, Deyu Meng, Zhi Wang

专题命中 视觉定位与Grounding :VLM(abstract);visual language model(abstract);分类 cs.CV

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00307 2024-11-04 cs.CV 70%

HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model

Khoa Vo, Thinh Phan, Kashu Yamazaki, Minh Tran, Ngan Le

专题命中 视觉定位与Grounding :visual reasoning(abstract);grounding(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20141 2024-10-31 cs.CV 70%

OpenDAS: Open-Vocabulary Domain Adaptation for 2D and 3D Segmentation

Gonca Yilmaz, Songyou Peng, Marc Pollefeys, Francis Engelmann, Hermann Blum

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16163 2024-10-22 cs.CV 70%

Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models

Yufei Zhan, Hongyin Zhao, Yousong Zhu, Fan Yang, Ming Tang, Jinqiao Wang

专题命中 视觉定位与Grounding :visual question answering(abstract);grounding(abstract);分类 cs.CV

Comments This work has been submitted to the IEEE for possible publication. Codes and data will be later released at https://github.com/jefferyZhan/Griffon

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12015 2024-10-11 cs.RO cs.CL cs.CV 70%

GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration

Naoki Wake, Atsushi Kanehira, Kazuhiro Sasabuchi, Jun Takamatsu, Katsushi Ikeuchi

专题命中 视觉定位与Grounding :vision language model(abstract);grounding(abstract);分类 cs.CV

Comments 8 pages, 10 figures, 3 tables. Published in IEEE Robotics and Automation Letters (RA-L) (in press). Last updated on September 26th, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05963 2024-10-10 cs.CV 70%

Training-Free Open-Ended Object Detection and Segmentation via Attention as Prompts

Zhiwei Lin, Yongtao Wang, Zhi Tang

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00436 2024-10-02 cs.RO cs.CV 70%

Task Success Prediction for Open-Vocabulary Manipulation Based on Multi-Level Aligned Representations

Miyu Goko, Motonari Kambara, Daichi Saito, Seitaro Otsuki, Komei Sugiura

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted for presentation at CoRL2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.20171 2024-08-27 cs.CV 70%

Diffusion Feedback Helps CLIP See Better

Wenxuan Wang, Quan Sun, Fan Zhang, Yepeng Tang, Jing Liu, Xinlong Wang

专题命中 视觉定位与Grounding :VLM(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02261 2024-08-06 cs.CV 70%

Cross-Domain Semantic Segmentation on Inconsistent Taxonomy using VLMs

Jeongkee Lim, Yusung Kim

专题命中 视觉定位与Grounding :vision language model(abstract);visual language model(abstract);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01433 2024-08-06 cs.CV cs.ET 70%

Evaluating and Enhancing Trustworthiness of LLMs in Perception Tasks

Malsha Ashani Mahawatta Dona, Beatriz Cabrero-Daniel, Yinan Yu, Christian Berger

专题命中 视觉定位与Grounding :LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Accepted in 27th IEEE International Conference on Intelligent Transportation Systems (ITSC) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00331 2024-08-02 cs.CV 70%

DECIDER: Leveraging Foundation Model Priors for Improved Model Failure Detection and Explanation

Rakshith Subramanyam, Kowshik Thopalli, Vivek Narayanaswamy, Jayaraman J. Thiagarajan

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments Accepted at ECCV (European Conference on Computer Vision) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21554 2024-08-01 cs.CV 70%

Conditioned Prompt-Optimization for Continual Deepfake Detection

Francesco Laiti, Benedetta Liberatori, Thomas De Min, Elisa Ricci

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments Accepted at ICPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14815 2024-07-23 cs.CV 70%

Designing A Sustainable Marine Debris Clean-up Framework without Human Labels

Raymond Wang, Nicholas R. Record, D. Whitney King, Tahiya Chowdhury

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments 9 pages, 6 figures, 2 tables, In Proceedings of the 7th ACM SIGCAS/SIGCHI Conference on Computing and Sustainable Societies

详情

展开后加载摘要…

URL PDF HTML 收藏