Modulating CNN Features with Pre-Trained ViT Representations for Open-Vocabulary Object Detection
专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV
Comments 10 pages, 3 figures, and 5 tables
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV
Comments project page: https://github.com/Liuziyu77/Visual-RFT
专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments 8 pages, 5 figures, Accpet by CVPR2025
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments Accepted by The Thirteenth International Conference on Learning Representations (ICLR 2025). Code is available at https://github.com/AuroraZengfh/Local-Prompt
Journal ref The Thirteenth International Conference on Learning Representations (ICLR 2025)
专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments ICLR 2025. Project page: https://mc-lan.github.io/Text4Seg/
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments Accepted by ECCV 2024
Journal ref European Conference on Computer Vision, 2024
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.AI
Comments Accepted by CIKM2024. The code and data can be found at https://github.com/AwellmanZha/M2ConceptBase
专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI
专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments Project page: https://t2v-compbench-2025.github.io/ Code: https://github.com/KaiyueSun98/T2V-CompBench/tree/V2
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments WACV workshop
专题命中 视觉定位与Grounding :visual reasoning(abstract);grounding(abstract);分类 cs.CV
Comments 18 pages, EMNLP 2024 main
专题命中 视觉定位与Grounding :vision-language model(abstract);visual question answering(abstract);分类 cs.CV
Comments Accepted by ICASSP 2025
专题命中 视觉定位与Grounding :grounding(abstract);MLLM(abstract);分类 cs.CV
Comments 5 pages, Accepted by ICASSP2025, full paper
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments 6 pages, 6 figures
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :VLM(abstract);visual language model(abstract);分类 cs.CV
Comments 13 pages, 6 figures
专题命中 视觉定位与Grounding :visual reasoning(abstract);grounding(abstract);分类 cs.CV
Comments Accepted to NeurIPS 2024
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :visual question answering(abstract);grounding(abstract);分类 cs.CV
Comments This work has been submitted to the IEEE for possible publication. Codes and data will be later released at https://github.com/jefferyZhan/Griffon
专题命中 视觉定位与Grounding :vision language model(abstract);grounding(abstract);分类 cs.CV
Comments 8 pages, 10 figures, 3 tables. Published in IEEE Robotics and Automation Letters (RA-L) (in press). Last updated on September 26th, 2024
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments Accepted by NeurIPS 2024
专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted for presentation at CoRL2024
专题命中 视觉定位与Grounding :VLM(abstract);multimodal large language model(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :vision language model(abstract);visual language model(abstract);分类 cs.CV
Comments ECCV 2024
专题命中 视觉定位与Grounding :LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV
Comments Accepted in 27th IEEE International Conference on Intelligent Transportation Systems (ITSC) 2024
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments Accepted at ECCV (European Conference on Computer Vision) 2024
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments Accepted at ICPR 2024
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV
Comments 9 pages, 6 figures, 2 tables, In Proceedings of the 7th ACM SIGCAS/SIGCHI Conference on Computing and Sustainable Societies