arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7387 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7387 篇

2507.15003 2025-07-22 cs.SE cs.AI cs.CE cs.LG 62%

The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering

Hao Li, Haoxiang Zhang, Ahmed E. Hassan

机构 * Queen's University(女王大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09795 2025-07-22 cs.CV cs.LG 62%

NegRefine: Refining Negative Label-Based Zero-Shot OOD Detection

Amirhossein Ansari, Ke Wang, Pulei Xiong

机构 * Simon Fraser University(西蒙弗雷泽大学) National Research Council Canada(加拿大国家研究理事会)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13363 2025-07-21 cs.CV cs.AI 62%

Just Add Geometry: Gradient-Free Open-Vocabulary 3D Detection Without Human-in-the-Loop

Atharv Goel, Mehar Khurana

机构 * Indraprastha Institute of Information Technology(印度理工学院信息技术研究所)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12761 2025-07-18 cs.CV cs.AI 62%

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation

Hanlei Shi, Leyuan Qu, Yu Liu, Di Gao, Yuhua Zheng, Taihao Li

机构 * Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(杭州高等研究院,中国科学院大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01504 2025-07-16 cs.CV cs.AI cs.CL 62%

Following the Clues: Experiments on Person Re-ID using Cross-Modal Intelligence

Robert Aufschläger, Youssef Shoeb, Azarm Nowzad, Michael Heigl, Fabian Bally, Martin Schramm

机构 * Deggendorf Institute of Technology(德格多夫技术学院) Continental AG(大陆集团)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments accepted for publication at the 2025 IEEE 28th International Conference on Intelligent Transportation Systems (ITSC 2025), taking place during November 18-21, 2025 in Gold Coast, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.02656 2025-07-16 cs.LG cs.AI cs.NE 62%

Invariant Representations with Stochastically Quantized Neural Networks

Mattia Cerrato, Marius Köppel, Roberto Esposito, Stefan Kramer

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments To appear in AAAI23

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08740 2025-07-10 cs.CV cs.AI cs.IR 62%

Hespi: A pipeline for automatically detecting information from hebarium specimen sheets

Robert Turnbull, Emily Fitzgerald, Karen Thompson, Joanne L. Birch

机构 * The University of Melbourne(墨尔本大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18116 2025-07-09 cs.CV cs.AI cs.CL 62%

Bayesian Optimization for Controlled Image Editing via LLMs

Chengkun Cai, Haoliang Liu, Xu Zhao, Zhongyu Jiang, Tianfang Zhang, Zongkai Wu, John Lee, Jenq-Neng Hwang, Lei Li

机构 * University of Edinburgh(爱丁堡大学) University of Manchester(曼彻斯特大学) University of Washington(华盛顿大学) Tsinghua University(清华大学) Skai Intelligence(Skai智能科技) University of Copenhagen(哥本哈根大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments 8 figures, accept at ACL2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04522 2025-07-08 cs.CV cs.AI cs.RO 62%

Grounded Gesture Generation: Language, Motion, and Space

Anna Deichler, Jim O'Regan, Teo Guichoux, David Johansson, Jonas Beskow

机构 * KTH Royal Institute of Technology(皇家理工学院) Sorbonne University(索邦大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted as a non-archival paper at the CVPR 2025 Humanoid Agents Workshop. Project page: https://groundedgestures.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07772 2025-07-08 cs.CV cs.LG 62%

Hallucinatory Image Tokens: A Training-free EAZY Approach on Detecting and Mitigating Object Hallucinations in LVLMs

Liwei Che, Tony Qingze Liu, Jing Jia, Weiyi Qin, Ruixiang Tang, Vladimir Pavlovic

机构 * Rutgers University(罗杰斯大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

Comments Accepted to ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02920 2025-07-08 cs.HC cs.AI cs.LG 62%

Visual-Conversational Interface for Evidence-Based Explanation of Diabetes Risk Prediction

Reza Samimi, Aditya Bhattacharya, Lucija Gosak, Gregor Stiglic, Katrien Verbert

机构 * University of Maribor(马拉堡大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments 18 pages, 5 figures, 7th ACM Conference on Conversational User Interfaces

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02906 2025-07-08 cs.CV cs.LG 62%

Enhancing Sports Strategy with Video Analytics and Data Mining: Automated Video-Based Analytics Framework for Tennis Doubles

Jia Wei Chen

机构 * Department of Information Systems and Analytics(信息系统与分析系) School of Computing(计算学院) National University of Singapore(新加坡国立大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

Comments B.Sc. thesis 59 pages, 26 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00018 2025-07-08 cs.LG cs.AI cs.CL 62%

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections

Bo Wang, Qinyuan Cheng, Runyu Peng, Rong Bao, Peiji Li, Qipeng Guo, Linyang Li, Zhiyuan Zeng, Yunhua Zhou, Xipeng Qiu

机构 * School of Computer Science, Fudan University(复旦大学计算机学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09320 2025-07-03 cs.CV cs.LG cs.RO 62%

2HandedAfforder: Learning Precise Actionable Bimanual Affordances from Human Videos

Marvin Heidinger, Snehal Jauhri, Vignesh Prasad, Georgia Chalvatzaki

机构 * Computer Science Department, Technische Universität Darmstadt(德意志联邦共和国达姆施塔特技术大学计算机科学系)

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV、cs.LG

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20960 2025-07-01 cs.CV cs.AI 62%

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

Yiman Zhang, Ziheng Luo, Qiangyu Yan, Wei He, Borui Jiang, Xinghao Chen, Kai Han

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室) University of Science and Technology of China(中国科学技术大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16994 2025-06-23 cs.CV cs.LG 62%

Prmpt2Adpt: Prompt-Based Zero-Shot Domain Adaptation for Resource-Constrained Environments

Yasir Ali Farrukh, Syed Wali, Irfan Khan, Nathaniel D. Bastian

机构 * Clean and Resilient Energy System Lab (CARES), Texas A&M University(清洁而坚韧能源系统实验室(CARES),德克萨斯A&M大学) Robotics Research Center, United States Military Academy(机器人研究中心,美国军事学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11261 2025-06-16 cs.RO cs.AI cs.CV 62%

Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation

Shizhe Chen, Ricardo Garcia, Paul Pacaud, Cordelia Schmid

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10503 2025-06-13 cs.CV cs.AI 62%

Semantic Localization Guiding Segment Anything Model For Reference Remote Sensing Image Segmentation

Shuyang Li, Shuang Wang, Zhuangzhuang Sun, Jing Xiao

机构 * School of Artificial Intelligence, Xidian University(西安电子科技大学人工智能学院) Shaanxi Satellite Application Center for Natural Resources(陕西省自然资源卫星应用中心)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10745 2025-06-10 cs.CV cs.AI cs.RO 62%

Unifying 2D and 3D Vision-Language Understanding

Ayush Jain, Alexander Swerdlow, Yuzhou Wang, Sergio Arnaud, Ada Martin, Alexander Sax, Franziska Meier, Katerina Fragkiadaki

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09401 2025-06-10 cs.CV cs.AI cs.RO 62%

MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations

Ruiyuan Lyu, Jingli Lin, Tai Wang, Shuai Yang, Xiaohan Mao, Yilun Chen, Runsen Xu, Haifeng Huang, Chenming Zhu, Dahua Lin, Jiangmiao Pang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Zhiyuan College, Shanghai Jiao Tong University(上海交通大学紫阳学院) CPII under InnoHK(创新香港下的CPII)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Follow-up of EmbodiedScan (camera-ready version). A multi-modal 3D dataset with the most-ever comprehensive language annotations for 3D-LLMs. Project page: https://tai-wang.github.io/mmscan/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06569 2025-06-10 cs.CV cs.AI 62%

Textile Analysis for Recycling Automation using Transfer Learning and Zero-Shot Foundation Models

Yannis Spyridis, Vasileios Argyriou

机构 * Department of Computer Science(计算机科学系) Department of Networks and Digital Media(网络与数字媒体系)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Journal ref IEEE DCOSS IoTi5 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05358 2025-06-09 cs.CV cs.AI cs.CR 62%

Can ChatGPT Perform Image Splicing Detection? A Preliminary Study

Souradip Nath

机构 * Arizona State University(亚利桑那州立大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18941 2025-06-06 cs.CV cs.LG 62%

LEMoN: Label Error Detection using Multimodal Neighbors

Haoran Zhang, Aparna Balagopalan, Nassim Oufattole, Hyewon Jeong, Yan Wu, Jiacheng Zhu, Marzyeh Ghassemi

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

Comments Published in ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17451 2025-06-06 cs.AI cs.CV cs.GR 62%

Generating by Understanding: Neural Visual Generation with Logical Symbol Groundings

Yifei Peng, Zijie Zha, Yu Jin, Zhexu Luo, Wang-Zhou Dai, Zhong Ren, Yao-Xiang Ding, Kun Zhou

机构 * State Key Laboratory of CAD&CG(CAD与CG国家重点实验室) National Key Laboratory for Novel Software Technology(新型软件技术国家实验室) Department of Computer and Information Science(计算机与信息科学系)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments KDD 2025 research track paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04277 2025-06-06 cs.CV cs.AI 62%

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Yi Lu, Jiawang Cao, Yongliang Wu, Bozheng Li, Licheng Tang, Yangguang Ji, Chong Wu, Jay Wu, Wenbo Zhu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted as ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02890 2025-06-04 cs.LG cs.AI cs.CL 62%

Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights

Jakub Krajewski, Marcin Chochowski, Daniel Korzekwa

机构 * NVIDIA IDEAS NCBR, University of Warsaw(IDEAS NCBR,华沙大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01600 2025-06-03 cs.RO cs.AI cs.CV 62%

WoMAP: World Models For Embodied Open-Vocabulary Object Localization

Tenny Yin, Zhiting Mei, Tao Sun, Lihan Zha, Emily Zhou, Jeremy Bao, Miyu Yamane, Ola Shorinwa, Anirudha Majumdar

机构 * Princeton University(普林斯顿大学) McGill University(麦吉尔大学)

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00925 2025-06-03 q-bio.BM cs.CV cs.LG 62%

ProtInvTree: Deliberate Protein Inverse Folding with Reward-guided Tree Search

Mengdi Liu, Xiaoxue Cheng, Zhangyang Gao, Hong Chang, Cheng Tan, Shiguang Shan, Xilin Chen

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学) AI Lab, Research Center for Industries of the Future, Westlake University(未来产业研究院人工智能实验室) Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高陵人工智能学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24214 2025-06-02 cs.CV cs.AI 62%

Benchmarking Foundation Models for Zero-Shot Biometric Tasks

Redwan Sony, Parisa Farmanifard, Hamzeh Alzwairy, Nitish Shukla, Arun Ross

机构 * Department of Computer Science and Engineering(计算机科学与工程系)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00850 2025-05-30 cs.LG cs.AI cs.CL 62%

GWQ: Gradient-Aware Weight Quantization for Large Language Models

Yihua Shao, Yan Gu, Siyu Chen, Haiyang Liu, Zixian Zhu, Zijian Ling, Minxi Yan, Ziyang Yan, Chenyu Zhang, Michele Magno, Haotong Qin, Yan Wang, Jingcai Guo, Ling Shao, Hao Tang

机构 * PKU(北京大学) CASIA(中国科学院自动化研究所) THU(清华大学) USTB(中国矿业大学) UNITN(意大利那不勒斯) ETHz(苏黎世联邦理工学院) PolyU(香港理工大学) UCAS(中国科学院大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏