arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7387 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7387 篇

2501.01895 2025-11-18 cs.RO cs.CV cs.LG 62%

EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation

Siyuan Huang, Liliang Chen, Pengfei Zhou, Shengcong Chen, Zhengkai Jiang, Yue Hu, Yue Liao, Peng Gao, Hongsheng Li, Maoqing Yao, Guanghui Ren

机构 * SJTU(上海交通大学) AgiBot Shanghai AI Lab(上海人工智能实验室) CUHK MMLab(香港大学多模态实验室) LV-NUS Lab(南洋理工大学-立陶宛实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

Comments Accepted by NeurIPS 2025. Website: https://sites.google.com/view/enerverse

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11702 2025-11-18 cs.CV cs.AI eess.IV 62%

Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric Refinement

Lian He, Meng Liu, Qilang Ye, Yu Zhou, Xiang Deng, Gangyi Ding

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11583 2025-11-18 cs.LG cs.AI cs.IR 62%

Parallel and Multi-Stage Knowledge Graph Retrieval for Behaviorally Aligned Financial Asset Recommendations

Fernando Spadea, Oshani Seneviratne

机构 * Rensselaer Polytechnic Institute(伦斯勒理工学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 3 figures, RAGE-KG 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11502 2025-11-17 cs.CV cs.AI 62%

PAS : Prelim Attention Score for Detecting Object Hallucinations in Large Vision--Language Models

Nhat Hoang-Xuan, Minh Vu, My T. Thai, Manish Bhattarai

机构 * Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室) University of Florida(佛罗里达大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11450 2025-11-17 cs.CV cs.LG 62%

VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentation

Maximilian Rokuss, Moritz Langenberg, Yannick Kirchhoff, Fabian Isensee, Benjamin Hamm, Constantin Ulrich, Sebastian Regnery, Lukas Bauer, Efthimios Katsigiannopulos, Tobias Norajitra, Klaus Maier-Hein

机构 * German Cancer Research Center, Division of Medical Image Computing(德国癌症研究中心,医学影像计算部) Faculty of Mathematics and Computer Science(数学与计算机科学系) Medical Faculty - Heidelberg University(海德堡大学医学系) Helmholtz Imaging(海德堡影像技术) Department of Radiation Oncology, Heidelberg University Hospital(海德堡大学医院放射肿瘤科) HIDSS4Health, Heidelberg(HIDSS4Health,海德堡) Pattern Analysis and Learning Group, Heidelberg University Hospital(海德堡大学医院模式分析与学习组)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10020 2025-11-14 cs.CV cs.AI 62%

Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation

Yuxin Jiang, Wei Luo, Hui Zhang, Qiyu Chen, Haiming Yao, Weiming Shen, Yunkang Cao

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07685 2025-11-12 cs.AI cs.CL cs.LG 62%

ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents

Manasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi, Clinton Wang, Ankit Aich, Huy Nghiem, Tahseen Rabbani, Ye Htet, Brian Jang, Sumana Basu, Aishwarya Balwani, Denis Peskoff, Marcos Ayestaran, Sean M. Hendryx, Brad Kenstler, Bing Liu

机构 * Scale AI University of Maryland(马里兰大学) University of Chicago(芝加哥大学) Washington University, St. Louis(圣路易斯华盛顿大学) McGill University(麦吉尔大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments 27 pages, 21 figures, pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07238 2025-11-11 cs.CV cs.AI 62%

Leveraging Text-Driven Semantic Variation for Robust OOD Segmentation

Seungheon Song, Jaekoo Lee

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 5 figure references, 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07068 2025-11-11 cs.CV cs.LG 62%

ClusterMine: Robust Label-Free Visual Out-Of-Distribution Detection via Concept Mining from Text Corpora

Nikolas Adaloglou, Diana Petrusheva, Mohamed Asker, Felix Michels, Markus Kollmann

机构 * Heinrich Heine University of Düsseldorf(海因里希-海涅大学杜塞尔多夫分校)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

Comments Accepted in WACV 2026. Code in https://github.com/HHU-MMBS/clustermine_wacv_official 9 Tables, 11 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01019 2025-11-07 cs.CL cs.AI cs.CE cs.LG physics.ao-ph 62%

OceanAI: A Conversational Platform for Accurate, Transparent, Near-Real-Time Oceanographic Insights

Bowen Chen, Jayesh Gajbhar, Gregory Dusek, Rob Redmon, Patrick Hogan, Paul Liu, DelWayne Bohnenstiehl, Dongkuan Xu, Ruoying He

机构 * North Carolina State University(北卡罗来纳州立大学) NOAA(国家海洋和大气管理局)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments A related presentation will be given at the AGU(American Geophysical Union) and AMS(American Meteorological Society) Annual Meetings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01449 2025-11-04 cs.CV cs.AI 62%

Privacy Preserving Ordinal-Meta Learning with VLMs for Fine-Grained Fruit Quality Prediction

Riddhi Jain, Manasi Patwardhan, Aayush Mishra, Parijat Deshpande, Beena Rai

机构 * TCS-Research Pune, India(TCS-研究 Pune,印度)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 1 figure, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00269 2025-11-04 cs.CV cs.AI 62%

FedReplay: A Feature Replay Assisted Federated Transfer Learning Framework for Efficient and Privacy-Preserving Smart Agriculture

Long Li, Jiajia Li, Dong Chen, Lina Pu, Haibo Yao, Yanbo Huang

机构 * Department of Electrical and Computer Engineering, The University of Alabama(电气与计算机工程系,阿拉巴马大学) Electrical and Computer Engineering, Michigan State University(电气与计算机工程,密歇根州立大学) Agricultural and Biological Engineering, Mississippi State University(农业与生物工程,密苏里州立大学) Department of Computer Science, University of Alabama(计算机科学系,阿拉巴马大学) USDA-ARS Genetics and Sustainbale Agriculture(美国农业部ARS基因与可持续农业)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10821 2025-11-04 cs.CV cs.AI cs.CL 62%

VideoExplorer: Think With Videos For Agentic Long-Video Understanding

Huaying Yuan, Zheng Liu, Junjie Zhou, Hongjin Qian, Yan Shu, Nicu Sebe, Ji-Rong Wen, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Beijing University of Posts and Telecommunications(北京邮电大学) Hong Kong Polytechnic University(香港理工大学) Peking University(北京大学) University of Trento(特伦托大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02820 2025-10-31 cs.AI cs.CL cs.LG 62%

AutoLibra: Agent Metric Induction from Open-Ended Human Feedback

Hao Zhu, Phil Cuvin, Xinkai Yu, Charlotte Ka Yee Yan, Jason Zhang, Diyi Yang

机构 * Stanford University(斯坦福大学) University of Toronto(多伦多大学) University of Pennsylvania(宾夕法尼亚大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments https://github.com/Open-Social-World/autolibra

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25368 2025-10-30 cs.LG cs.AI cs.NE 62%

Position: Biology is the Challenge Physics-Informed ML Needs to Evolve

Julien Martinelli

机构 * ELLIS Institute Finland(芬兰ELLIS研究所) Department of Computer Science, Aalto University(艾尔沃斯大学计算机科学系)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25051 2025-10-30 cs.CV cs.LG 62%

Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models

Shunjie-Fabian Zheng, Hyeonjun Lee, Thijs Kooi, Ali Diba

机构 * Department of Medicine I, LMU University Hospital, LMU Munich, Germany(慕尼黑大学医学院第一医学部,LMU慕尼黑大学医院) Lunit Inc.(Lunit公司)

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV、cs.LG

Comments Accepted to Computer Vision for Automated Medical Diagnosis (CVAMD) Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24737 2025-10-30 eess.SP cs.AI cs.LG 62%

Cardi-GPT: An Expert ECG-Record Processing Chatbot

Koustav Mallick, Neel Singh, Mohammedreza Hajiarbabi

机构 * Department of Computer Science Purdue University Fort Wayne(计算机科学系 Purdue 大学 Fort Wayne)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Journal ref SoutheastCon 2025 352-357

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17394 2025-10-28 cs.CV cs.AI 62%

HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMs

Zhaolin Cai, Fan Li, Ziwei Zheng, Yanjun Qin

机构 * Xinjiang University(新疆大学) Xi'an Jiaotong University(西安交通大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20759 2025-10-28 cs.CV cs.AI 62%

PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding

Ansel Blume, Jeonghwan Kim, Hyeonjeong Ha, Elen Chatikyan, Xiaomeng Jin, Khanh Duy Nguyen, Nanyun Peng, Kai-Wei Chang, Derek Hoiem, Heng Ji

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California Los Angeles(加州大学洛杉矶分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025 Spotlight; project page: https://wjdghks950.github.io/partonomy.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21763 2025-10-28 cs.CV cs.AI 62%

Proportion and Perspective Control for Flow-Based Image Generation

Julien Boudier, Hugo Caselles-Dupré

机构 * Obvious Research(Obvious研究)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Technical report after open-source release

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20229 2025-10-24 cs.CV cs.AI cs.CL 62%

Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context

Ge Zheng, Jiaye Qian, Jiajin Tang, Sibei Yang

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) ShanghaiTech University(上海理工大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 4101-4113

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25033 2025-10-24 cs.CV cs.LG 62%

VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning

Wenhao Li, Qiangchang Wang, Xianjing Meng, Zhibin Wu, Yilong Yin

机构 * School of Software, Shandong University(山东大学软件学院) Shenzhen Loop Area Institute(深圳河套学院) School of Computing and Artificial Intelligence, Shandong University of Finance and Economics(山东财经大学计算机与人工智能学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17875 2025-10-22 cs.CV cs.AI 62%

3D Weakly Supervised Semantic Segmentation via Class-Aware and Geometry-Guided Pseudo-Label Refinement

Xiaoxu Xu, Xuexun Liu, Jinlong Li, Yitian Yuan, Qiudan Zhang, Lin Ma, Nicu Sebe, Xu Wang

机构 * College of Computer Science, Beihang University(北京航空航天大学计算机科学学院) College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) Department of Information Engineering and Computer Science, University of Trento(特伦托大学信息工程与计算机科学系) Meituan(美团)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17218 2025-10-21 cs.CV cs.AI 62%

When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions

Zhuo Cao, Heming Du, Bingqing Zhang, Xin Yu, Xue Li, Sen Wang

机构 * The University of Queensland, Australia(昆士兰大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16972 2025-10-21 cs.CV cs.AI 62%

The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA

Quanzhu Niu, Dengxian Gong, Shihao Chen, Tao Zhang, Yikang Zhou, Haobo Yuan, Lu Qi, Xiangtai Li, Shunping Ji

机构 * Wuhan University(武汉大学) University of California, Merced(加州大学默塞德分校) Nanyang Technological University(南洋理工大学)

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV、cs.AI

Comments The 1st place report of 7th LSVOS challenge RVOS track in ICCV 2025. The code is released in Sa2VA repository: https://github.com/bytedance/Sa2VA

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15998 2025-10-21 cs.LG cs.AI 62%

AMStraMGRAM: Adaptive Multi-cutoff Strategy Modification for ANaGRAM

Nilo Schwencke, Cyriaque Rousselot, Alena Shilova, Cyril Furtlehner

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13564 2025-10-20 cs.CV cs.AI 62%

HumorDB: Can AI understand graphical humor?

Vedaant Jain, Felipe dos Santos Alves Feitosa, Gabriel Kreiman

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of São Paulo(圣保罗大学) Harvard Medical School(哈佛医学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 10 main figures, 4 additional appendix figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21976 2025-10-16 cs.CV cs.AI 62%

Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning

Zilun Zhang, Zian Guan, Tiancheng Zhao, Haozhan Shen, Tianyu Li, Yuxiang Cai, Zhonggen Su, Zhaojun Liu, Jianwei Yin, Xiang Li

机构 * College of Computer Science and Technology of Zhejiang University(浙江大学计算机科学与技术学院) Polytechnic Institute of Zhejiang University(浙江大学Polytechnic学院) Om AI Research(Om AI研究机构) Binjiang Research Institute of Zhejiang University(浙江大学滨江研究机构) School of Software Engineering of Zhejiang University(浙江大学软件工程学院) School of Mathematical Sciences of Zhejiang University(浙江大学数学科学学院) China Academy of Space Technology(中国航天科技研究院) University of Bristol(布里斯托大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19972 2025-10-16 cs.CV cs.AI cs.CL 62%

GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity

Seongheon Park, Sharon Li

机构 * Department of Computer Sciences University of Wisconsin-Madison(计算机科学系威斯康星大学麦迪逊分校)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08775 2025-10-13 cs.CV cs.AI 62%

Re-Identifying Kākā with AI-Automated Video Key Frame Extraction

Paula Maddigan, Andrew Lensen, Rachael C. Shaw

机构 * Centre for Data Science and Artificial Intelligence, and School of Engineering and Computer Science(数据科学与人工智能中心,工程与计算机科学学院) Victoria University of Wellington(惠灵顿维多利亚大学) School of Biological Sciences(生物科学学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏