arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7387 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7387 篇

2510.05096 2025-10-10 cs.CV cs.AI cs.CL cs.MA cs.MM 62%

Paper2Video: Automatic Video Generation from Scientific Papers

Zeyu Zhu, Kevin Qinghong Lin, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(展示实验室,新加坡国立大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Project Page: https://showlab.github.io/Paper2Video/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08713 2025-10-10 cs.LG cs.AI 62%

ProtoECGNet: Case-Based Interpretable Deep Learning for Multi-Label ECG Classification with Contrastive Learning

Sahil Sethi, David Chen, Thomas Statchen, Michael C. Burkhart, Nipun Bhandari, Bashar Ramadan, Brett Beaulieu-Jones

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments Accepted to PMLR 298, 10th Machine Learning for Healthcare Conference (MLHC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03363 2025-10-09 cs.CV cs.AI eess.IV 62%

Unified Unsupervised Anomaly Detection via Matching Cost Filtering

Zhe Zhang, Mingxiu Cai, Gaochang Wu, Jing Zhang, Lingqiao Liu, Dacheng Tao, Tianyou Chai, Xiatian Zhu

机构 * State Key Laboratory of Synthetical Automation for Process Industries, Northeastern University, Shenyang, China(合成过程工业综合自动化国家重点实验室,东北大学,沈阳,中国) University of Surrey(Surrey大学) School of Computer Science, Wuhan University(武汉大学计算机学院) School of Computer Science, The University of Adelaide(阿德莱德大学计算机学院) College of Computing & Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) Surrey Institute for People-Centred Artificial Intelligence, and Centre for Vision, Speech and Signal Processing, University of Surrey(Surrey人本人工智能研究所,以及视觉、语音和信号处理中心,Surrey大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 63 pages (main paper and supplementary material), 39 figures, 58 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06512 2025-10-09 cs.CV cs.AI 62%

LogSTOP: Temporal Scores over Prediction Sequences for Matching and Retrieval

Avishree Khare, Hideki Okamoto, Bardh Hoxha, Georgios Fainekos, Rajeev Alur

机构 * seas.upenn.edu(宾夕法尼亚大学塞as学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09650 2025-10-06 cs.CV cs.LG cs.MM cs.RO eess.IV 62%

HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios

Kunyu Peng, Junchao Huang, Xiangsheng Huang, Di Wen, Junwei Zheng, Yufan Chen, Kailun Yang, Jiamin Wu, Chongqing Hao, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Beijing Institute of Technology(北京理工大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Hunan University(湖南大学) Shanghai AI Lab(上海人工智能实验室) HEBUST

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV、cs.LG

Comments Accepted to NeurIPS 2025. The dataset and code are available at https://github.com/KPeng9510/HopaDIFF

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05288 2025-10-03 cs.CV cs.AI cs.RO 62%

PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes

Ahmed Abdelreheem, Filippo Aleotti, Jamie Watson, Zawar Qureshi, Abdelrahman Eldesokey, Peter Wonka, Gabriel Brostow, Sara Vicente, Guillermo Garcia-Hernando

机构 * Niantic Spatial KAUST(科威特科学与技术研究中心) UCL(伦敦大学学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025. Project page: https://nianticlabs.github.io/placeit3d/

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14073 2025-10-03 stat.ML cs.AI cs.LG math.PR 62%

Neural Network Parameter-optimization of Gaussian pmDAGs

Mehrzad Saremi

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments 52 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22646 2025-10-02 cs.CV cs.AI cs.CL 62%

Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs

Xingyu Fu, Siyi Liu, Yinuo Xu, Pan Lu, Guangqiuse Hu, Tianbo Yang, Taran Anantasagar, Christopher Shen, Yikai Mao, Yuanzhe Liu, Keyush Shah, Chung Un Lee, Yejin Choi, James Zou, Dan Roth, Chris Callison-Burch

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Project Page: https://deeptracereward.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25379 2025-10-01 cs.LG cs.AI 62%

Let Physics Guide Your Protein Flows: Topology-aware Unfolding and Generation

Yogesh Verma, Markus Heinonen, Vikas Garg

机构 * Aalto University(阿alto大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24192 2025-09-30 cs.CV cs.AI 62%

Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection

Sojung An, Kwanyong Park, Yong Jae Lee, Donghyun Kim

机构 * Korea University(韩国大学) University of Seoul(首尔大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 23 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24107 2025-09-30 cs.AI cs.LG 62%

Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval and Synthesis for SLMs

Shreyas Singh, Kunal Singh, Pradeep Moturi

机构 * Fractal AI Research(Fractal AI研究)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15182 2025-09-30 cs.CL cs.AI cs.LG 62%

ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection

Jeonghye Kim, Sojeong Rhee, Minbeom Kim, Dohyung Kim, Sangmook Lee, Youngchul Sung, Kyomin Jung

机构 * KAIST(韩国科学技术院) Seoul National University(首尔国立大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments Accepted to EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23612 2025-09-30 cs.CV cs.AI 62%

InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects

Xinhao Cai, Minghang Zheng, Xin Jin, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机技术研究院) Beijing Electronic Science and Technology Institute(北京电子科学技术研究所) Wangxuan Institute of Computer Technology(北京大学计算机技术研究院) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12341 2025-09-30 cs.CV cs.AI 62%

Semantic Discrepancy-aware Detector for Image Forgery Identification

Ziye Wang, Minghang Yu, Chunyan Xu, Zhen Cui

机构 * Nanjing University of Science and Technology(南京理工大学) Beijing Normal University(北京师范大学)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20524 2025-09-26 cs.CV cs.AI 62%

InstructVTON: Optimal Auto-Masking and Natural-Language-Guided Interactive Style Control for Inpainting-Based Virtual Try-On

Julien Han, Shuwen Qiu, Qi Li, Xingzi Xu, Mehmet Saygin Seyfioglu, Kavosh Asadi, Karim Bouyarmane

机构 * Amazon(亚马逊公司) University of California, Los Angeles (UCLA)(加州大学洛杉矶分校) Duke University(杜克大学)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV、cs.AI

Comments Submitted to CVPR 2025 and Published at CVPR 2025 AI for Content Creation workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16421 2025-09-26 cs.CV cs.AI 62%

AHA -- Predicting What Matters Next: Online Highlight Detection Without Looking Ahead

Aiden Chang, Celso De Melo, Stephanie M. Lukin

机构 * University of Southern California(南加州大学) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted at NeurIPS 2025, 32 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08705 2025-09-26 cs.CV cs.AI 62%

Instance-aware Image Colorization with Controllable Textual Descriptions and Segmentation Masks

Yanru An, Ling Gui, Chunlei Cai, Tianxiao Ye, JIangchao Yao, Guangtao Zhai, Qiang Hu, Xiaoyun Zhang

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 视觉定位与Grounding :visual language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10787 2025-09-26 cs.CV cs.LG 62%

Lightweight Modular Parameter-Efficient Tuning for Open-Vocabulary Object Detection

Bilal Faye, Hanane Azzag, Mustapha Lebbah

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19017 2025-09-24 cs.LG cs.AI 62%

Fully Learnable Neural Reward Machines

Hazem Dewidar, Elena Umili

机构 * La Sapienza University of Rome(拉维亚大学罗马分校) Peoples' Friendship University of Russia (RUDN University)(俄罗斯人民友谊大学) Joint Institute for Nuclear Research(联合核子研究所) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) University of Skövde(斯德哥尔摩大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21535 2025-09-23 eess.IV cs.CV cs.LG 62%

Exploring the Design Space of 3D MLLMs for CT Report Generation

Mohammed Baharoon, Jun Ma, Congyu Fang, Augustin Toma, Bo Wang

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所) Department of Biomedical Informatics, Harvard Medical School(生物医学信息学系,哈佛医学院) Peter Munk Cardiac Centre, University Health Network(皮特·蒙克心脏中心,大学健康网络) Medical Biophysics, University of Toronto(医学生物物理系,多伦多大学) Department of Computer Science, University of Toronto(计算机科学系,多伦多大学) Department of Laboratory Medicine and Pathobiology, University of Toronto(实验室医学与病理学系,多伦多大学) AI Hub, University Health Network(人工智能中心,大学健康网络)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20742 2025-09-18 cs.CV cs.AI cs.CL 62%

Structured Preference Optimization for Vision-Language Long-Horizon Task Planning

Xiwen Liang, Min Lin, Weiqi Ruan, Rongtao Xu, Yuecheng Liu, Jiaqi Chen, Bingqian Lin, Yuzheng Zhuang, Xiaodan Liang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13255 2025-09-17 cs.CV cs.AI cs.IR eess.IV 62%

ResidualViT for Efficient Temporally Dense Video Encoding

Mattia Soldan, Fabian Caba Heilbron, Bernard Ghanem, Josef Sivic, Bryan Russell

机构 * KAUST(科威特科学与技术王国(KAUST)) CIIRC CTU(布拉格技术大学智能信息与计算研究中心(CIIRC CTU)) Adobe Research(Adobe研究)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19331 2025-09-17 cs.CV cs.AI cs.CL 62%

Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation

Luca Barsellotti, Lorenzo Bianchi, Nicola Messina, Fabio Carrara, Marcella Cornia, Lorenzo Baraldi, Fabrizio Falchi, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) ISTI-CNR(意大利国家研究委员会ISTI) University of Pisa(比萨大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02074 2025-09-10 cs.CV cs.AI 62%

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges

Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma

机构 * Department of Computer Science, Iowa State University(计算机科学系,爱荷华州立大学) Department of Civil, Construction and Environmental Engineering, Iowa State University(土木、建设与环境工程系,爱荷华州立大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19787 2025-09-09 cs.LG cs.AI 62%

CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives

Armin Saghafian, Amirmohammad Izadi, Negin Hashemi Dijujin, Mahdieh Soleymani Baghshah

机构 * Sharif University of Technology(谢里夫理工大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments Accepted to TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05751 2025-09-09 cs.CV cs.AI 62%

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation

Bingrui Zhao, Lin Yuanbo Wu, Xiangtian Fan, Deyin Liu, Lu Zhang, Ruyi He, Jialie Shen, Ximing Li

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04894 2025-09-08 cs.CV cs.LG 62%

SynGen-Vision: Synthetic Data Generation for training industrial vision models

Alpana Dubey, Suma Mani Kuriakose, Nitish Bhardwaj

机构 * Accenture Labs(埃森哲实验室)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04180 2025-09-05 cs.CV cs.AI 62%

VisioFirm: Cross-Platform AI-assisted Annotation Tool for Computer Vision

Safouane El Ghazouali, Umberto Michelucci

机构 * TOELT LLC AI lab(TOELT LLC人工智能实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04549 2025-09-03 cs.CV cs.AI cs.MM 62%

MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning

Quang-Trung Truong, Yuk-Kwan Wong, Vo Hoang Kim Tuyen Dang, Rinaldi Gotama, Duc Thanh Nguyen, Sai-Kit Yeung

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Ho Chi Minh University of Science(胡志明市科学大学) Deakin University(德肯大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Published at ACMMM2025 (Dataset track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18322 2025-08-27 cs.CV cs.AI 62%

Structures Meet Semantics: Multimodal Fusion via Graph Contrastive Learning

Jiangfeng Sun, Sihao He, Zhonghong Ou, Meina Song

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments 9 pages,7 figures,conference

详情

展开后加载摘要…

URL PDF HTML 收藏