arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7370 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7370 篇

2601.21199 2026-01-30 cs.CV cs.AI 73%

Thinker: A vision-language foundation model for embodied intelligence

Thinker:一个用于具身智能的视觉-语言基础模型

Baiyu Pan, Daqin Luo, Junpeng Yang, Jiyuan Wang, Yixuan Zhang, Hailin Shi, Jichao Jiao

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 Thinker通过构建大规模数据集和改进输入方式,在机器人感知与推理任务中实现了最先进的性能。

Comments IROS 2025, 4 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04655 2026-01-29 cs.CV cs.AI 73%

X-SAM: From Segment Anything to Any Segmentation

X-SAM:从Segment Anything到Any Segmentation

Hao Wang, Limeng Qiao, Zequn Jie, Zhijian Huang, Chengjian Feng, Qingfang Zheng, Lin Ma, Xiangyuan Lan, Xiaodan Liang

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 X-SAM通过引入统一框架和VGD分割任务,提升多模态大语言模型的像素级视觉理解能力,实现更广泛的分割任务整合。

Comments AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19222 2026-01-28 cs.CV cs.AI 73%

UniPCB: A Unified Vision-Language Benchmark for Open-Ended PCB Quality Inspection

UniPCB: 一个统一的视觉-语言基准用于开放式PCB质量检测

Fuxiang Sun, Xi Jiang, Jiansheng Wu, Haigang Zhang, Feng Zheng, Jinfeng Yang

机构 * Shenzhen Polytechnic University(深圳职业技术大学) Southern University of Science and Technology(南方科技大学) University of Science and Technology Liaoning(辽宁科技大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 UniPCB提出一个统一的视觉-语言基准用于开放式PCB质量检测,通过系统化流程和PCB-GPT模型提升缺陷定位性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11031 2026-01-27 cs.LG cs.AI cs.CL 73%

Prefill-Guided Thinking for zero-shot detection of AI-generated images

预填充引导的思考用于AI生成图像的零样本检测

Zoher Kachwala, Danishjeet Singh, Danielle Yang, Filippo Menczer

机构 * Observatory on Social Media(社会媒体观察所) Indiana University(印第安纳大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.AI、cs.LG

AI总结 本文提出预填充引导思考方法,通过引导视觉-语言模型推理提升AI生成图像的零样本检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17844 2026-01-27 cs.HC cs.AI cs.LG 73%

RAICL: Retrieval-Augmented In-Context Learning for Vision-Language-Model Based EEG Seizure Detection

RAICL:基于视觉-语言模型的EEG癫痫检测的检索增强上下文学习

Siyang Li, Zhuoya Wang, Xiyan Gui, Xiaoqing Chen, Ziwei Wang, Yaozhi Wen, Dongrui Wu

机构 * Ministry of Education Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(教育部图像处理与智能控制重点实验室,人工智能与自动化学院,华中科技大学) State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences(脑认知与脑启发智能技术国家重点实验室,自动化研究所,中国科学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.AI、cs.LG

AI总结 RAICL通过利用视觉-语言模型分析EEG波形图,实现了更高效的癫痫检测,无需重新训练,具有广泛临床应用前景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11808 2026-01-27 cs.CV cs.AI cs.CL cs.CY cs.MM 73%

Labels or Input? Rethinking Augmentation in Multimodal Hate Detection

标签还是输入?重新思考多模态仇恨检测中的增强

Sahajpreet Singh, Kokil Jaidka, Subhayan Mukerjee

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

AI总结 本文提出通过提示优化、微调和自动化数据增强改进小型模型,开发多模态增强框架以提升隐含仇恨检测性能。

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00662 2026-01-27 cs.CV cs.CL cs.LG 73%

Mitigating the Modality Gap: Few-Shot Out-of-Distribution Detection with Multi-modal Prototypes and Image Bias Estimation

弥合模态差距:基于多模态原型和图像偏差估计的少样本分布外检测

Yimu Wang, Evelien Riddell, Adrian Chow, Sean Sedwards, Krzysztof Czarnecki

机构 * University of Waterloo(滑铁卢大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.LG

AI总结 本文提出SUPREME框架,通过引入多模态原型和图像偏差估计,有效缓解图像与文本之间的模态差距,提升少样本分布外检测性能。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13142 2026-01-21 cs.CV cs.AI cs.CL 73%

TVWorld: Foundations for Remote-Control TV Agents

TVWorld: 电视遥控的基石

Zhantao Ma, Quanfeng Lu, Shuai Zhong, Dahai Yu, Ping Luo, Michael K. Ng

机构 * The University of Hong Kong(香港大学) Hong Kong Baptist University(香港 Baptist 大学) TCL Corporate Research (Hong Kong) Co., Ltd(TCL 香港企业研究有限公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 TVWorld提出了一种基于图的电视导航抽象,开发了TVWorld-N和TVWorld-G两个基准测试,通过拓扑感知训练框架TVTheseus实现了68.3%的成功率,超越现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24097 2026-01-01 cs.CV cs.AI cs.CL cs.MM 73%

Factorized Learning for Temporally Grounded Video-Language Models

分解学习用于时间感知的视频-语言模型

Wenzheng Zeng, Difei Gao, Mike Zheng Shou, Hwee Tou Ng

机构 * National University of Singapore(新加坡国立大学)

专题命中 视觉定位与Grounding :VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 本文提出D$^2$VLM框架,通过分解学习方法提升视频-语言模型在时间定位和文本响应任务中的性能,引入证据标记和FPO算法以优化学习过程。

Comments ICCV 2025 paper. This arXiv version updates Figure 1 to include the concurrent work Qwen2.5-VL to ensure consistency with Table 1

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23244 2025-12-30 cs.CV cs.AI 73%

ViLaCD-R1: A Vision-Language Framework for Semantic Change Detection in Remote Sensing

ViLaCD-R1: 一种用于遥感语义变化检测的视觉-语言框架

Xingwei Ma, Shiyang Feng, Bo Zhang, Bin Wang

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

AI总结 ViLaCD-R1通过多图像推理器和掩码引导解码器,提升遥感变化检测的语义识别与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14312 2025-12-17 cs.CV cs.AI 73%

From YOLO to VLMs: Advancing Zero-Shot and Few-Shot Detection of Wastewater Treatment Plants Using Satellite Imagery in MENA Region

从YOLO到VLMs:利用卫星图像在中东和北非地区推进零样本和少样本废水处理厂检测

Akila Premarathna, Kanishka Hewageegana, Garcia Andarcia Mariangel

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

AI总结 本研究利用VLMs替代YOLOv8,通过零样本和少样本方法高效识别中东和北非地区废水处理厂,提升遥感应用的可扩展性。

Comments 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04728 2025-12-05 cs.CV cs.AI 73%

Measuring the Unspoken: A Disentanglement Model and Benchmark for Psychological Analysis in the Wild

测量未言之: 一种解耦模型和基准用于野外心理分析

Yigui Feng, Qinglin Wang, Haotian Mo, Yang Liu, Ke Liu, Gencheng Liu, Xinhai Chen, Siqi Shen, Songzhu Mei, Jie Liu

机构 * College of Computer Science, National University of Defense Technology(计算机科学学院,国防科技大学) Shien-Ming Wu School of Intelligent Engineering, South China University of Technology(智能工程学院,华南理工大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 本文提出MIND模型和PRISM基准,通过解耦算法和专家标注数据提升野外对话心理分析的准确性与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02835 2025-12-03 cs.CV cs.AI cs.CL 73%

ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning

ReVSeg:通过强化学习激励视频分割的推理链

Yifan Li, Yingda Yin, Lingting Zhu, Weikai Chen, Shengju Qian, Xin Wang, Yanwei Fu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) LIGHTSPEED

专题命中 视觉定位与Grounding :vision language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 ReVSeg通过强化学习优化视频分割的多步推理链,实现可解释的推理轨迹和先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00287 2025-12-02 cs.RO cs.AI cs.CV 73%

RealAppliance: Let High-fidelity Appliance Assets Controllable and Workable as Aligned Real Manuals

RealAppliance: 让高保真家电资产可控且可操作,如对齐的现实手册

Yuzheng Gao, Yuxing Long, Lei Kang, Yuchong Guo, Ziyan Yu, Shangqing Mao, Jiyao Zhang, Ruihai Wu, Dongjiang Li, Hui Shen, Hao Dong

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 RealAppliance通过高保真家电资产和对齐手册的基准测试,推动家电操控技术的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18164 2025-11-25 cs.CV cs.AI 73%

Nested Unfolding Network for Real-World Concealed Object Segmentation

嵌套展开网络用于现实世界隐蔽物体分割

Chunming He, Rihan Zhang, Dingming Zhang, Fengyang Xiao, Deng-Ping Fan, Sina Farsiu

机构 * Duke University(杜克大学) Nankai University(南开大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

AI总结 本文提出嵌套展开网络(NUN),通过解耦恢复与分割并引入自一致性损失,提升现实世界隐蔽物体分割的性能。

Comments 6 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13442 2025-11-19 cs.CV cs.AI 73%

Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline

Rui Zuo, Qinyue Tong, Zhe-Ming Lu, Ziqian Lu

机构 * Zhejiang University(浙江大学) Zhejiang Sci-Tech University(浙江科技学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00411 2025-11-18 cs.CV cs.AI 73%

Does Bigger Mean Better? Comparitive Analysis of CNNs and Biomedical Vision Language Modles in Medical Diagnosis

Ran Tong, Jiaqi Liu, Tong Wang, Xin Hu, Su Liu, Lanruo Wang, Jiexi Xu

机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校) Independent Researcher(独立研究者) Duke University(杜克大学) University of Michigan Ann Arbor(密歇根大学安娜堡分校) Georgia Institute of Technology(佐治亚理工学院) University of California, Irvine(加州大学 Irvine 分校)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments 6pages,3 figures.Uunder review of International Conference on Artificial Intelligence, Computer, Data Sciences and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03662 2025-11-17 cs.CV cs.AI cs.RO 73%

Zero-Shot Temporal Interaction Localization for Egocentric Videos

Erhang Zhang, Junyi Ma, Yin-Dong Zheng, Yixuan Zhou, Hesheng Wang

机构 * IRMV Lab, the Department of Automation, Shanghai Jiao Tong University(IRMV实验室,自动化系,上海交通大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Accepted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09958 2025-11-14 cs.CV cs.AI 73%

Zero-Shot Referring Expression Comprehension via Vison-Language True/False Verification

Jeffrey Liu, Rongbin Hu

专题命中 视觉定位与Grounding :VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05565 2025-11-11 cs.CV cs.AI 73%

In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy

Shreyan Ganguly, Angona Biswas, Jaydeep Rade, Md Hasibul Hasan Hasib, Nabila Masud, Nitish Singla, Abhipsa Dash, Ushashi Bhattacharjee, Aditya Balu, Anwesha Sarkar, Adarsh Krishnamurthy, Soumik Sarkar

机构 * Iowa State University(爱荷华州立大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03757 2025-11-07 cs.LG cs.AI 73%

Laugh, Relate, Engage: Stylized Comment Generation for Short Videos

Xuan Ouyang, Senan Wang, Bouzhou Wang, Siyuan Xiahou, Jinrong Zhou, Yuekang Li

机构 * University of New South Wales(新南威尔士大学) University of Sydney(悉尼大学) The University of Hong Kong(香港大学) University of Southern California(南加州大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27164 2025-11-03 cs.CV cs.AI 73%

Generating Accurate and Detailed Captions for High-Resolution Images

Hankyeol Lee, Gawon Seo, Kyounggyu Lee, Dogun Kim, Kyungwoo Song, Jiyoung Jung

机构 * Department of Artificial Intelligence, University of Seoul(首尔大学人工智能系) Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系) Department of Applied Statistics, Yonsei University(延世大学应用统计系)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Work conducted in 2024; released for archival purposes

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26151 2025-10-31 cs.CV cs.AI 73%

MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction

Shunjie-Fabian Zheng, Hyeonjun Lee, Thijs Kooi, Ali Diba

机构 * Department of Medicine I, LMU University Hospital, LMU Munich, Germany(慕尼黑大学医学部第一部门,LMU大学医院,慕尼黑,德国) Lunit Inc.(Lunit公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Accepted to Computer Vision for Automated Medical Diagnosis (CVAMD) Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25616 2025-10-30 cs.LG cs.AI cs.RO 73%

Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization

Nikita Kachaev, Mikhail Kolosov, Daniil Zelezetsky, Alexey K. Kovalev, Aleksandr I. Panov

机构 * Cognitive AI Lab(认知人工智能实验室) Cognitive AI Lab, IAI MIPT(认知人工智能实验室,IAI MIPT)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.AI、cs.LG

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13891 2025-10-17 cs.LG cs.AI 73%

K-frames: Scene-Driven Any-k Keyframe Selection for long video understanding

Yifeng Yao, Yike Yun, Jing Wang, Huishuai Zhang, Dongyan Zhao, Ke Tian, Zhihao Wang, Minghui Qiu, Tao Wang

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) Bytedance(字节跳动)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17473 2025-10-17 cs.CV cs.AI cs.CL 73%

InfoDet: A Dataset for Infographic Element Detection

Jiangning Zhu, Yuxing Zhou, Zheng Wang, Juntao Yao, Yima Gu, Yuhui Yuan, Shixia Liu

机构 * BNRist, Tsinghua University(清华大学信息与技术研究院)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments Submitted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10560 2025-10-14 cs.CL cs.AI cs.CV 73%

BitMar: Low-Bit Multimodal Fusion with Episodic Memory for Edge Devices

Euhid Aman, Esteban Carlin, Hsing-Kuo Pao, Giovanni Beltrame, Ghaluh Indah Permata Sari, Yie-Tarng Chen

机构 * NTUST(国立台湾科技大学) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments 6 pages, BabyLM Workshop, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10264 2025-10-14 cs.CV cs.AI 73%

MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs

Haonan Ge, Yiwei Wang, Ming-Hsuan Yang, Yujun Cai

机构 * Department of Computer Science and Engineering, University of California at Merced(计算机科学与工程系,加州大学默塞德分校)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21447 2025-10-13 cs.CV cs.AI 73%

Multimodal Language Models See Better When They Look Shallower

Haoran Chen, Junyan Lin, Xinghao Chen, Yue Fan, Jianfeng Dong, Xin Jin, Hui Su, Jinlan Fu, Xiaoyu Shen

机构 * Zhejiang Gongshang University(浙江工商大学) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) Institute of Digital Twin, Eastern Institute of Technology, Ningbo(数字孪生研究院,东部技术研究所,宁波) Meituan Inc.(美团公司) National University of Singapore(新加坡国立大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 6 figures, accepted by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17374 2025-09-30 cs.CV cs.AI cs.IR 73%

From Drawings to Decisions: A Hybrid Vision-Language Framework for Parsing 2D Engineering Drawings into Structured Manufacturing Knowledge

Muhammad Tayyab Khan, Lequn Chen, Zane Yong, Jun Ming Tan, Wenhe Feng, Seung Ki Moon

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Preprint submitted to Elsevier

详情

展开后加载摘要…

URL PDF HTML 收藏