arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 4454 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 4454 篇

2510.25191 2026-03-05 cs.RO 82%

SoraNav: Adaptive UAV Task-Centric Navigation via Zeroshot VLM Reasoning

SoraNav: 通过零样本VLM推理实现适应性无人机任务导向导航

Hongyu Song, Rishabh Dev Yadav, Cheng Guo, Wei Pan

机构 * Department of Computer Science, The University of Manchester(计算机科学系,曼彻斯特大学)

专题命中 视觉推理 :VLM(title,abstract);vision-language model(abstract)

AI总结 SoraNav通过零样本VLM推理实现无人机任务导向导航,解决空间-语义差距问题,提升导航成功率与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21628 2026-03-04 cs.CL 82%

RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoning

RuCL:基于分层评分标准的课程学习用于多模态大语言模型推理

Yukun Chen, Jiaming Li, Longze Chen, Ze Gong, Jingpeng Li, Zhen Qin, Hengyu Chang, Ancheng Xu, Zhihao Yang, Hamid Alinejad-Rokny, Qiang Qu, Bo Zheng, Min Yang

机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) University of Chinese Academy of Sciences(中国科学院大学) Alibaba Group(阿里巴巴集团) School of Biomedical Engineering, UNSW Sydney(新南威尔士大学生物医学工程学院)

专题命中 视觉推理 :multimodal large language model(title,abstract);visual reasoning(abstract)

AI总结 RuCL通过分层评分标准课程学习提升多模态大语言模型的推理能力,实现7.83%的准确率提升。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06749 2026-03-03 cs.CV cs.AI cs.CL cs.LG 82%

Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Vision-R1: 促进多模态大语言模型推理能力的激励方法

Wenxuan Huang, Bohan Jia, Zijie Zhai, Shaosheng Cao, Zheyu Ye, Fei Zhao, Zhe Xu, Xu Tang, Yao Hu, Shaohui Lin

机构 * East China Normal University(华东师范大学) The Chinese University of Hong Kong(香港中文大学) Xiaohongshu Inc.(小红书公司)

专题命中 视觉推理 :multimodal large language model(title);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 Vision-R1通过构建高质量多模态CoT数据集和渐进性思维抑制训练策略,提升多模态推理能力,在数学推理基准中取得显著成绩。

Comments Accepted to ICLR 2026. Code is available at https://github.com/Osilly/Vision-R1

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20330 2026-02-25 cs.CV cs.AI cs.LG 82%

Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking

视觉-语言模型中的电路追踪:理解多模态思维的内部机制

Jingcheng Yang, Tianhu Xiong, Shengyi Qian, Klara Nahrstedt, Mingyuan Wu

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出首个透明电路追踪框架,用于分析视觉-语言模型的多模态推理机制,揭示不同视觉特征电路在数学推理和跨模态关联中的作用。

Comments To appear in the Findings of CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16157 2026-02-19 cs.HC 82%

Peeking Ahead of the Field Study: Exploring VLM Personas as Support Tools for Embodied Studies in HCI

提前窥视田野研究:探索VLM人设作为人机交互领域具身研究的支持工具

Xinyue Gui, Ding Xia, Mark Colley, Yuan Li, Vishal Chauhan, Anubhav Anubhav, Zhongyi Zhou, Ehsan Javanmardi, Stela Hanbyeol Seo, Chia-Ming Chang, Manabu Tsukada, Takeo Igarashi

专题命中 视觉推理 :VLM(title,abstract);vision-language model(abstract)

AI总结 本文提出利用VLM人设模拟田野研究结果,验证其在具身研究中的应用潜力,发现其在响应模式上接近人类但缺乏多样性。

Comments Accepted to CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15237 2026-02-18 cs.HC 82%

Ground-Truth Depth in Vision Language Models: Spatial Context Understanding in Conversational AI for XR-Robotic Support in Emergency First Response

视觉语言模型中的真实深度:用于虚拟现实机器人支持的紧急急救中的空间情境理解

Rodrigo Gutierrez Maquilon, Marita Hueber, Georg Regal, Manfred Tscheligi

专题命中 视觉推理 :vision language model(title,abstract);VLM(abstract)

AI总结 本文提出了一种结合深度传感与视觉语言模型的原型,用于增强紧急急救中的空间情境理解,通过深度增强提高准确性和稳定性,同时降低工作量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17587 2026-02-03 cs.CV cs.AI cs.LG 82%

HalluRNN: Mitigating Hallucinations via Recurrent Cross-Layer Reasoning in Large Vision-Language Models

HalluRNN: 通过大规模视觉-语言模型中的递归跨层推理缓解幻觉

Le Yu, Kaishen Wang, Jianlong Xiong, Yue Cao, Lei Zhang, Zhang Yi Tao He

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 HalluRNN通过递归跨层推理模块缓解大规模视觉-语言模型中的幻觉问题,提升模型稳定性和性能。

Comments 6 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00211 2026-02-03 cs.CV cs.AI cs.LG 82%

Interpretable Unsupervised Deformable Image Registration via Confidence-bound Multi-Hop Visual Reasoning

可解释的无监督变形图像配准:通过置信度界限多跳视觉推理

Zafar Iqbal, Anwar Ul Haq, Srimannarayana Grandhi

机构 * School of Engineering and Technology(工程与技术学院)

专题命中 视觉推理 :visual reasoning(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出了一种基于多跳视觉推理的可解释无监督变形图像配准框架,通过局部空间细化和跨参考注意力机制,实现高精度且透明的医学图像配准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22714 2026-02-02 cs.LG cs.AI cs.CV 82%

Vision-Language Models Unlock Task-Centric Latent Actions

视觉-语言模型解锁任务导向的潜在动作

Alexander Nikulin, Ilya Zisman, Albina Klepach, Denis Tarasov, Alexander Derevyagin, Andrei Polubarov, Lyubaykin Nikita, Vladislav Kurenkov

机构 * Innopolis University(因诺波利斯大学) Research Center for Trusted Artificial Intelligence, ISP RAS(可信人工智能研究所以及ISP俄罗斯科学院)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出利用视觉-语言模型的常识推理能力,通过可提示的表示分离可控变化与噪声,提升潜在动作质量,从而提高下游任务性能。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17474 2025-12-22 eess.AS cs.SD 82%

Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models

基于商业自动语音识别和多模态大语言模型的口吃语音零样本识别

Ali Alsayegh, Tariq Masood

专题命中 视觉推理 :multimodal large language model(title,abstract);MLLM(abstract)

AI总结 本研究评估了商业ASR和MLLM在口吃语音识别中的性能,发现GPT-4o在重度口吃情况下显著降低WER,而Gemini变体表现下降,为辅助语音接口技术选择提供实证依据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11464 2025-12-15 cs.CV cs.AI cs.LG 82%

Exploring MLLM-Diffusion Information Transfer with MetaCanvas

探索MLLM-扩散信息传输与MetaCanvas

Han Lin, Xichen Pan, Ziqi Huang, Ji Hou, Jialiang Wang, Weifeng Chen, Zecheng He, Felix Juefei-Xu, Junzhe Sun, Zhipeng Fan, Ali Thabet, Mohit Bansal, Chu Wang

机构 * Meta Superintelligence Labs(Meta 超智能实验室) New York University(纽约大学) Nanyang Technological University(南洋理工大学)

专题命中 视觉推理 :MLLM(title);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 MetaCanvas通过使MLLMs在潜在空间中推理和规划,提升多模态生成的精确度和结构化控制。

Comments Project page: https://metacanvas.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09349 2025-12-11 cs.RO 82%

COVLM-RL: Critical Object-Oriented Reasoning for Autonomous Driving Using VLM-Guided Reinforcement Learning

COVLM-RL:基于VLM引导强化学习的自动驾驶中关键对象导向推理

Lin Li, Yuxin Cai, Jianwu Fang, Jianru Xue, Chen Lv

机构 * School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与航空航天工程学院) National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, National Engineering Research Center for Visual Information and Applications, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学人机混合增强智能国家重点实验室、视觉信息与应用国家工程研究中心、人工智能与机器人研究院)

专题命中 视觉推理 :VLM(title,abstract);vision-language model(abstract)

AI总结 COVLM-RL通过结合关键对象导向推理与VLM引导强化学习,提升了自动驾驶系统的泛化能力和训练效率。

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07177 2025-12-09 cs.RO 82%

Using Vision-Language Models as Proxies for Social Intelligence in Human-Robot Interaction

利用视觉-语言模型作为社交智能代理在人机交互中的应用

Fanjun Bu, Melina Tsai, Audrey Tjokro, Tapomayukh Bhattacharjee, Jorge Ortiz, Wendy Ju

机构 * Cornell University, Cornell Tech(康奈尔大学,康奈尔科技) Cornell University(康奈尔大学) Rutgers University(罗格斯大学) Cornell Tech, Jacobs Technion Cornell Institute(康奈尔科技,雅各布斯技术学院康奈尔研究所)

专题命中 视觉推理 :vision-language model(title,abstract);VLM(abstract)

AI总结 本文提出利用视觉-语言模型作为代理,通过轻量级感知检测器触发,实现机器人在人机交互中的社交响应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20531 2025-11-26 cs.AI cs.CV cs.LG 82%

Beyond Generation: Multi-Hop Reasoning for Factual Accuracy in Vision-Language Models

超越生成:面向视觉语言模型事实准确性的多跳推理

Shamima Hossain

机构 * Department of Computer Science, Brac University(布鲁尔大学计算机科学系) bKash Limited(bKash公司)

专题命中 视觉推理 :vision-language model(title);visual language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出一种基于知识图谱的多跳推理框架,提升视觉语言模型在事实准确性上的表现,通过多步骤推理增强模型的逻辑推理能力。

Comments Accepted as poster at NewInML Workshop ICML, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12008 2025-11-18 cs.AI cs.CV cs.LG 82%

Adaptive Diagnostic Reasoning Framework for Pathology with Multimodal Large Language Models

Yunqi Hong, Johnson Kao, Liam Edwards, Nein-Tzu Liu, Chung-Yen Huang, Alex Oliveira-Kowaleski, Cho-Jui Hsieh, Neil Y. C. Lin

机构 * Computer Science Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校计算机科学系) Mechanical and Aerospace Engineering Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校机械与航空航天工程系) Department of Pathology, Tri-Service General Hospital, National Defense Medical Center, Taipei, Taiwan(台湾国防医学院三军总医院病理部) Department of Pathology, National Taiwan University Hospital, Taipei, Taiwan(台湾国立台湾大学医院病理部) Department of Pathology, David Geffen School of Medicine, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校大卫·Geffen医学院病理部) Bioengineering Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校生物工程系) Institute for Quantitative and Computational Biosciences, University of California, CA, USA(加州大学定量与计算生物科学研究所)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11831 2025-11-18 cs.AI cs.CV cs.LG 82%

TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models

Wenhao Zhou, Hao Zheng, Rong Zhao

机构 * Center for Brain-Inspired Computing Research (CBICR)(脑启发计算研究中心) Department of Precision Instruments(精密仪器系) IDG/McGovern Institute for Brain Research(IDG/麦戈文脑研究学院)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21572 2025-11-14 cs.CL 82%

Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling

Shengwu. Xiong, Tianyu. Zou, Cong. Wang, Xuelong Li

机构 * Interdisciplinary Artificial Intelligence Research Institute, Wuhan College(交叉学科人工智能研究 institute,武汉学院) School of Computer and Artificial Intelligence, Wuhan University of Technology(计算机与人工智能学院,武汉理工大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Sanya Science and Education Innovation Park, Wuhan University of Technology(三亚科学教育创新园,武汉理工大学) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) School of Mathematics and Statistics, Northwestern Polytechnical University(数学与统计学院,西北工业大学) Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究所(TeleAI))

专题命中 视觉推理 :MLLM(title,abstract);multimodal large language model(abstract)

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17042 2025-09-23 cs.RO 82%

Orchestrate, Generate, Reflect: A VLM-Based Multi-Agent Collaboration Framework for Automated Driving Policy Learning

Zengqi Peng, Yusen Xie, Yubin Wang, Rui Yang, Qifeng Chen, Jun Ma

机构 * Robotics and Autonomous Systems Thrust, The Hong Kong University of Science and Technology (Guangzhou)(机器人与自主系统方向,香港科学与技术大学(广州)) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科学与技术大学) Cheng Kar-Shun Robotics Institute, The Hong Kong University of Science and Technology(陈家骏机器人研究所,香港科学与技术大学)

专题命中 视觉推理 :VLM(title,abstract);vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06729 2025-09-17 cs.RO 82%

STRIVE: Structured Representation Integrating VLM Reasoning for Efficient Object Navigation

Haokun Zhu, Zongtai Li, Zhixuan Liu, Wenshan Wang, Ji Zhang, Jonathan Francis, Jean Oh

机构 * Carnegie Mellon University(卡内基梅隆大学) Bosch Center for AI(博世人工智能中心)

专题命中 视觉推理 :VLM(title,abstract);vision-language model(abstract)

Comments We remove OSG and CogNav from Table. 1 for a fair comparison

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18381 2025-08-27 cs.CL 82%

Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models

Yuchun Fan, Yilin Wang, Yongyu Mu, Lei Huang, Bei Li, Xiaocheng Feng, Tong Xiao, Jingbo Zhu

专题命中 视觉推理 :vision-language model(title,abstract);visual reasoning(abstract)

Comments Accepted by EMNLP 2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11918 2025-08-19 cs.RO 82%

ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models

Zhichen Lou, Kechun Xu, Zhongxiang Zhou, Rong Xiong

机构 * State Key Laboratory of Industrial Control Technology(工业控制技术国家重点实验室) Institute of Cyber-Systems and Control(网络系统与控制研究所) Zhejiang University(浙江大学) Zhejiang Humanoid Robot Innovation Center Co., Ltd.(浙江人形机器人创新中心有限公司)

专题命中 视觉推理 :vision-language model(title,abstract);VLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10371 2025-08-15 cs.RO 82%

Few-shot Vision-based Human Activity Recognition with MLLM-based Visual Reinforcement Learning

Wenqi Zheng, Yutaka Arakawa

机构 * Graduate School(研究生院) Faculty of Information Science(信息科学系) Electrical Engineering(电气工程) Kyushu University(九州大学) JAPAN(日本)

专题命中 视觉推理 :MLLM(title,abstract);multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14805 2025-08-13 cs.CV cs.AI cs.CL cs.LG cs.MM 82%

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?

Yang Yao, Lingyu Li, Jiaxin Song, Chiyu Chen, Zhenqi He, Yixu Wang, Xin Wang, Tianle Gu, Jie Li, Yan Teng, Yingchun Wang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学) Shanghai Jiao Tong University(上海交通大学) The Hong Kong University of Science and Technology(香港科学与技术大学) Fudan University(复旦大学) Tsinghua University(清华大学)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03733 2025-08-07 cs.LG cs.AI cs.CL cs.CV 82%

CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning

Wenjie Li, Yujie Zhang, Haoran Sun, Yueqi Li, Fanrui Zhang, Mengzhe Xu, Victoria Borja Clausich, Sade Mellin, Renhao Yang, Chenrun Wang, Jethro Zih-Shuo Wang, Shiyi Yao, Gen Li, Yidong Xu, Hanyu Wang, Yilin Huang, Angela Lin Wang, Chen Shi, Yin Zhang, Jianan Guo, Luqi Yang, Renxuan Li, Yang Xu, Jiawei Liu, Yao Zhang, Lei Liu, Carlos Gutiérrez SanRomán, Lei Wang

机构 * College of Health Science and Technology, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院健康科学与技术学院) Shanghai Innovation Institute(上海创新研究院) Clinical Center for Sports Medicine, Department of Orthopaedics, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院骨科临床中心) School of Basic Medical Sciences, Intelligent Medicine Institute, Fudan University(复旦大学基础医学学院) Department of Hematology, The First Affiliated Hospital, College of Medicine, Zhejiang University(浙江大学医学院第一附属医院血液科) MoE Key Laboratory of Brain-Inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学脑启发智能感知与认知教育部重点实验室) Department of Public Health and Primary Care, University of Cambridge(剑桥大学公共卫生与初级保健学院) Department of Medicine, Faculty of Health Sciences, Universidad CEU Cardenal Herrera(CEU卡德纳尔-赫尔曼大学健康科学学院医学系) Faculty of Medicine, University of Helsinki(赫尔辛基大学医学院) X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院X-LANCE实验室) Department of Hepatobiliary Surgery, National Cancer Center / National Clinical Research Center for Cancer / Cancer Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College(中国医学科学院肿瘤医院肝胆外科) Department of Surgery, The Ohio State University Wexner Medical Center, The James Comprehensive Cancer Center(俄亥俄州立大学韦克斯纳医学中心外科部,詹姆斯综合癌症中心) Ningbo Institute of Technology, Beihang University(北航宁波理工学院)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02886 2025-08-06 cs.CL 82%

Coherent Multimodal Reasoning with Iterative Self-Evaluation for Vision-Language Models

Wenjie Luo, Ruocheng Li, Shanshan Zhu, Julian Perry

机构 * Fujian University of Technology(福建工程大学) Delta University for Science and Technology(科技大学)

专题命中 视觉推理 :vision-language model(title,abstract);LLaVA(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15266 2025-07-22 cs.RO cs.SY eess.SY 82%

VLM-UDMC: VLM-Enhanced Unified Decision-Making and Motion Control for Urban Autonomous Driving

Haichao Liu, Haoren Guo, Pei Liu, Benshan Ma, Yuxiang Zhang, Jun Ma, Tong Heng Lee

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) National University of Singapore(新加坡国立大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 视觉推理 :VLM(title,abstract);vision-language model(abstract)

Comments 14 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09535 2025-07-15 eess.SP 82%

Reframing SAR Target Recognition as Visual Reasoning: A Chain-of-Thought Dataset with Multimodal LLMs

Chaoran Li, Xingguo Xu, Siyuan Mu

专题命中 视觉推理 :visual reasoning(title,abstract);multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.00928 2025-07-08 cs.LG cs.AI cs.CV 82%

Learning Differentiable Logic Programs for Abstract Visual Reasoning

Hikaru Shindo, Viktor Pfanschilling, Devendra Singh Dhami, Kristian Kersting

专题命中 视觉推理 :visual reasoning(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments Published at Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07936 2025-06-10 cs.CV cs.AI cs.CL cs.LG 82%

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Chengyue Huang, Yuchen Zhu, Sichen Zhu, Jingyun Xiao, Moises Andrade, Shivang Chopra, Zsolt Kira

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17641 2025-06-10 cs.RO 82%

Scene Exploration by Vision-Language Models

Venkatesh Sripada, Samuel Carter, Frank Guerin, Amir Ghalamzan

机构 * University of Surrey(萨里大学)

专题命中 视觉推理 :vision-language model(title,abstract);VLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏