arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26117 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 4478 篇

2604.17233 2026-04-21 cs.CV cs.AI 82%

Enhancing Zero-shot Personalized Image Aesthetics Assessment with Profile-aware Multimodal LLM

通过基于资料的多模态大语言模型增强零样本个性化图像审美评估

Chun Wang, Chenfeng Wei, Chenyang Liu, Weihong Deng

机构 * Mashang Consumer Finance Co., Ltd.(马莎消费金融有限公司) Xi'an Jiaotong-Liverpool University(西安交通大学利物浦大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视觉推理 :MLLM(summary_cn,abstract);分类 cs.CV、cs.AI

AI总结 本文提出P-MLLM,通过用户资料引导的多模态大语言模型实现零样本个性化图像审美评估,实验表明其在无历史数据情况下表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17054 2026-04-21 cs.CV cs.AI 82%

mEOL: Training-Free Instruction-Guided Multimodal Embedder for Vector Graphics and Image Retrieval

mEOL:无需训练的指令引导多模态嵌入器用于矢量图形和图像检索

Kyeong Seon Kim, Baek Seong-Eun, Lee Jung-Mok, Tae-Hyun Oh

机构 * KAIST(韩国科学技术院) POSTECH

专题命中 视觉推理 :MLLM(abstract,abstract_cn);visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出无需训练的指令引导多模态嵌入框架,通过多模态大语言模型将文本、位图和SVG代码映射到对齐的嵌入空间,利用模态特定指令和结构化SVG提示实现嵌入方向控制,构建首个文本到SVG检索基准,展示出优于传统基线的性能。

Comments Round 1 early acceptance to WACV 2026, Project page: https://scene-the-ella.github.io/meol

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05497 2026-04-08 cs.AI cs.CV 82%

Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models

Thinking Diffusion: 通过惩罚和引导在扩散多模态语言模型中抑制和引导视觉推理

Keuntae Kim, Mingyu Kang, Yong Suk Choi

机构 * Department of Computer Science, Hanyang University(汉阳大学计算机科学系) Department of Artificial Intelligence, Hanyang University(汉阳大学人工智能系)

专题命中 视觉推理 :vision-language model(abstract);visual reasoning(abstract);grounding(abstract);multimodal large language model(abstract)

AI总结 本文针对扩散多模态语言模型在视觉推理中生成过早答案的问题,提出位置与步数惩罚和视觉推理引导方法,提升推理准确率并加快推理速度。

Comments CVPR 2026 - main

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17474 2026-08-17 eess.AS cs.SD 82%

Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models

基于商业自动语音识别和多模态大语言模型的口吃语音零样本识别

Ali Alsayegh, Tariq Masood

专题命中 视觉推理 :multimodal large language model(title,abstract);MLLM(abstract)

AI总结 本研究评估了商业ASR和MLLM在口吃语音识别中的性能,发现GPT-4o在重度口吃情况下显著降低WER,而Gemini变体表现下降,为辅助语音接口技术选择提供实证依据。

Journal ref International Journal of Intelligent Systems, 2026; 2026:6065038

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12939 2026-07-13 cs.RO 版本更新 82%

RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics

RoboStream: 用记忆融合时空推理提升视觉语言模型在机器人中的应用

Yuzhi Huang, Jie Wu, Weijue Bu, Ziyi Xiong, Gaoyang Jiang, Ye Li, Kangye Ji, Shuzhao Xie, Yue Huang, Chenglei Wu, Jingyan Jiang, Zhi Wang

机构 * Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院,清华大学) YXGN Robotics(YXGN机器人) China University of Mining and Technology(中国矿业大学) Shenzhen Technology University(深圳技术大学) Huazhong University of Science and Technology(华中科技大学) Xiamen University(厦门大学)

专题命中 视觉推理 :vision-language model(title);VLM(abstract);grounding(abstract)

AI总结 RoboStream通过时空融合令牌和因果时空图实现持久记忆,解决机器人长时间任务中的几何锚定和状态追踪问题,提升长期任务表现。

Comments Accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07251 2026-07-09 cs.CL 新提交 82%

Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models

视觉语言模型中使用空间指示表达式的多语言能力评估

Kaito Watanabe, Taisei Yamamoto, Tomoki Doi, Hitomi Yanaka

机构 * The University of Tokyo(东京大学) Riken(理化学研究所) Tohoku University(东北大学)

专题命中 视觉推理 :vision-language model(title,abstract);grounding(abstract)

AI总结 研究聚焦视觉语言模型的空间推理能力,通过开发基准评估其使用四种语言空间指示表达式的多语言能力,实验发现测试模型使用指示词方式与人类不同,特别是在依距离选指示词方面。

Comments Accepted to ACL SRW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12169 2026-06-11 cs.CV cs.AI cs.CL cs.LG 新提交 82%

OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

OpenMedReason: 医学视觉语言模型的科学推理监督

Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi

机构 * York University(约克大学) Vector Institute(向量研究所) University of British Columbia(不列颠哥伦比亚大学) University of Toronto(多伦多大学) Unity Health Toronto / St. Michael’s Hospital(多伦多联合健康/圣迈克尔医院) University Health Network(大学健康网络) Arc Institute(弧研究所) Queen's University(女王大学)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 提出OpenMedReason,一个包含约45万图像-问题-答案实例的大规模开放医学推理语料库,其推理轨迹主要来自生物医学科学文章,并配套基准OpenMedReason-Bench进行细粒度评估,在监督微调和强化对齐中有效提升模型性能。

Comments 42 pages, 9 figures, 24 tables. Dataset and code: https://huggingface.co/datasets/neginb/OpenMedReason

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22766 2026-06-09 cs.CL 版本更新 82%

Imagination Helps Visual Reasoning, But Not Yet in Latent Space

想象力有助于视觉推理,但尚未在潜在空间中实现

You Li, Chi Chen, Yanghao Li, Fanhu Zeng, Kaiyu Huang, Jinan Xu, Maosong Sun

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 视觉推理 :visual reasoning(title,abstract);multimodal large language model(abstract)

AI总结 通过因果中介分析发现,多模态大语言模型中的潜在推理存在输入-潜在和潜在-答案两个关键断连,表明其有效性有限,并提出显式想象方法CapImagine,在视觉推理任务中表现更优。

Comments ICML 2026 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05744 2026-06-05 cs.CL 82%

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models

PlanBench-V: 面向视觉语言模型的空间规划地图基准

Minxin Chen, He Zhu, Junyou Su, Wen Wang, Yijie Deng, Wenjia Zhang

机构 * Behavioral and Spatial AI Lab(行为与空间人工智能实验室) Tongji University(同济大学) Peking University(北京大学) College of Architecture and Urban Planning(建筑与城市规划学院)

专题命中 视觉推理 :vision-language model(title,abstract);VLM(abstract_cn)

AI总结 为评估视觉语言模型在空间规划地图解读中的能力,构建了专家标注数据集SPMD,并提出基于感知、推理、关联、实施四阶段认知框架的基准PlanBench-V,实验表明当前模型在实施类任务上存在显著局限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08726 2026-04-13 cs.RO 82%

Task-Aware Bimanual Affordance Prediction via VLM-Guided Semantic-Geometric Reasoning

基于VLM引导的语义-几何推理的任务感知双臂 affordance 预测

Fabian Hahne, Vignesh Prasad, Georgia Chalvatzaki, Jan Peters, Alap Kshirsagar

机构 * Department of Computer Science, Technical University of Darmstadt(达姆施塔特工业大学计算机科学系) German Research Center for AI (DFKI)(德国人工智能研究中心) Centre for Cognitive Science, Technical University of Darmstadt(达姆施塔特工业大学认知科学中心) Hessian Center for Artificial Intelligence (Hessian.AI), Darmstadt(黑森州人工智能中心) Robotics Institute Germany (RIG)(德国机器人研究所)

专题命中 视觉推理 :VLM(title,abstract);vision-language model(abstract)

AI总结 本文提出一种任务感知双臂 affordance 预测框架,通过VLM实现跨物体类别和任务描述的泛化,融合多视角RGB-D数据生成全局6自由度抓取候选,结合语义和几何过滤提升任务成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01547 2026-03-24 cs.CV cs.AI cs.LG 82%

Vision-language models lag human performance on physical dynamics and intent reasoning

视觉-语言模型在物理动态和意图推理上仍逊于人类表现

Tianjun Gu, Jingyu Gong, Zhizhong Zhang, Yuan Xie, Lizhuang Ma, Xin Tan, Athanasios V

机构 * East China Normal University(东华大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) University of Agder(阿格德大学)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出Teleo-Spatial Intelligence(TSI)以连接时空变化与目标导向结构,通过EscherVerse大规模开放世界资源评估,发现视觉-语言模型在开放世界中仍无法达到人类水平的意图推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14786 2026-03-17 cs.RO 82%

CORAL: COntextual Reasoning And Local Planning in A Hierarchical VLM Framework for Underwater Monitoring

CORAL:在分层视觉语言框架中实现上下文推理与局部规划用于水下监测

Zhenqi Wu, Yuanjie Lu, Xuesu Xiao, Xiaomin Lin

专题命中 视觉推理 :VLM(title,abstract);vision-language model(abstract)

AI总结 CORAL框架通过分离高层语义推理与底层反应控制,提升水下监测的覆盖率和安全性,减少碰撞并降低对VLM的调用次数。

Comments Submitted to IROS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11414 2026-03-13 cs.CL cond-mat.mtrl-sci 82%

MaterialFigBENCH: benchmark dataset with figures for evaluating college-level materials science problem-solving abilities of multimodal large language models

MaterialFigBENCH:用于评估多模态大语言模型在大学级材料科学问题解决能力的基准数据集

Michiko Yoshitake, Yuta Suzuki, Ryo Igarashi, Yoshitaka Ushiku, Keisuke Nagato

专题命中 视觉推理 :multimodal large language model(title,abstract);visual reasoning(abstract)

AI总结 MaterialFigBench是一个评估多模态大语言模型在材料科学问题中图表解读能力的基准数据集,通过137个问题测试模型在视觉理解和数值精度上的表现,揭示了当前模型在图表处理方面的不足。

Comments 27 pages, 4 tables, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05017 2026-03-06 cs.RO 82%

Direct Contact-Tolerant Motion Planning With Vision Language Models

直接接触容忍的运动规划与视觉语言模型

He Li, Jian Sun, Chengyang Li, Guoliang Li, Qiyu Ruan, Shuai Wang, Chengzhong Xu

机构 * State Key Laboratory of Internet of Things for Smart City (SKL-IOTSC), University of Macau(物联网智能城市国家重点实验室,澳门大学) Shenzhen Institutes of Advanced Technology (SIAT), Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Department of Electrical and Computer Engineering, The University of Hong Kong(香港大学电子与计算机工程系)

专题命中 视觉推理 :vision language model(title);vision-language model(abstract);VLM(abstract)

AI总结 本文提出直接接触容忍运动规划方法,利用视觉语言模型实现接触感知导航,提升机器人在拥挤环境中的鲁棒性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25191 2026-03-05 cs.RO 82%

SoraNav: Adaptive UAV Task-Centric Navigation via Zeroshot VLM Reasoning

SoraNav: 通过零样本VLM推理实现适应性无人机任务导向导航

Hongyu Song, Rishabh Dev Yadav, Cheng Guo, Wei Pan

机构 * Department of Computer Science, The University of Manchester(计算机科学系,曼彻斯特大学)

专题命中 视觉推理 :VLM(title,abstract);vision-language model(abstract)

AI总结 SoraNav通过零样本VLM推理实现无人机任务导向导航,解决空间-语义差距问题,提升导航成功率与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21628 2026-03-04 cs.CL 82%

RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoning

RuCL:基于分层评分标准的课程学习用于多模态大语言模型推理

Yukun Chen, Jiaming Li, Longze Chen, Ze Gong, Jingpeng Li, Zhen Qin, Hengyu Chang, Ancheng Xu, Zhihao Yang, Hamid Alinejad-Rokny, Qiang Qu, Bo Zheng, Min Yang

机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) University of Chinese Academy of Sciences(中国科学院大学) Alibaba Group(阿里巴巴集团) School of Biomedical Engineering, UNSW Sydney(新南威尔士大学生物医学工程学院)

专题命中 视觉推理 :multimodal large language model(title,abstract);visual reasoning(abstract)

AI总结 RuCL通过分层评分标准课程学习提升多模态大语言模型的推理能力,实现7.83%的准确率提升。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06749 2026-03-03 cs.CV cs.AI cs.CL cs.LG 82%

Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Vision-R1: 促进多模态大语言模型推理能力的激励方法

Wenxuan Huang, Bohan Jia, Zijie Zhai, Shaosheng Cao, Zheyu Ye, Fei Zhao, Zhe Xu, Xu Tang, Yao Hu, Shaohui Lin

机构 * East China Normal University(华东师范大学) The Chinese University of Hong Kong(香港中文大学) Xiaohongshu Inc.(小红书公司)

专题命中 视觉推理 :multimodal large language model(title);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 Vision-R1通过构建高质量多模态CoT数据集和渐进性思维抑制训练策略,提升多模态推理能力,在数学推理基准中取得显著成绩。

Comments Accepted to ICLR 2026. Code is available at https://github.com/Osilly/Vision-R1

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20330 2026-02-25 cs.CV cs.AI cs.LG 82%

Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking

视觉-语言模型中的电路追踪:理解多模态思维的内部机制

Jingcheng Yang, Tianhu Xiong, Shengyi Qian, Klara Nahrstedt, Mingyuan Wu

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出首个透明电路追踪框架,用于分析视觉-语言模型的多模态推理机制,揭示不同视觉特征电路在数学推理和跨模态关联中的作用。

Comments To appear in the Findings of CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16157 2026-02-19 cs.HC 82%

Peeking Ahead of the Field Study: Exploring VLM Personas as Support Tools for Embodied Studies in HCI

提前窥视田野研究:探索VLM人设作为人机交互领域具身研究的支持工具

Xinyue Gui, Ding Xia, Mark Colley, Yuan Li, Vishal Chauhan, Anubhav Anubhav, Zhongyi Zhou, Ehsan Javanmardi, Stela Hanbyeol Seo, Chia-Ming Chang, Manabu Tsukada, Takeo Igarashi

专题命中 视觉推理 :VLM(title,abstract);vision-language model(abstract)

AI总结 本文提出利用VLM人设模拟田野研究结果,验证其在具身研究中的应用潜力,发现其在响应模式上接近人类但缺乏多样性。

Comments Accepted to CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15237 2026-02-18 cs.HC 82%

Ground-Truth Depth in Vision Language Models: Spatial Context Understanding in Conversational AI for XR-Robotic Support in Emergency First Response

视觉语言模型中的真实深度:用于虚拟现实机器人支持的紧急急救中的空间情境理解

Rodrigo Gutierrez Maquilon, Marita Hueber, Georg Regal, Manfred Tscheligi

专题命中 视觉推理 :vision language model(title,abstract);VLM(abstract)

AI总结 本文提出了一种结合深度传感与视觉语言模型的原型,用于增强紧急急救中的空间情境理解,通过深度增强提高准确性和稳定性,同时降低工作量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17587 2026-02-03 cs.CV cs.AI cs.LG 82%

HalluRNN: Mitigating Hallucinations via Recurrent Cross-Layer Reasoning in Large Vision-Language Models

HalluRNN: 通过大规模视觉-语言模型中的递归跨层推理缓解幻觉

Le Yu, Kaishen Wang, Jianlong Xiong, Yue Cao, Lei Zhang, Zhang Yi Tao He

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 HalluRNN通过递归跨层推理模块缓解大规模视觉-语言模型中的幻觉问题,提升模型稳定性和性能。

Comments 6 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00211 2026-02-03 cs.CV cs.AI cs.LG 82%

Interpretable Unsupervised Deformable Image Registration via Confidence-bound Multi-Hop Visual Reasoning

可解释的无监督变形图像配准:通过置信度界限多跳视觉推理

Zafar Iqbal, Anwar Ul Haq, Srimannarayana Grandhi

机构 * School of Engineering and Technology(工程与技术学院)

专题命中 视觉推理 :visual reasoning(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出了一种基于多跳视觉推理的可解释无监督变形图像配准框架,通过局部空间细化和跨参考注意力机制,实现高精度且透明的医学图像配准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22714 2026-02-02 cs.LG cs.AI cs.CV 82%

Vision-Language Models Unlock Task-Centric Latent Actions

视觉-语言模型解锁任务导向的潜在动作

Alexander Nikulin, Ilya Zisman, Albina Klepach, Denis Tarasov, Alexander Derevyagin, Andrei Polubarov, Lyubaykin Nikita, Vladislav Kurenkov

机构 * Innopolis University(因诺波利斯大学) Research Center for Trusted Artificial Intelligence, ISP RAS(可信人工智能研究所以及ISP俄罗斯科学院)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出利用视觉-语言模型的常识推理能力,通过可提示的表示分离可控变化与噪声,提升潜在动作质量,从而提高下游任务性能。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11464 2025-12-15 cs.CV cs.AI cs.LG 82%

Exploring MLLM-Diffusion Information Transfer with MetaCanvas

探索MLLM-扩散信息传输与MetaCanvas

Han Lin, Xichen Pan, Ziqi Huang, Ji Hou, Jialiang Wang, Weifeng Chen, Zecheng He, Felix Juefei-Xu, Junzhe Sun, Zhipeng Fan, Ali Thabet, Mohit Bansal, Chu Wang

机构 * Meta Superintelligence Labs(Meta 超智能实验室) New York University(纽约大学) Nanyang Technological University(南洋理工大学)

专题命中 视觉推理 :MLLM(title);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 MetaCanvas通过使MLLMs在潜在空间中推理和规划,提升多模态生成的精确度和结构化控制。

Comments Project page: https://metacanvas.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09349 2025-12-11 cs.RO 82%

COVLM-RL: Critical Object-Oriented Reasoning for Autonomous Driving Using VLM-Guided Reinforcement Learning

COVLM-RL:基于VLM引导强化学习的自动驾驶中关键对象导向推理

Lin Li, Yuxin Cai, Jianwu Fang, Jianru Xue, Chen Lv

机构 * School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与航空航天工程学院) National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, National Engineering Research Center for Visual Information and Applications, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学人机混合增强智能国家重点实验室、视觉信息与应用国家工程研究中心、人工智能与机器人研究院)

专题命中 视觉推理 :VLM(title,abstract);vision-language model(abstract)

AI总结 COVLM-RL通过结合关键对象导向推理与VLM引导强化学习,提升了自动驾驶系统的泛化能力和训练效率。

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07177 2025-12-09 cs.RO 82%

Using Vision-Language Models as Proxies for Social Intelligence in Human-Robot Interaction

利用视觉-语言模型作为社交智能代理在人机交互中的应用

Fanjun Bu, Melina Tsai, Audrey Tjokro, Tapomayukh Bhattacharjee, Jorge Ortiz, Wendy Ju

机构 * Cornell University, Cornell Tech(康奈尔大学,康奈尔科技) Cornell University(康奈尔大学) Rutgers University(罗格斯大学) Cornell Tech, Jacobs Technion Cornell Institute(康奈尔科技,雅各布斯技术学院康奈尔研究所)

专题命中 视觉推理 :vision-language model(title,abstract);VLM(abstract)

AI总结 本文提出利用视觉-语言模型作为代理,通过轻量级感知检测器触发,实现机器人在人机交互中的社交响应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20531 2025-11-26 cs.AI cs.CV cs.LG 82%

Beyond Generation: Multi-Hop Reasoning for Factual Accuracy in Vision-Language Models

超越生成:面向视觉语言模型事实准确性的多跳推理

Shamima Hossain

机构 * Department of Computer Science, Brac University(布鲁尔大学计算机科学系) bKash Limited(bKash公司)

专题命中 视觉推理 :vision-language model(title);visual language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出一种基于知识图谱的多跳推理框架,提升视觉语言模型在事实准确性上的表现,通过多步骤推理增强模型的逻辑推理能力。

Comments Accepted as poster at NewInML Workshop ICML, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12008 2025-11-18 cs.AI cs.CV cs.LG 82%

Adaptive Diagnostic Reasoning Framework for Pathology with Multimodal Large Language Models

Yunqi Hong, Johnson Kao, Liam Edwards, Nein-Tzu Liu, Chung-Yen Huang, Alex Oliveira-Kowaleski, Cho-Jui Hsieh, Neil Y. C. Lin

机构 * Computer Science Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校计算机科学系) Mechanical and Aerospace Engineering Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校机械与航空航天工程系) Department of Pathology, Tri-Service General Hospital, National Defense Medical Center, Taipei, Taiwan(台湾国防医学院三军总医院病理部) Department of Pathology, National Taiwan University Hospital, Taipei, Taiwan(台湾国立台湾大学医院病理部) Department of Pathology, David Geffen School of Medicine, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校大卫·Geffen医学院病理部) Bioengineering Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校生物工程系) Institute for Quantitative and Computational Biosciences, University of California, CA, USA(加州大学定量与计算生物科学研究所)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11831 2025-11-18 cs.AI cs.CV cs.LG 82%

TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models

Wenhao Zhou, Hao Zheng, Rong Zhao

机构 * Center for Brain-Inspired Computing Research (CBICR)(脑启发计算研究中心) Department of Precision Instruments(精密仪器系) IDG/McGovern Institute for Brain Research(IDG/麦戈文脑研究学院)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21572 2025-11-14 cs.CL 82%

Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling

Shengwu. Xiong, Tianyu. Zou, Cong. Wang, Xuelong Li

机构 * Interdisciplinary Artificial Intelligence Research Institute, Wuhan College(交叉学科人工智能研究 institute,武汉学院) School of Computer and Artificial Intelligence, Wuhan University of Technology(计算机与人工智能学院,武汉理工大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Sanya Science and Education Innovation Park, Wuhan University of Technology(三亚科学教育创新园,武汉理工大学) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) School of Mathematics and Statistics, Northwestern Polytechnical University(数学与统计学院,西北工业大学) Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究所(TeleAI))

专题命中 视觉推理 :MLLM(title,abstract);multimodal large language model(abstract)

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏