arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-04-17 至 2026-04-17 共收录 17 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 17 篇

2604.14656 2026-04-17 cs.AI cs.CL cs.CV 85%

Rethinking Patient Education as Multi-turn Multi-modal Interaction

重新思考患者教育作为多轮多模态交互

Zonghai Yao, Zhipeng Tang, Chengtao Lin, Xiong Luo, Benlu Wang, Juncheng Huang, Chin Siang Ong, Hong Yu

机构 * VA Bedford Health Care(VA贝德福德医疗中心) UMass Amherst(马萨诸塞大学阿默斯特分校) UMass Lowell(马萨诸塞大学洛厄尔分校) Yale University(耶鲁大学) National University of Singapore(新加坡国立大学) Yale School of Medicine(耶鲁医学院)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出MedImageEdu基准,通过多轮多模态交互提升患者教育效果,评估咨询过程和最终响应质量,发现多模态模型在视觉 grounding、安全性和情绪互动方面存在不足。

Comments Equal contribution for the first two authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06559 2026-04-17 cs.CV 83%

ArrowGEV: Grounding Events in Video via Learning the Arrow of Time

ArrowGEV:通过学习时间之箭来视频中接地事件

Fangxu Yu, Ziyao Lu, Liqiang Niu, Fandong Meng, Jie Zhou

机构 * School of Artificial Intelligence, Nanjing University, China(南京大学人工智能学院,中国) Pattern Recognition Center, WeChat AI, Tencent Inc, China(腾讯人工智能研究院,中国)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision language model(abstract);分类 cs.CV

AI总结 本文提出ArrowGEV框架,通过学习时间方向性提升视频中事件接地与时间理解能力,通过区分时间敏感和时间不敏感事件增强模型鲁棒性与泛化能力。

Comments Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15602 2026-04-17 cs.CV 83%

TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies?

TennisTV: 多模态大语言模型能否理解网球发球?

Zhongyuan Bao, Lejun Zhang

机构 * New York University(纽约大学) Tandon School of Engineering(坦顿工程学院) Shanghai, China(上海,中国) New York, USA(纽约,美国)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV

AI总结 TennisTV是首个全面评估网球视频理解的基准,通过模拟发球过程中的连续击球事件,评估17种大语言模型在网球视频理解中的表现,发现帧采样密度需平衡且需提升时间定位能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03307 2026-04-17 cs.CV cs.AI 82%

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators

V-Reflection:将多模态大语言模型从被动观察者转变为主动提问者

Jiazhou Zhou, Yucheng Chen, Hongyang Li, Qing Jiang, Hu Zhou, Ying-Cong Chen, Lei Zhang

机构 * AI Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)人工智能方向) International Digital Economy Academy(国际数字经济学院) MedVisAI Lab, Lee Kong Chian School of Medicine, Nanyang Technological University, and Centre of AI in Medicine(医学视觉人工智能实验室,南洋理工大学Lee Kong Chian医学院,以及人工智能在医学中的中心) South China University of Technology(华南理工大学) Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(香港理工大学电子与电气工程系)

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出V-Reflection框架,通过'思考后再观察'的视觉反思机制,使多模态大语言模型能主动提问视觉特征空间,提升细粒度任务的感知能力。

Comments Main paper 14 pages with supplementary 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04585 2026-04-17 cs.CV 81%

SAM3-I: Segment Anything with Instructions

SAM3-I: 基于指令的分割

Jingjing Li, Yue Feng, Yuchen Guo, Jincai Huang, Wei Ji, Qi Bi, Yongri Piao, Miao Zhang, Xiaoqi Zhao, Qiang Chen, Shihao Zou, Huchuan Lu, Li Cheng

机构 * University of Alberta(阿尔伯塔大学) Tencent WeChat(腾讯微信) Northwestern University(西北大学) SUSTech(四川大学) Yale University(耶鲁大学) Utrecht University(乌得勒支大学) Dalian University of Technology(大连理工大学) SIAT, Chinese Academy of Sciences(中科院深圳先进技术研究院)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);分类 cs.CV

AI总结 SAM3-I通过整合概念级 grounding 和指令级推理,提升分割性能,支持复杂自然语言指令。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14846 2026-04-17 cs.CV cs.AI 79%

Zero-Shot Retail Theft Detection via Orchestrated Vision Models: A Model-Agnostic, Cost-Effective Alternative to Trained Single-Model Systems

通过协调视觉模型实现零样本零售盗窃检测:一种无需训练单个模型的低成本替代方案

Haileab Yagersew

机构 * Paza AI

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出Paza框架,通过协调多个现有模型实现零样本零售盗窃检测,降低训练成本,利用多信号预过滤减少昂贵视觉语言模型调用次数,提升效率和准确性。

Comments 16 pages, 3 figures, Code to be released at https://github.com/xHaileab/Paza-AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15946 2026-04-17 cs.LG cs.AI cs.CR 79%

Fall into a Pit, Gain in a Wit: Cognitive-Guided Harmful Meme Detection via Misjudgment Risk Pattern Retrieval

跌入陷阱,获得智慧:通过误判风险模式检索的认知引导有害迷因检测

Wenshuo Wang, Ziyou Jiang, Junjie Wang, Mingyang Li, Jie Huang, Yuekai Huang, Zhiyuan Chang, Feiyan Duan, Qing Wang

机构 * State Key Laboratory of Complex System Modeling and Simulation Technology(复杂系统建模与仿真技术国家重点实验室) Science and Technology on Integrated Information System Laboratory(集成信息系统技术研究所) Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出PatMD方法,通过学习并主动缓解潜在误判风险,识别有害迷因的深层误判风险模式,提升多模态大语言模型的检测能力,实验显示在5项有害检测任务中,F1-score和准确率均显著提升。

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04567 2026-04-17 cs.CV 77%

All Changes May Have Invariant Principles: Improving Ever-Shifting Harmful Meme Detection via Design Concept Reproduction

所有变化可能都有不变原则:通过设计概念再现改进永动有害迷因检测

Ziyou Jiang, Mingyang Li, Junjie Wang, Yuekai Huang, Jie Huang, Zhiyuan Chang, Zhaoyang Li, Qing Wang

机构 * State Key Laboratory of Complex System Modeling and Simulation Technology(复杂系统建模与仿真技术国家重点实验室) Science and Technology on Integrated Information System Laboratory Institute of Software Chinese Academy of Sciences(软件研究所信息集成系统技术研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出RepMD方法,通过设计概念再现检测不断变化的有害迷因,利用攻击树定义设计概念图,并指导多模态大语言模型提高检测准确率。

Comments 19 pages, 11 figures, 9 tables accepted by ACL 2026 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15225 2026-04-17 cs.HC 67%

UrbanClipAtlas: A Visual Analytics Framework for Event and Scene Retrieval in Urban Videos

UrbanClipAtlas:面向城市视频中事件和场景检索的视觉分析框架

Joel Perca, Luis Sante, Juanpablo Heredia, Joao Rulff, Claudio Silva, Jorge Poco

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract)

AI总结 本文提出UrbanClipAtlas框架,结合RAG、实体提取和视频定位技术,实现城市视频中事件和场景的高效检索与解释,通过知识图谱和增强聊天界面提升分析效率。

Comments 12 pages and 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15134 2026-04-17 cs.CV 57%

How to Correctly Make Mistakes: A Framework for Constructing and Benchmarking Mistake Aware Egocentric Procedural Videos

如何正确犯错:一个用于构建和基准测试错误感知第一人称视频的框架

Olga Loginova, Frank Keller

机构 * University of Trento(特伦托大学) University of Edinburgh(爱丁堡大学) Italy(意大利) United Kingdom(大不列颠岛)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本文提出PIE-V框架,通过引入可控的人类合理偏差来构建和评估错误感知的第一人称视频,结合心理学启发的错误计划、纠正计划、LLM作家和评委,提升视频的合理性和一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14141 2026-04-17 cs.CV 57%

Geometric Context Transformer for Streaming 3D Reconstruction

流式3D重建的几何上下文变换器

Lin-Zhuo Chen, Jian Gao, Yihang Chen, Ka Leong Cheng, Yipengjing Sun, Liangxiao Hu, Nan Xue, Xing Zhu, Yujun Shen, Yao Yao, Yinghao Xu

机构 * technology.robbyant.com(robbyant技术公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本文提出LingBot-Map,基于几何上下文变换器架构,通过精心设计的注意力机制实现流式3D重建,兼顾几何精度、时间一致性和计算效率,达到20FPS的稳定高效推理。

Comments Project page: https://technology.robbyant.com/lingbot-map Code: https://github.com/robbyant/lingbot-map

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14624 2026-04-17 cs.SE cs.AI 57%

Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks

询问何为关键:面向软件工程任务的奖励驱动澄清

Sanidhya Vijayvargiya, Vijay Viswanathan, Graham Neubig

机构 * Language Technologies Institute, Carnegie Mellon University, Pittsburgh, USA(语言技术研究所,卡内基梅隆大学,匹兹堡,美国)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文研究软件工程任务中有效澄清问题的方法,通过量化影响任务成功的信息类型和用户能提供的信息,提出基于Shapley属性和分布比较的多阶段强化学习奖励机制,训练出CLARITI模块,在未明确问题上表现优于GPT-5,生成更少问题。

Comments 28 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14576 2026-04-17 cs.AI 57%

Enhancing Mental Health Counseling Support in Bangladesh using Culturally-Grounded Knowledge

在孟加拉国增强心理健康咨询支持:基于文化根基的知识

Md Arid Hasan, Azhagu Meena SP, Aditya Khan, Abu Md Akteruzzaman Bhuiyan, Helal Uddin Ahmed, Joysree Debi, Farig Sadeque, Annie En-Shiun Lee, Syed Ishtiaque Ahmed

机构 * University of Toronto(多伦多大学) mPower Social Enterprises Ltd(mPower社会企业有限公司) National Institute of Mental Health(心理健康国家研究所) Sajida Foundation(萨吉达基金会) BRAC University(BRAC大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文探讨如何通过系统整合临床验证的知识提升LLM在心理健康咨询中的表现,比较了检索增强生成与知识图谱方法,证明结构化专业知识能提升咨询相关性与实用性。

Comments submitted to CLPsych 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14500 2026-04-17 cs.AI 57%

Geometric Metrics for MoE Specialization: From Fisher Information to Early Failure Detection

几何度量用于MoE专业化:从Fisher信息到早期故障检测

Dongxin Guo, Jikun Wu, Siu Ming Yiu

机构 * The University of Hong Kong(香港大学) Brain Investing Limited

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出基于信息几何的框架,通过Fisher信息度量分析MoE专业化动态,提出FSI和FHS指标,有效预测训练失败并提升模型性能。

Comments 6 pages, 2 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14386 2026-04-17 cs.GT cs.AI 57%

Coalition Formation in LLM Agent Networks: Stability Analysis and Convergence Guarantees

在LLM代理网络中的联盟形成:稳定性分析与收敛保证

Dongxin Guo, Jikun Wu, Siu-Ming Yiu

机构 * Department of Compute Science, The University of Hong Kong(香港大学计算机科学系) Brain Investing Limited(Brain Investing有限公司) Stellaris AI Limited(Stellaris AI有限公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出首个基于享乐博弈理论的LLM联盟形成框架,分析了LLM代理的有限理性特性,并通过实验验证了CoalT协议在联盟稳定性上的优越性。

Comments 15 pages including supplementary material, 2 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14261 2026-04-17 cs.CL cs.AI 57%

ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated Agents

ReviewGrounder:通过规则引导和工具集成代理提升评审的实质内容

Zhuofeng Li, Yi Lu, Dongfu Jiang, Haoxiang Zhang, Yuyang Bai, Chuan Li, Yu Wang, Shuiwang Ji, Jianwen Xie, Yu Zhang

机构 * Texas A&M University(德克萨斯A&M大学) University of Waterloo(滑铁卢大学) UC San Diego(南加州大学) Lambda University of Oregon(俄勒冈大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出ReviewGrounder框架,通过规则引导和工具集成提升AI会议评审的实质内容,实验表明其在8个维度上优于基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14222 2026-04-17 cs.IR cs.AI 57%

Adaptive Query Routing: A Tier-Based Framework for Hybrid Retrieval Across Financial, Legal, and Medical Documents

自适应查询路由:一种分层框架,用于金融、法律和医疗文档的混合检索

Afshan Hashmi

机构 * TRDC, Tuwaiq Academy(TRDC,图瓦伊克学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出了一种分层框架,通过比较三种检索架构在金融、法律和医疗领域的表现,揭示了不同检索方法在不同复杂度层级上的性能差异,强调了树状推理和混合检索在跨参考和多段查询中的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏