arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-04-24 至 2026-04-24 共收录 7 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7 篇

2604.21134 2026-04-24 cs.CL 91%

Beyond Pixels: Introspective and Interactive Grounding for Visualization Agents

超越像素:可视化代理的 introspective 和 interactive 地基

Yiyang Lu, Woong Shin, Ahmad Maroof Karimi, Feiyi Wang, Jie Ren, Evgenia Smirni

机构 * William & Mary(威廉玛丽学院) Oak Ridge National Laboratory(橡树岭国家实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(summary_cn,abstract);vision-language model(abstract,abstract_cn)

AI总结 本文提出IVG框架,结合规范引导的反思与视图引导的交互,解决VLM在可视化任务中的误读与歧义问题,通过iPlotBench验证,提升问答准确度至0.81,并展示在自主探索与实时协作中的能力。

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15594 2026-04-24 cs.LG cs.CV 62%

Analytical Softmax Temperature Setting from Feature Dimensions for Model- and Domain-Robust Classification

基于特征维度的分析softmax温度设置用于模型和领域鲁棒分类

Tatsuhito Hasegawa, Shunsuke Sakai

机构 * Graduate School of Engineering, University of Fukui(福井大学工学研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

AI总结 本文提出基于特征维度确定softmax温度的理论方法,通过优化温度系数和批量归一化层,实现无训练的温度设置,提升分类鲁棒性。

Comments 22 pages, 11 figures, under review

Journal ref Neural Computing and Applications 37, 27985-28016, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21144 2026-04-24 cs.CL cs.AI cs.HC 57%

Using Machine Mental Imagery for Representing Common Ground in Situated Dialogue

利用机器心智 imagery 代表情境对话中的共同背景

Biswesh Mohapatra, Giovanni Duca, Laurent Romary, Justine Cassell

机构 * Inria University of Trento(特伦托大学) Inria, Carnegie Mellon University(Inria 和卡内基梅隆大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出一种主动视觉支架框架,通过将对话状态转化为持久的视觉历史,减少代表模糊并提高对话理解。

Comments Work under review. Biswesh Mohapatra and Giovanni Duca both contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20860 2026-04-24 cs.IR cs.AI 57%

RealRoute: Dynamic Query Routing System via Retrieve-then-Verify Paradigm

RealRoute:通过检索-验证范式实现动态查询路由系统

Jiahe Liu, Qinkai Yu, Jingcheng Niu, Xi Zhu, Zirui He, Zhen Xiang, Fan Yang, Jinman Zhao

机构 * Technical University of Denmark(技术大学) University of Exeter(埃克塞特大学) University of Toronto(多伦多大学) Rutgers University(罗格斯大学) NJIT(新 jersey 工业技术学院) University of Georgia(佐治亚大学) Wake Forest University(威克森林大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 RealRoute引入一种鲁棒的检索-验证机制,通过并行源无关检索和动态验证器确保证据完整性,优于预测基线,在多跳RAG推理任务中表现更佳。

Comments 12 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20848 2026-04-24 cs.IR cs.AI 57%

MATRAG: Multi-Agent Transparent Retrieval-Augmented Generation for Explainable Recommendations

MATRAG:多智能体透明检索增强生成用于可解释推荐

Sushant Mehta

机构 * ACM

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 MATRAG通过多智能体协作与知识图谱增强检索,提升推荐系统的透明性和可解释性,实验表明其在准确率和可信度上均优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20948 2026-04-24 cs.HC 50%

Can Virtual Agents Care? Designing an Empathetic and Personalized LLM-Driven Conversational Agent

虚拟代理能关心吗?设计一个具有同理心和个性化的LLM驱动对话代理

Truong Le Minh Toan, Dieu Bang Mach, Tan Duy Le, Nguyen Tan Viet Tuyen

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出一种框架,通过检索增强架构、结构化记忆和多模态交互,设计出具有同理心和个性化支持的虚拟代理,以提升心理健康支持的质量和可靠性。

Comments Accepted manuscript version to be presented at the SCI-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20865 2026-04-24 cs.CY 50%

Advances in Art: Orthogonal Disruption and the Beauty in Schematics

艺术进展:正交颠覆与图示之美

Sergio Alvarez-Telena, Marta Diez-Fernandez

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出正交艺术,一种对人工智能的辩证回应而非其服务的艺术学科。通过技术图示作为主要媒介,探索生成与概念空间的新型艺术实践,促进人文学科与艺术、工程、哲学交叉领域的跨学科素养。

Comments 10 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏