arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-02-27 至 2026-02-27 共收录 14 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 14 篇

2602.22727 2026-02-27 cs.CV 83%

HulluEdit: Single-Pass Evidence-Consistent Subspace Editing for Mitigating Hallucinations in Large Vision-Language Models

HulluEdit: 单次通过证据一致子空间编辑用于缓解大视觉-语言模型中的幻觉

Yangguang Lin, Quan Fang, Yufei Li, Jiachen Sun, Junyu Gao, Jitao Sang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beijing Jiaotong University(北京交通大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV

AI总结 HulluEdit通过单次通过的正交子空间编辑方法,有效缓解大视觉-语言模型中的幻觉问题,实现了最先进的幻觉减少效果。

Comments accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06008 2026-02-27 cs.CV cs.AI 81%

Detection and Measurement of Hailstones with Multimodal Large Language Models

利用多模态大语言模型检测和测量冰雹

Moritz Alker, David C. Schedl, Andreas Stöckl

机构 * Digital Media Lab University of Applied Sciences Upper Austria(数字媒体实验室 上奥地利应用科学大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV、cs.AI

AI总结 利用多模态大语言模型检测和测量冰雹,通过社交媒体图像实现快速评估

Comments 6 pages, 5 figures, accepted at The 2nd International Conference on Electrical and Computer Engineering Researches

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18207 2026-02-27 cs.CV cs.AI 76%

From Open Vocabulary to Open World: Teaching Vision Language Models to Detect Novel Objects

从开放词汇到开放世界:教会视觉语言模型检测新物体

Zizhao Li, Zhengkang Xiang, Joseph West, Kourosh Khoshelham

机构 * The University of Melbourne Parkville, VIC, Australia(墨尔本大学帕克维尔分校)

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV、cs.AI

AI总结 本文提出了一种开放世界框架,使OVD模型能够检测新物体,通过引入OWEL和MSCAL方法提升模型对远超出分布物体的识别能力。

Comments Accepted by BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22426 2026-02-27 cs.CV cs.LG 73%

SimpleOCR: Rendering Visualized Questions to Teach MLLMs to Read

SimpleOCR: 将可视化问题呈现以教导大语言模型阅读

Yibo Peng, Peng Xia, Ding Zhong, Kaide Zeng, Siwei Han, Yiyang Zhou, Jiaqi Liu, Ruiyi Zhang, Huaxiu Yao

机构 * UNC-Chapel Hill(北卡罗来纳大学教堂山分校) Carnegie Mellon University(卡内基梅隆大学) University of Michigan(密歇根大学) Adobe Research(Adobe研究)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.LG

AI总结 SimpleOCR通过强制模型视觉参与,解决多模态大语言模型在图像中阅读文本的性能问题,提升模型的视觉文本提取能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16782 2026-02-27 cs.LG cs.AI cs.CY 62%

What Is the Point of Equality in Machine Learning Fairness? Beyond Equality of Opportunity

机器学习公平性中平等的意义:超越机会的平等

Youjin Kong

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 本文探讨机器学习公平性中平等的意义,提出超越机会平等的平等主义框架,以解决结构性不平等和代表性伤害问题。

Comments Presented at ACM FAccT 2025; Forthcoming in ACM Journal on Responsible Computing

Journal ref ACM J. Responsib. Comput. 3, 1, Article 4 (March 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22549 2026-02-27 cs.CV cs.AI 62%

DrivePTS: A Progressive Learning Framework with Textual and Structural Enhancement for Driving Scene Generation

DrivePTS: 一种结合文本和结构增强的渐进式学习框架用于驾驶场景生成

Zhechao Wang, Yiming Zeng, Lufan Ma, Zeqing Fu, Chen Bai, Ziyao Lin, Cheng Lu

机构 * XPeng Motors

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 DrivePTS通过渐进式学习、多视角文本生成和频率引导结构损失,提升驾驶场景生成的保真度和可控性,生成稀有场景并增强泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22345 2026-02-27 cs.LG cs.AI 62%

Structure and Redundancy in Large Language Models: A Spectral Study via Random Matrix Theory

大语言模型的结构与冗余:通过随机矩阵理论的谱研究

Davide Ettori

机构 * Laurea Magistrale in Computer Science Engineering - Ingegneria Informatica(计算机科学工程硕士课程 - 信息工程)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI、cs.LG

AI总结 本研究通过谱几何和随机矩阵理论,提出EigenTrack用于检测大语言模型幻觉和分布外行为,以及RMT-KD用于压缩深度网络,提升模型可靠性与效率。

Comments Executive Summary of Master Thesis in Computer Science Engineering, Politecnico di Milano

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23339 2026-02-27 cs.CV 57%

Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?

检索与分割:在开放词汇分割中,几个示例是否足以弥合监督差距?

Tilemachos Aravanis, Vladan Stojnić, Bill Psomas, Nikos Komodakis, Giorgos Tolias

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出了一种基于检索增强的测试时间适配器,通过融合文本和视觉支持特征来实现开放词汇分割,有效缩小了零样本与监督分割间的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22963 2026-02-27 cs.AI 57%

FactGuard: Agentic Video Misinformation Detection via Reinforcement Learning

FactGuard:通过强化学习进行代理视频虚假信息检测

Zehao Li, Hongwei Yu, Hao Jiang, Qiang Sheng, Yilong Xu, Baolong Bi, Yang Li, Zhenlong Yuan, Yujun Cai, Zhaoqi Wang

机构 * Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China(中国科学院计算技术研究所) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学) University of Science and Technology Beijing(北京科技大学) The University of Queensland, Brisbane, Australia(昆士兰大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.AI

AI总结 FactGuard通过强化学习和代理框架,提升视频虚假信息检测的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11620 2026-02-27 cs.AI 57%

A Mind Cannot Be Smeared Across Time

意识无法跨越时间被抹平

Michael Timothy Bennett

机构 * Michael Timothy Bennett(独立研究者)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 作者提出意识统一性需要客观同时实例化,而非时间顺序实现,指出软件意识在严格顺序子系统中无法实现,强调硬件对意识的重要性。

Comments Forthcoming in the proceedings of the AAAI 2026 Spring Symposium on Machine Consciousness: Integrating Theory, Technology, and Philosophy

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06139 2026-02-27 cs.CV 57%

Deforming Videos to Masks: Flow Matching for Referring Video Segmentation

视频变形到掩码:用于指认视频分割的流匹配

Zanyi Wang, Dengyang Jiang, Liuzhuozheng Li, Sizhe Dang, Chengzu Li, Harry Yang, Guang Dai, Mengmeng Wang, Jingdong Wang

机构 * SGIT AI Lab, State Grid Corporation of China(国网信通研究院) University of California, San Diego(加州大学圣地亚哥分校) The Hong Kong University of Science and Technology(香港科技大学) The University of Tokyo(东京大学) University of Cambridge(剑桥大学) Zhejiang University of Technology(浙江工业大学) Baidu(百度)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本文提出FlowRVS框架,将指认视频分割视为条件连续流问题,通过学习视频整体表示到目标掩码的直接语言引导变形,在多个基准上取得新突破。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22219 2026-02-27 cs.IR cs.AI cs.CL 57%

Comparative Analysis of Neural Retriever-Reranker Pipelines for Retrieval-Augmented Generation over Knowledge Graphs in E-commerce Applications

电商应用场景下知识图谱检索增强生成中神经检索-排序流水线的比较分析

Teri Rumble, Zbyněk Gazdík, Javad Zarrin, Jagdeep Ahluwalia

机构 * Abertay University(阿伯泰大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文研究电商场景下知识图谱检索增强生成中神经检索-排序流水线的设计与比较,通过实验验证了优化配置在提升检索性能上的有效性。

Comments This manuscript is under review at the Springer journal Knowledge and Information Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22208 2026-02-27 cs.CV 57%

Solaris: Building a Multiplayer Video World Model in Minecraft

Solaris:在Minecraft中构建多玩家视频世界模型

Georgy Savva, Oscar Michel, Daohan Lu, Suppakit Waiwitlikhit, Timothy Meehan, Dhairya Mishra, Srivats Poddar, Jack Lu, Saining Xie

机构 * New York University(纽约大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 Solaris通过多玩家数据系统和分阶段训练方法,在Minecraft中构建了能模拟多视角观测的视频世界模型,提升了多代理交互的建模能力。

Comments Project website: https://solaris-wm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23075 2026-02-27 cs.CL cs.IR 50%

CiteLLM: An Agentic Platform for Trustworthy Scientific Reference Discovery

CiteLLM:一个用于可信科学引文发现的代理平台

Mengze Hong, Di Jiang, Chen Jason Zhang, Zichang Guo, Yawen Li, Jun Chen, Shaobo Cui, Zhiyang Su

机构 * Hong Kong Polytechnic University(香港理工大学) Beijing University of Posts and Telecommunications(北京邮电大学) Swiss Federal Technology Institute of Lausanne (EPFL)(洛桑联邦理工学院) Hong Kong University of Science and Technology (HKUST)(香港科学大学)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 CiteLLM通过在LaTeX编辑器中嵌入LLM工具,实现可信的科学引文发现,确保引文的准确性和可靠性。

Comments Accepted by TheWebConf 2026 Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏