arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-06-30 至 2026-06-30 共收录 132 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 22 篇

2606.28338 2026-06-30 cs.IR 50%

Memory Shot for Long-Term Dialogue

长期对话的记忆快照

Chunyi Peng, Haidong Xin, Xuanshuo Sheng, Xin Dai, Zhenghao Liu, Shuo Wang, Yukun Yan, Zulong Chen, Yu Gu, Ge Yu

专题命中 视觉推理 :visual reasoning(abstract)

AI总结 提出MemShot方法,通过将对话片段渲染为结构化视觉记忆单元,利用模型内部视觉推理能力关联关键情节,实现高效长期对话建模,在LoCoMo和LongMemEval上取得稳定性能并实现70倍加速。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉定位与Grounding 41 篇

2606.29649 2026-06-30 cs.CL 90%

Resolution Thresholds in VLM Detection of Harmful ASCII Art Across Construction Modes and Languages

VLM检测有害ASCII艺术的分辨率阈值:跨构建模式与语言的研究

Yikai Hua, Peter West

机构 * Department of Computer Science(计算机科学系) The University of British Columbia(不列颠哥伦比亚大学)

专题命中 视觉定位与Grounding :VLM(title,title_cn);vision-language model(abstract)

AI总结 研究图像分辨率如何影响VLM检测有害ASCII艺术,发现检测率在特定分辨率阈值以上急剧下降,且基于单词的模式最难检测,揭示了VLM内容审核系统的系统性漏洞。

Comments 13 pages, 9 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29267 2026-06-30 cs.CV 90%

Enhancing Part-Level Point Grounding for Any Open-Source MLLMs

增强任意开源多模态大语言模型的部件级点定位能力

Jin-Cheng Jhang, Fu-En Wang, Xin Yang, Nan Qiao, Lu Xia, Min Sun, Cheng-Hao Kuo

机构 * National Tsing Hua University(国立清华大学) Amazon(亚马逊)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 提出一种通用方法,通过冻结原模型参数并引入Q-Synth模块和注意力到点解码器,为任意开源MLLM赋予精确的2D部件级点定位能力,显著提升部件级定位精度。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30247 2026-06-30 cs.CL 90%

Grounding LLM Reasoning under Incomplete Graph Evidence

在不完全图证据下 grounding LLM 推理

Jiaqi Li, Fanghui Song

机构 * Tianjin Normal University, College of Computer and Information Engineering(天津师范大学计算机与信息工程学院) Harbin Institute of Technology, School of Mathematics(哈尔滨工业大学数学学院)

专题命中 视觉定位与Grounding :grounding(title,title_cn)

AI总结 本文提出在不完全知识图谱证据下,通过KL正则化变形LLM先验实现软grounding,并给出稳定性界限,适用于GraphRAG、KGQA等场景。

Comments A theoretical perspective about Grounding LLM Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28401 2026-06-30 cs.CV cs.LG 90%

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs

视觉驱动的偏好合成用于缓解VLM中的幻觉

Yunhun Nam, Jongheon Jeong

专题命中 视觉定位与Grounding :VLM(title_cn,summary_cn);vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.LG

AI总结 提出ViPSy框架,通过视觉线索构建策略对齐且视觉 grounded 的偏好数据,显著降低VLM幻觉率,在AMBER和Object HalBench上分别降低35.7%和24.5%,并提升通用视觉基准性能。

Comments 29 pages; Code is available at https://github.com/yunpal/ViPSy

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28385 2026-06-30 cs.RO cs.AI 86%

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

RoboGaze: 通过结构化视觉-语言分析评估机器人世界模型

Minh-Loi Nguyen, Nghiem Tuong Diep, Hung Khang Nguyen, Minh Le, Doanh Le Thien, Hoang H. Tran, Dung D. Le, Vu N. Duong, Daniel Sonntag, An Thai Le, Duy Minh Ho Nguyen, Vien Anh Ngo, Tran Van Nhiem

机构 * Ho Chi Minh City University of Science, Vietnam National University(越南国家科学大学胡志明市大学) VinRobotics VinUniversity German Research Center for AI (DFKI)(人工智能研究中心(DFKI)) University of Stuttgart(斯图加特大学) Max Planck Research School for Intelligent Systems(智能系统马克斯·普朗克研究学校) Technische Universität Darmstadt(达姆施塔特技术大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);grounding(abstract);分类 cs.AI

AI总结 提出RoboGaze,一种无需训练的多智能体VLM框架,通过任务-场景对齐、维度专家路由和批评验证三阶段流水线,对机器人操作视频进行结构化可解释评估,显著提升描述F1和时序对齐性能。

Comments First version 29 pages, 7 figures. Project webpage: https://robogaze-eval.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30632 2026-06-30 cs.RO cs.AI cs.CV 86%

GROW$^2$: Grounding Which and Where for Robot Tool Use

GROW$^2$:为机器人工具使用确定哪个物体和哪个部位

Yuhong Deng, Yuyao Liu, David Hsu

机构 * National University of Singapore(新加坡国立大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 提出GROW$^2$方法,通过分层语义和几何接地实现开放世界工具选择与部位定位,在零样本泛化中优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29023 2026-06-30 cs.CV cs.AI 86%

Efficient Spatio-Temporal Grounding with Multimodal Large Models via Second-Level Tracking and RL Verification

基于二级跟踪和强化学习验证的多模态大模型高效时空定位

Tianshu Zhang, Yan Wang, Ji Qi, Lijie Wen

机构 * Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 提出一种从帧级到秒级跟踪的流水线,结合跨秒平滑、链式思维轨迹合成和强化学习优化,在长视频中实现高效且准确的时空定位。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07683 2026-06-30 cs.CV cs.AI 84%

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding

TAR:基于时间锚的视频时间定位推理

Chaohong Guo, Xun Mo, Yongwei Nie, Fei Ma, Xuemiao Xu, Chengjiang Long

机构 * South China University of Technology(华南理工大学) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳)) Bytedance Inc.(字节跳动有限公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 TAR通过引入时间锚机制,提升视频时间定位任务中推理过程的可信度和自主性,采用自举方法生成高质量推理轨迹,实现更精确的最终预测。

Comments Accepted by ECCV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28920 2026-06-30 cs.CV 83%

ExACT: Exemplar-Driven Calibrated Refinement for Training-Free Visual Grounding in Remote Sensing Images

ExACT: 基于示例驱动的校准精化用于遥感图像中免训练的视觉定位

Zixiao Zhang, Lingling Li, Pei He, Xu Liu, Licheng Jiao

机构 * Xidian University(西安电子科技大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 提出ExACT框架,通过一次性视觉提示机制弥合多模态大语言模型在遥感视觉定位中的模态差距,实现免训练的精确像素级定位。

Comments 11 pages, 8 figures, supplementary material included

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28520 2026-06-30 cs.CV cs.CL 83%

Detecting Clinical Hallucinations in LVLMs via Counterfactual Visual Grounding Uncertainty

通过反事实视觉基础不确定性检测LVLM中的临床幻觉

Xiao Song, Haonan Qin, Zhaoxu Zhang, Jiong Zhang, Yuqi Fang, Caifeng Shan

机构 * School of Intelligent Science and Technology, Nanjing University(南京大学智能科学与技术学院) National Institute of Healthcare Data Science at Nanjing University(南京大学医疗数据科学国家研究院) School of Biomedical Engineering, Nanjing University(南京大学生物医学工程学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

AI总结 提出一种基于反事实视觉基础不确定性的框架,通过提取实体并对比正反事实定位结果计算不确定性分数,无需修改模型即可检测LVLM在临床图像中的幻觉。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27122 2026-06-30 cs.CV 83%

InterPartAbility: Phrase-Region Grounding for Interpretable Text-to-Image Person Re-Identification

InterPartAbility: 基于文本引导的部分匹配用于可解释的人员重识别

Shakeeb Murtaza, Aryan Shukla, Rajarshi Bhattacharya, Maguelonne Heritier, Eric Granger

机构 * LIVIA, Dept. of Systems Engineering, ETS Montreal, Canada(LIVIA系统工程系,蒙特利尔ÉTS学院,加拿大) Genetec Inc.(Genetec公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本文提出InterPartAbility,通过显式部分匹配和短语-区域绑定提升TI-ReID的可解释性,引入PPIM模块实现概念级指导,生成 grounded 解释图谱,实验表明在CUHK-PEDES等基准上达到SOTA可解释性性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26553 2026-06-30 cs.CV 83%

SemConFlow: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching

HolisticSemGes: 语义 grounding 的整体共语手势生成与对比流匹配

Lanmiao Liu, Esam Ghaleb, Aslı Özyürek, Zerrin Yumak

机构 * Max Planck Institute for Psycholinguistics(马克斯·普朗克心理语言学研究所) Utrecht University(乌得勒支大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn);分类 cs.CV

AI总结 本文提出一种基于对比流匹配的共语手势生成模型,通过引入不匹配的音频-文本条件作为负样本,确保跨模态一致性,并在BEAT2和SHOW数据集上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28724 2026-06-30 cs.CV cs.AI 82%

CCRC: A Change-Aware Captioning and Reasoning Chain for Image Change Captioning and Segmentation

CCRC:面向图像变化描述与分割的变化感知描述与推理链

Jinhong Hu, Xiaoping Wang, Shuyin Huang, Guojin Zhong, Kaitai Liu, Kai Lu

机构 * College of Computer Science and Electronic Engineering, Hunan University, China(湖南大学计算机科学与电子工程学院,中国) Guangxi Minzu University, China(广西民族大学,中国) National University of Defense Technology, China(国防科学技术大学,中国)

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 提出变化感知描述与推理链(CCRC),通过双链框架解耦语义推理与空间分割,实现图像变化描述与分割(ICCS)任务,在合成和真实基准上达到最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29847 2026-06-30 cs.CV 81%

See Only When Needed: Context-Aware Attention Intervention for Mitigating Hallucinations in LVLMs

仅在需要时看:用于缓解LVLM中幻觉的上下文感知注意力干预

Yuqing Lei, Wenbo Lyu, Yingjun Du, Xiantong Zhen, Cees G. M. Snoek, Ling Shao

机构 * University of Chinese Academy of Sciences(中国科学院大学) University of Amsterdam(阿姆斯特丹大学) United Imaging Healthcare Co., Ltd.(联合影像医疗科技股份有限公司)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 提出训练自由的上下文感知注意力干预(CAI),通过两轴选择性(看哪里和何时干预)在解码时针对性增强视觉 grounding,有效缓解物体幻觉并保持语言流畅性。

Journal ref ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23294 2026-06-30 cs.CV 79%

Towards Long-Form Spatio-Temporal Video Grounding

面向长形式时空视频定位

Xin Gu, Bing Fan, Jiali Yao, Zhipeng Zhang, Yan Huang, Cheng Han, Heng Fan, Libo Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出ART-STVG方法,针对长时视频中的目标定位问题,通过自回归Transformer架构和记忆银行设计,提升对长视频中时空信息的处理能力,实验表明其在长形式STVG任务中优于现有方法。

Comments 22 pages, 11 figures. Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01215 2026-06-30 cs.CV cs.AI cs.CL cs.MM 79%

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs

将神经符号程序蒸馏到3D多模态大语言模型中

Wentao Mo, Yang Liu

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV、cs.AI

AI总结 提出APEIRIA,通过三阶段课程学习将符号推理模式蒸馏到3D多模态大语言模型中,实现透明推理与开放词汇空间推理的统一。

Comments To appear in ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30220 2026-06-30 cs.CV 77%

From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA

从准确性到视觉依赖性:审计与过滤交通视频问答中的模态崩溃

Sena Korkut, María Alejandra Bravo Sarmiento, Sanghwan Kim, Zeynep Akata

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract_cn);grounding(abstract);分类 cs.CV

AI总结 针对交通视频问答中模型依赖文本捷径而非视觉证据的问题,提出盲差距和视觉增益两个数据集级诊断指标,以及实例级捷径分数实现无训练过滤,提升视觉基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28696 2026-06-30 cs.AI 74%

COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

COMPASS:在统一多模态模型中锚定构图意图引导

Ziqi Zhou, Weize Quan, Mining Tan, Zhihan Chen, Dandan Zheng, Jingdong Chen, Jun Zhou, Weiming Dong, Dong-Ming Yan

机构 * University of Edinburgh(爱丁堡大学) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS)(多模态人工智能系统国家重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) Ant Group(蚂蚁集团)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

AI总结 提出统一框架COMPASS,通过共享专家令牌τ_c实现构图感知与生成的双向锚定,结合MoE骨干和Comp-11数据集,显著提升细粒度构图理解与可控生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30393 2026-06-30 cs.CV 70%

SADL: What to Ignore? A Benchmark for Subject-Aware Distractor Localization

SADL:该忽略什么?面向主体感知的干扰物定位基准

Cao-Tri Nguyen, Nguyen-Khoa Luong, Vinh-Tiep Nguyen, Minh-Triet Tran

机构 * University of Science(科学大学) Vietnam National University Ho Chi Minh City(越南胡志明市国家大学) University of Information Technology(信息技术大学) John Von Neumann Institute(约翰·冯·诺依曼研究所)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract_cn);分类 cs.CV

AI总结 提出SADL基准,包含1,800个主体感知案例,用于评估视觉语言模型在干扰物定位中的排除校准能力,揭示模型过度排除的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28343 2026-06-30 cs.IR cs.AI 70%

The Crowded Embedding Space: A Mean-Field Mechanism for Emergent Marginalization in Retrieval-Augmented Agents

拥挤的嵌入空间:检索增强型智能体中涌现边缘化的平均场机制

Shwan Ashrafi, Dan Roth

专题命中 视觉定位与Grounding :grounding(abstract,abstract_cn);分类 cs.AI

AI总结 研究检索增强型智能体中密集检索导致少数内容被边缘化的问题,通过静态和动态分析揭示平均场机制,证明局部相关性目标会系统性地排斥少数兴趣。

Journal ref ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21526 2026-06-30 cs.CV 70%

VIGIL: Part-Grounded Structured Reasoning for Generalizable Deepfake Detection

VIGIL:基于部分的结构化推理用于通用深度伪造检测

Xinghan Li, Junhao Xu, Jingjing Chen

机构 * Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究所) Shanghai Key Laboratory of Multimodal Embodied AI(上海市多模态具身人工智能重点实验室)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 VIGIL通过计划-检验流程,结合部分级证据和结构化推理,提升深度伪造检测的可解释性和泛化能力。

Comments Project Page: https://vigil.best

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12598 2026-06-30 cs.CV 70%

Neural Gate: Mitigating Privacy Risks in LVLMs via Neuron-Level Gradient Gating

神经门:通过神经层面梯度门缓解LVLMs的隐私风险

Xiangkui Cao, Jie Zhang, Meina Kan, Shiguang Shan, Xilin Chen

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);LLaVA(abstract);分类 cs.CV

AI总结 本文提出Neural Gate方法,通过神经层面模型编辑提升LVLMs对隐私问题的拒绝率,增强隐私保护同时保持模型原有功能。

Comments Accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29091 2026-06-30 cs.LG cs.AI cs.DB 62%

Statistically Indistinguishable, Operationally Distinct: A Formal Barrier for Tabular Foundation Models

统计不可区分,操作上截然不同:表格基础模型的形式化障碍

Tassilo Klein, Johannes Hoffart

机构 * SAP SE(SAP公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 本文提出操作图灵测试,证明仅基于值的表格模型无法区分合法与违规数据库状态,而操作审计特征可突破该障碍,揭示可识别性而非模型容量是根本限制。

Comments Accepted at the 2nd ICML Workshop on Foundation Models for Structured Data, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28896 2026-06-30 eess.IV cs.AI cs.LG 62%

A Task-Driven and Quality-Assured Agent Framework for SAR Data Generation

面向SAR数据生成的任务驱动与质量保障智能体框架

Xuanting Wu, Fan Zhanga, Fei Ma, Ling Guan, Guochun Ma, Yongsheng Zhou

机构 * College of Information Science and Technology, Beijing University of Chemical Technology(北京化工大学信息科学与技术学院) Science and Technology on Electromagnetic Scattering Laboratory, Beijing Institute of Environmental Features(北京环境特征研究院电磁散射实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 提出SAGA框架,通过模式约束规划与多维度评估器,实现面向任务的SAR数据增强,提升可靠性、可复现性及下游任务效用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02327 2026-06-30 cs.CV cs.AI 62%

Steerable Visual Representations

可操控的视觉表示

Jona Ruthardt, Manu Gaur, Deva Ramanan, Makarand Tapaswi, Yuki M. Asano

机构 * University of Technology Nuremberg(纽伦堡工业大学) Carnegie Mellon University(卡内基梅隆大学) International Institute of Information Technology, Hyderabad(海得拉巴国际信息技术学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出可操控的视觉表示,通过自然语言指导视觉特征,提升视觉任务的灵活性和泛化能力。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30638 2026-06-30 cs.CV 57%

Open-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D Detectors

开放词汇与指代分割:利用2D检测器实现3D高斯

Jameel Hassan, Yasiru Ranasinghe, Vishal Patel

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 提出GaussDet方法,利用离散的2D开放词汇检测器与指代表达能力,通过视图聚合语义标签分布实现3D场景的开放词汇分割和指代分割,在零样本设置下显著提升指代分割性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30458 2026-06-30 cs.CV 57%

Cross-Resolution Semantic Transfer for Robust Text-to-Image Retrieval in Low-Resolution Surveillance

跨分辨率语义迁移用于低分辨率监控下的鲁棒文本-图像检索

Wenjie Qian, Bin Yang, Xiao Wang, Wenke Huang, Ling Mei, Xin Xu, Mang Ye

机构 * School of Computer Science and Technology, Wuhan University of Science and Technology(武汉科技大学计算机科学与技术学院) School of Computer Science, National Engineering Research Center for Multimedia Software, Wuhan University(武汉大学计算机学院,国家多媒体软件工程技术研究中心) Hubei Province Key Laboratory of Intelligent Information Processing and Real-time Industrial System, Wuhan University of Science and Technology(湖北省智能信息处理与实时工业系统重点实验室,武汉科技大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 针对低分辨率监控场景中文本-图像检索的可靠性崩溃和排序漂移问题,提出CLIP框架CRST,通过分辨率条件推理、文本引导精炼和跨分辨率邻域迁移,在三个数据集上平均提升超低分辨率Rank-1和mAP分别5.7%和5.3%。

Comments 10 pages,8 figures,conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30288 2026-06-30 cs.CV 57%

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context

VisReflect: 用于长视觉上下文中细粒度感知的潜在视觉反射

Xiaoqian Shen, Mohamed Elhoseiny

机构 * King Abdullah University of Science and Technology(卡布斯大学)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

AI总结 针对长视觉上下文中细粒度感知的挑战,提出VisReflect框架,通过潜在空间中的连续视觉反射选择性强调相关区域,无需显式定位或额外前向传播,在图像和视频基准上分别提升4.1%和1.8%,并减少约44%推理时间。

Comments Accepted to ECCV 2026; Project page: https://xiaoqian-shen.github.io/VisReflect

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29850 2026-06-30 cs.CV 57%

Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction

面向具身AI的高效视觉指代:智能体驱动数据合成、跨块注意力与迭代校正

Zijian Hong, Qi Lv, Yuxiang Xie, Jianming Xing, Xiang Deng, Weili Guan, Liqiang Nie

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 提出PointArena 2026解决方案,通过智能体驱动数据合成、确定性可操控数据流水线、跨块注意力与迭代校正模块,实现77.2%总体准确率,排名第二。

详情

展开后加载摘要…

URL PDF HTML 收藏