arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

2026-02-27 至 2026-02-27 共收录 16
2602.23361 2026-02-27 cs.CV

VGG-T$^3$: Offline Feed-Forward 3D Reconstruction at Scale

VGG-T³:大规模离线前馈3D重建

Sven Elflein, Ruilong Li, Sérgio Agostinho, Zan Gojcic, Laura Leal-Taixé, Qunjie Zhou, Aljosa Osep

机构 * NVIDIA Vector Institute(向量研究所) University of Toronto(多伦多大学)

AI总结 VGG-T³通过测试时训练将变长键值空间蒸馏为固定大小的MLP,实现大规模离线前馈3D重建,速度比基线方法快11.6倍,并在重建精度和视觉定位能力上表现优异。

Comments CVPR 2026, Project page: https://research.nvidia.com/labs/dvl/projects/vgg-ttt

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23359 2026-02-27 cs.CV cs.AI

SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generation

SeeThrough3D: 3D布局生成中的遮挡感知控制

Vaibhav Agrawal, Rishubh Parihar, Pradhaan Bhat, Ravi Kiran Sarvadevabhatla, R. Venkatesh Babu

机构 * IIIT Hyderabad(IIIT海得拉尔) IISc Bengaluru(IISc班加罗尔)

AI总结 SeeThrough3D通过引入遮挡感知的3D场景表示和掩码自注意力机制,实现了对3D布局生成中遮挡关系的精确建模与控制。

Comments Project page: https://seethrough3d.github.io. Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23229 2026-02-27 cs.CV

Large Multimodal Models as General In-Context Classifiers

大多模态模型作为通用上下文分类器

Marco Garosi, Matteo Farina, Alessandro Conti, Massimiliano Mancini, Elisa Ricci

AI总结 本文提出CIRCLE方法,通过伪标签迭代优化,使LMM在开放世界分类中超越VLM,展示LMM作为统一分类器的潜力。

Comments CVPR Findings 2026. Project website at https://circle-lmm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23120 2026-02-27 cs.CV

TriLite: Efficient Weakly Supervised Object Localization with Universal Visual Features and Tri-Region Disentanglement

TriLite: 基于通用视觉特征和三区域解耦的高效弱监督物体定位

Arian Sabaghi, José Oramas

机构 * University of Antwerp(安特卫普大学) sqIRL/IDLab imec Antwerp(Antwerp imec)

AI总结 TriLite通过自监督学习和三区域解耦实现高效弱监督物体定位,无需端到端训练,参数更高效且训练更简单。

Comments This paper consists of 8 pages including 6 figures. Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22917 2026-02-27 cs.CV

Towards Multimodal Domain Generalization with Few Labels

迈向少标签多模态领域泛化的研究

Hongzhao Li, Hao Dong, Hualei Wan, Shupan Li, Mingliang Xu, Muhammad Haris Khan

机构 * Zhengzhou University(郑州大学) ETH Zürich(苏黎世联邦理工学院) MBZUAI(马克斯·普朗克智能系统研究所)

AI总结 本文提出SSMDG框架,通过三个关键组件实现少标签多模态领域泛化,提升跨模态鲁棒性和领域不变性。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22862 2026-02-27 cs.RO cs.CV

GraspLDP: Towards Generalizable Grasping Policy via Latent Diffusion

GraspLDP: 通过潜在扩散实现通用抓取策略

Enda Xiang, Haoxiang Ma, Xinzhu Ma, Zicheng Liu, Di Huang

机构 * State Key Laboratory of Complex and Critical Software Environment(复杂与关键软件环境国家重点实验室) School of Computer Science and Engineering(计算机科学与工程学院)

AI总结 GraspLDP通过引入潜在扩散框架和抓取先验知识,提升了抓取策略的精度和泛化能力,实验表明其在动态抓取任务中表现优异。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22819 2026-02-27 cs.CV

Face Time Traveller : Travel Through Ages Without Losing Identity

面容时间旅行者:无需丢失身份即可穿越年龄

Purbayan Kar, Ayush Ghadiya, Vishal Chudasama, Pankaj Wasnik, C. V. Jawahar

机构 * Sony Research India(索尼印度研究实验室) IIIT Hyderabad(Hyderabad 研究学院)

AI总结 FaceTT通过基于扩散的框架实现高保真身份一致的年龄变换,引入了Face-Attribute-Aware Prompt Refinement策略、无调优的Angular Inversion方法和自适应注意力控制机制,提升了年龄变换的精度和稳定性。

Comments Accepted at CVPR 2026 (Findings Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22727 2026-02-27 cs.CV

HulluEdit: Single-Pass Evidence-Consistent Subspace Editing for Mitigating Hallucinations in Large Vision-Language Models

HulluEdit: 单次通过证据一致子空间编辑用于缓解大视觉-语言模型中的幻觉

Yangguang Lin, Quan Fang, Yufei Li, Jiachen Sun, Junyu Gao, Jitao Sang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beijing Jiaotong University(北京交通大学)

AI总结 HulluEdit通过单次通过的正交子空间编辑方法,有效缓解大视觉-语言模型中的幻觉问题,实现了最先进的幻觉减少效果。

Comments accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22716 2026-02-27 cs.CV cs.AI

SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMs

SoPE: 基于球坐标的位置嵌入:增强3D大视觉-语言模型的空间感知

Guanting Ye, Qiyan Zhao, Wenhao Yu, Liangyu Yuan, Mingkai Li, Xiaofeng Zhang, Jianmin Ji, Yanyong Zhang, Qing Jiang, Ka-Veng Yuen

机构 * University of Macau(澳门大学) University of Science and Technology of China(中国科学技术大学) Shanghai Jiaotong University(上海交通大学) Hefei University of Technology(合肥工业大学) National University of Singapore(新加坡国立大学)

AI总结 SoPE通过基于球坐标的位置嵌入提升3D LVLMs的空间感知能力,结合多尺度频率混合策略,增强几何表示的一致性和表达性。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20903 2026-02-27 cs.CV

TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering

TextPecker: 通过奖励结构异常量化提升视觉文本渲染

Hanshen Zhu, Yuliang Liu, Xuecheng Wu, An-Lan Wang, Hao Feng, Dingkang Yang, Chao Feng, Can Huang, Jingqun Tang, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) ByteDance(字节跳动)

AI总结 TextPecker通过结构异常感知强化学习策略提升视觉文本渲染的结构忠实度和语义对齐度。

Comments Accepted by CVPR 2026; Code: https://github.com/CIawevy/TextPecker

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21058 2026-02-27 cs.CV

Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control

超越像素模拟:通过诊断语义标记和原型控制生成病理图像

Minghao Han, Yichen Liu, Yizhou Liu, Zizhi Chen, Jingqun Tang, Xuecheng Wu, Dingkang Yang, Lihua Zhang

机构 * College of Intelligent Robotics and Advanced Manufacturing(智能机器人与先进制造学院) Fudan University(复旦大学) Fysics Intelligence Technologies Co., Ltd.(Fysics智能科技有限公司) University of Science and Technology Beijing(北京科技大学) ByteDance(字节跳动) School of Computer Science and Technology(计算机科学与技术学院) Xi’an Jiaotong University(西安交通大学)

AI总结 UniPath通过诊断语义标记和原型控制实现可控的病理图像生成,取得SOTA性能,包括Patho-FID 80.9和98.7%的细粒度语义控制。

Comments accepted by CVPR 2026; 32 pages, 17 figures, and 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18897 2026-02-27 cs.CV

Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs

超越标签:基于推理增强的LMMs的无词汇细粒度识别

Dmitry Demidov, Zaigham Zaheer, Zongyan Han, Omkar Thawakar, Rao Anwer

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

AI总结 FiNDR通过推理增强的LMMs实现无词汇细粒度识别,提出三步流程生成候选标签、过滤排名并构建轻量级分类器,取得显著性能提升。

Journal ref CVPR 2026 (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10981 2026-02-27 cs.CV

CLIP-Free, Label Free, Unsupervised Concept Bottleneck Models

无需CLIP、无需标签的无监督概念瓶颈模型

Fawaz Sammani, Jonas Fischer, Nikos Deligiannis

机构 * Max-Planck-Institut für Informatik(马克斯·普朗克研究所) ETRO Department, Vrije Universiteit Brussel(自由大学布鲁塞尔分校ETRO部门)

AI总结 无需CLIP和标签的无监督概念瓶颈模型,通过无监督方式推导线性分类器,提升图像分类和零样本描述性能。

Comments CVPR 2026 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22620 2026-02-27 cs.CV

Coded-E2LF: Coded Aperture Light Field Imaging from Events

编码的E2LF:从事件中获取4D光场

Tomoya Tsuchida, Keita Takahashi, Chihiro Tsutake, Toshiaki Fujii, Hajime Nagahara

机构 * Nagoya University, Japan(名古屋大学,日本) Osaka University, Japan(大阪大学,日本)

AI总结 本文提出Coded-E2LF方法,通过编码孔径和事件-only相机实现4D光场的像素级精确重建,首次证明仅使用事件数据即可获得高精度光场。

Comments accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22594 2026-02-27 cs.CV

Causal Motion Diffusion Models for Autoregressive Motion Generation

因果运动扩散模型用于自回归运动生成

Qing Yu, Akihisa Watanabe, Kent Fujiwara

机构 * LY Corporation(LY公司) Waseda University(早稻田大学)

AI总结 本文提出因果运动扩散模型,通过因果扩散变换器实现高质量运动生成,提升语义保真度与时间平滑度,同时减少推理延迟。

Comments Accepted to CVPR 2026, Project website: https://yu1ut.com/CMDM-HP/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22376 2026-02-27 cs.CV cs.AI

AeroDGS: Physically Consistent Dynamic Gaussian Splatting for Single-Sequence Aerial 4D Reconstruction

AeroDGS:基于物理的动态高斯点溅技术用于单序列航空4D重建

Hanyang Liu, Rongjun Qin

AI总结 AeroDGS通过物理引导的4D高斯点溅技术,解决单目无人机视频中动态物体的4D重建问题,提升重建精度与稳定性。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏