arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

2026-04-03 至 2026-04-03 共收录 39
2601.15475 2026-04-03 cs.CV

Seeing through Light and Darkness: Sensor-Physics Grounded Deblurring HDR NeRF from Single-Exposure Images and Events

透过光与暗:基于传感器物理的去模糊HDR NeRF从单曝光图像和事件

Yunshan Qi, Lin Zhu, Nan Bao, Yifan Zhao, Jia Li

机构 * State Key Laboratory of Virtual Reality Technology and Systems, SCSE & QRI, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室) School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)

AI总结 本文提出基于传感器物理的NeRF框架,通过单曝光模糊LDR图像和对应事件实现高动态范围和锐利3D表示的合成,利用CRF模型提升HDR和去模糊效果。

Comments Accepted by the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2026. Project Page: https://icvteam.github.io/See-NeRF.html. Our code and datasets are publicly available at https://github.com/iCVTEAM/See-NeRF

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11508 2026-04-03 cs.CV

ReScene4D: Temporally Consistent Semantic Instance Segmentation of Evolving Indoor 3D Scenes

ReScene4D: 时空一致的动态室内3D场景语义实例分割

Emily Steiner, Jianhao Zheng, Henry Howard-Jenkins, Chris Xie, Iro Armeni

机构 * Stanford University(斯坦福大学) Meta Reality Labs Research(Meta现实实验室)

AI总结 ReScene4D提出了一种新的4D室内语义实例分割方法,通过时空对比损失、掩码和序列化技术,在无需密集观测的情况下实现实例跟踪和性能提升,引入t-mAP指标评估时空一致性。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21681 2026-04-03 cs.CV

Seeing without Pixels: Perception from Camera Trajectories

无需像素的感知:从相机轨迹进行感知

Zihui Xue, Kristen Grauman, Dima Damen, Andrew Zisserman, Tengda Han

机构 * Google DeepMind(谷歌DeepMind) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文探讨通过相机轨迹而非像素感知视频内容的可能性,提出CamFormer框架,证明轨迹信息能有效用于视频内容理解及多种下游任务。

Comments Accepted by CVPR 2026, Project website: https://sites.google.com/view/seeing-without-pixels

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18123 2026-04-03 cs.CV cs.AI cs.CL cs.LG

Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models

偏差是子空间,而非坐标:视觉-语言模型中事后去偏的几何重思

Dachuan Zhao, Weiyue Li, Zhenda Shen, Yushu Qiu, Bowen Xu, Haoyu Chen, Yongchao Chen

机构 * Harvard University(哈佛大学) MIT(麻省理工学院)

AI总结 本文提出SPD框架,通过几何方法识别并去除线性可解码的偏子空间,提升视觉-语言模型的公平性与任务性能。

Comments Accepted at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22652 2026-04-03 cs.RO cs.CV

Pixel Motion Diffusion is What We Need for Robot Control

我们为机器人控制需要像素运动扩散

E-Ro Nguyen, Yichi Zhang, Kanchana Ranasinghe, Xiang Li, Michael S. Ryoo

机构 * Stony Brook University(石溪大学)

AI总结 DAWN框架通过结构化像素运动表示连接高层运动意图与底层机器人动作,实现端到端可训练系统,展示多任务性能和现实迁移能力。

Comments Accepted to CVPR 2026. Project page: https://eronguyen.github.io/DAWN

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12957 2026-04-03 cs.CV

Adaptive Reinforcement for Open-ended Medical Reasoning via Semantic-Guided Reward Collapse Mitigation

通过语义引导的奖励崩溃缓解实现自适应的开放性医疗推理

Yizhou Liu, Dingkang Yang, Zizhi Chen, Minghao Han, Xukun Zhang, Keliang Liu, Jingwei Wei, Lihua Zhang

机构 * College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院) Fysics Intelligence Technologies Co., Ltd. (Fysics AI)(菲斯克智能科技有限公司(菲斯克AI)) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

AI总结 本文提出ARMed框架,通过监督微调和强化优化提升开放性医疗VQA的推理一致性和事实准确性,实验表明其在六个医疗VQA基准上显著提升准确性和泛化能力。

Comments Accept to 2026 CVPR Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10739 2026-04-03 cs.MM eess.IV

HippoMM: Hippocampal-inspired Multimodal Memory for Long Audiovisual Event Understanding

HippoMM:基于海马体的多模态记忆用于长音频视频事件理解

Yueqian Lin, Jingyang Zhang, Qinsi Wang, Hancheng Ye, Yuzhe Fu, Yudong Liu, Hai "Helen" Li, Yiran Chen

AI总结 HippoMM通过整合事件分割、记忆巩固和分层记忆检索,实现长音频视频事件理解,达到78.2%的准确率,比基线模型快5倍。

Comments Accepted at CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05808 2026-04-03 cs.CV

Fast Sphericity and Roundness approximation in 2D and 3D using Local Thickness

二维和三维中基于局部厚度的快速球形度和圆度近似

Pawel Tomasz Pieta, Peter Winkel Rasumssen, Anders Bjorholm Dahl, Anders Nymark Christensen

机构 * Technical University of Denmark(丹麦技术大学)

AI总结 本文提出基于局部厚度算法的新型方法,用于高效计算二维和三维图像中物体的球形度和圆度,通过简化表面面积计算和避免复杂角曲率确定过程,提升计算效率。

Comments Accepted at CVMI (CVPR 2025 Workshop)

Journal ref Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 4667-4677

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04727 2026-04-03 eess.IV cs.CV

Learning to Translate Noise for Robust Image Denoising

学习噪声翻译以实现稳健的图像去噪

Inju Ha, Donghun Ryou, Seonguk Seo, Bohyung Han

机构 * Seoul National University(首尔大学) Meta

AI总结 本文提出噪声翻译框架,通过将复杂噪声转换为高斯噪声以提升图像去噪的鲁棒性和泛化能力,采用预训练的去噪网络实现一致的去噪效果。

Comments Project page: https://hij1112.github.io/learning-to-translate-noise/ Accepted to CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏