arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

International Conference on Computer Vision · 会议 · Computer Vision

共收录 4770
2507.07316 2026-05-14 cs.LG cs.CR

AdeptHEQ-FL: Adaptive Homomorphic Encryption for Federated Learning of Hybrid Classical-Quantum Models with Dynamic Layer Sparing

AdeptHEQ-FL: 适应性同态加密用于混合经典-量子模型联邦学习的框架

Md Abrar Jahin, Taufikur Rahman Fuad, M. F. Mridha, Nafiz Fahad, Md. Jakir Hossen

机构 * University of Southern California(南加州大学) Islamic University of Technology(伊斯兰科技大学) American International University-Bangladesh(孟加拉国美国国际大学) Multimedia University(多媒体大学)

AI总结 本文提出AdeptHEQ-FL框架,结合混合CNN-PQC架构、自适应准确度加权聚合方案、选择性同态加密和动态层冻结技术,提升联邦学习在非独立同分布环境下的性能、隐私保护与通信效率。

Comments Accepted in 1st International Workshop on ICCV'25 BISCUIT (Biomedical Image and Signal Computing for Unbiasedness, Interpretability, and Trustworthiness)

Journal ref 1st International Workshop on BISCUIT at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09479 2026-05-13 cs.CV cs.GR cs.LG

DiFaReli++: Diffusion Face Relighting with Consistent Cast Shadows

DiFaReli++: 基于扩散的面部光照重建与一致阴影生成

Puntawat Ponglertnapakorn, Nontawat Tritrong, Supasorn Suwajanakorn

机构 * School of Information Science and Technology, Vidyasirimedhi Institute of Science and Technology(信息科学与技术学院,维达亚西里米迪科学技术研究所)

AI总结 本文提出一种单视图面部光照重建方法,通过条件扩散隐式模型实现无光照地面真值训练,实现真实光照下的阴影一致性。

Comments Published in IEEE TPAMI (vol. 48, no. 5, May 2026). This is an extended version of the ICCV 2023 paper (DiFaReli)

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 5, pp. 5068-5082, May 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23617 2026-05-12 cs.CV cs.AI cs.GR cs.LG

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory

一条轨迹,一个标记:通过全景子对象轨迹实现的 grounded 视频标记化

Chenhao Zheng, Jieyu Zhang, Mohammadreza Salehi, Ziqi Gao, Vishnu Iyengar, Norimasa Kobori, Quan Kong, Ranjay Krishna

机构 * University of Washington(华盛顿大学) Allen Institute for Artificial Intelligence(人工智能艾伦研究所) Woven by Toyota, Inc(丰田公司)

AI总结 本文提出 grounded 视频标记化方法,通过全景子对象轨迹组织标记,减少冗余并保持时间一致性,TrajViT 在多个视频理解基准上优于空间-时间 ViT,具有更高的效率和性能。

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01662 2026-05-05 cs.CV

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models

视频主动感知:基于视觉-语言模型的高效推理时长视频理解

Martin Q. Ma, Willis Guo, Aditya Agrawal, Ankit Gupta, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency

机构 * Carnegie Mellon University(卡内基梅隆大学) MIT(麻省理工学院)

AI总结 本文提出视频主动感知方法,通过主动感知理论提升长视频问答性能,实现帧效率提升5.6倍,优于现有模型。

Comments ICCV 2025 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22699 2026-05-04 cs.CV

Image-Guided Shape-from-Template Using Mesh Inextensibility Constraints

基于网格不可伸长约束的图像引导形状从模板方法

Thuy Tran, Ruochen Chen, Shaifali Parashar

机构 * CNRS(法国国家科学研究中心) École Centrale de Lyon(里昂中央理工大学) INSA Lyon(里昂国立应用科学学院) Université Claude Bernard Lyon 1(里昂一大学) LIRIS(图像研究所)

AI总结 本文提出一种无监督的形状从模板方法,利用图像观测和网格不可伸长约束,实现比现有无监督方法快400倍的重建速度,并在细节生成和严重遮挡处理上表现更优。

Comments Accepted to ICCV 2025. Total 13 pages, 9 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28196 2026-05-01 cs.CV

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation

HERMES++:迈向统一的驾驶世界模型用于3D场景理解和生成

Xin Zhou, Dingkang Liang, Xiwu Chen, Feiyang Tan, Dingyuan Zhang, Hengshuang Zhao, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Mach Drive University of Hong Kong(香港大学)

AI总结 本文提出HERMES++,一种统一的驾驶世界模型,整合3D场景理解和未来几何预测。通过BEV表示、LLM增强世界查询和当前到未来链接等设计,提升驾驶场景的生成与理解能力。

Comments Extended version of ICCV 25 paper HERMES, Code: https://github.com/H-EmbodVis/HERMESV2, Project page: https://h-embodvis.github.io/HERMESV2/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10210 2026-04-29 cs.CV

TARS: Traffic-Aware Radar Scene Flow Estimation

TARS:面向交通的雷达场景流估计

Jialong Wu, Marco Braun, Dominic Spata, Matthias Rottmann

机构 * University of Wuppertal(乌尔姆大学) Osnabrück University(奥斯纳布吕克大学) Aptiv Services Deutschland GmbH(Aptiv Services德国公司)

AI总结 本文提出TARS方法,通过交通层面的运动刚性提升雷达场景流估计,结合目标检测与场景流联合优化,构建交通向量场以实现交通层面的场景理解。

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 26075-26084

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10171 2026-04-23 cs.CV cs.ET

SynSpill: Improved Industrial Spill Detection With Synthetic Data

SynSpill:基于合成数据的改进工业泄漏检测

Aaditya Baranwal, Abdul Mueez, Jason Voelker, Guneet Bhatia, Shruti Vyas

机构 * University of Central Florida(佛罗里达中央大学) Siemens Energy(西门子能源)

AI总结 针对工业泄漏检测中数据稀缺问题,提出基于高质量合成数据生成的框架,通过参数高效微调提升VLMs和检测器性能,实现安全关键领域的有效应用。

Comments Accepted at ICCV (VISION'25 Workshop) 2025

Journal ref 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), pp. 1425-1434

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08508 2026-04-23 cs.CV cs.CL

Re:Verse -- Can Your VLM Read a Manga?

Re:Verse -- 能读懂漫画吗?

Aaditya Baranwal, Madhav Kataria, Naitik Agrawal, Yogesh S Rawat, Shruti Vyas

机构 * University of Central Florida(中央佛罗里达大学) Indian Institute of Technology, Jodhpur(印度理工学院,朱达浦尔) Indian Institute of Technology, Varanasi(印度理工学院,瓦拉纳西)

AI总结 本文通过分析漫画叙事理解,揭示现有VLM在时间因果和跨面板连贯性上的不足,提出新的评估框架,系统研究长篇叙事理解能力。

Comments Accepted (oral) at ICCV (AISTORY Workshop) 2025

Journal ref 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), pp. 3820-3830

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16719 2026-04-23 cs.CV cs.LG

Learn2Synth: Learning Optimal Data Synthesis Using Hypergradients for Brain Image Segmentation

Learn2Synth: 利用超梯度学习最优数据合成用于脑图像分割

Xiaoling Hu, Xiangrui Zeng, Oula Puonti, Juan Eugenio Iglesias, Bruce Fischl, Yael Balbastre

机构 * Massachusetts General Hospital and Harvard Medical School(麻省总医院和哈佛医学院) Danish Research Centre for Magnetic Resonance, Copenhagen University Hospital(丹麦磁共振研究中心,哥本哈根大学医院) Centre for Medical Image Computing, University College London(医学影像计算中心,伦敦大学学院) Computer Science and AI Laboratory, Massachusetts Institute of Technology(计算机科学与人工智能实验室,麻省理工学院) Department of Experimental Psychology, University College London(实验心理学系,伦敦大学学院)

AI总结 通过超梯度学习优化数据合成参数,提升脑图像分割网络在真实数据上的性能,避免依赖真实数据训练。

Comments 16 pages, 5 figures. Accepted by ICCV'25. Bruce Fischl and Yael Balbastre are co-senior authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14461 2026-04-21 cs.CV

Ouroboros: Single-step Diffusion Models for Cycle-consistent Forward and Inverse Rendering

Ouroboros: 单步扩散模型用于循环一致的正向与反向渲染

Shanlin Sun, Yifan Wang, Hanwen Zhang, Yifeng Xiong, Qin Ren, Ruogu Fang, Xiaohui Xie, Chenyu You

机构 * University of California, Irvine(加州大学伊维特分校) Stony Brook University(石溪大学) Huazhong University of Science and Technology(华中科技大学) University of Florida(佛罗里达大学)

AI总结 Ouroboros通过双单步扩散模型实现正反向渲染的互促,扩展了内在分解到室内外场景,并引入循环一致性机制,实验显示其在多样场景中表现优异且推理速度显著提升。

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13401 2026-04-21 cs.CV

AIM 2025 Rip Current Segmentation (RipSeg) Challenge Report

AIM 2025 涨潮分离(RipSeg)挑战报告

Andrei Dumitriu, Florin Miron, Florin Tatui, Radu Tudor Ionescu, Radu Timofte, Aakash Ralhan, Florin-Alexandru Vasluianu, Shenyang Qian, Mitchell Harley, Imran Razzak, Yang Song, Pu Luo, Yumei Li, Cong Xu, Jinming Chai, Kexin Zhang, Licheng Jiao, Lingling Li, Siqi Yu, Chao Zhang, Kehuan Song, Fang Liu, Puhua Chen, Xu Liu, Jin Hu, Jinyang Xu, Biao Liu

机构 * Computer Vision Lab, CAIDAS & IFI, University of Würzburg, Germany(计算机视觉实验室,CAIDAS与IFI,乌尔姆大学,德国)

AI总结 本文报告了AIM 2025 RipSeg挑战赛,旨在提升静止图像中涨潮分离的自动分割技术。挑战赛基于最大的涨潮数据集RipVIS,聚焦单类实例分割,通过F1、F2、AP50和AP[50:95]等指标评估,揭示了涨潮分割的现状与未来方向。

Comments Challenge report paper from AIM Workshop at ICCV 2025

Journal ref 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23104 2026-04-15 cs.CV

DC-TTA: Divide-and-Conquer Framework for Test-Time Adaptation of Interactive Segmentation

DC-TTA:用于交互分割测试时间适应的分而治之框架

Jihun Kim, Hoyong Kwon, Hyeokjun Kweon, Wooseong Jeong, Kuk-Jin Yoon

机构 * KAIST(韩国科学技术院) Chung-Ang University(Chung-Ang 大学)

AI总结 本文提出DC-TTA框架,通过分而治之策略改进交互分割中用户交互的适应性,提升复杂场景下的分割性能。

Comments accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14075 2026-04-13 cs.CV cs.CL

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models

通过蒸馏和强化学习生长多头Twig以加速大视觉-语言模型

Zhenwei Shao, Mingyang Wang, Weijun Zhang, Zhou Yu, Wenwen Pan, Yan Yang, Tao Wei, Hongyuan Zhang, Jun Yu

机构 * Zhejiang Key Laboratory of Space Information Sensing and Transmission, School of Computer Science, Hangzhou Dianzi University(浙江省空间信息感知与传输重点实验室,计算机学院,杭州电子科技大学) Li Auto Inc.(理想汽车) School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)智能科学与工程学院)

AI总结 本文提出TwigVLM,通过在基础VLM早期层上生长轻量模块Twig,结合 Twig 引导的 token 剪枝和自推测解码策略,实现更高的准确性和速度。实验表明, TwigVLM 在剪枝88.9%的视觉token后仍保持96%的原始性能,并在生成长响应时达到154%的速度提升。

Comments An extended version of our ICCV paper at ICCV2025/html/Shao_Growing_a_Twig_to_Accelerate_Large_Vision-Language_Models_ICCV_2025_paper.html" target="_blank" rel="noopener">https://openaccess.thecvf.com/content/ICCV2025/html/Shao_Growing_a_Twig_to_Accelerate_Large_Vision-Language_Models_ICCV_2025_paper.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21435 2026-04-08 cs.CV cs.AI

MedShift: Implicit Conditional Transport for X-Ray Domain Adaptation

MedShift:隐式条件传输用于X射线领域迁移

Francisco Caetano, Christiaan Viviers, Peter H. N. De With, Fons van der Sommen

机构 * Eindhoven University of Technology(埃因霍温理工大学)

AI总结 本文提出MedShift,一种基于流匹配和Schrodinger Bridges的统一类条件生成模型,用于解决合成与真实X射线图像之间的跨领域翻译问题,通过学习共享的领域无关潜在空间实现高质量无配对图像转换。

Comments Accepted at the ICCV 2025 AIM Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03334 2026-04-07 cs.CV

Bridging the Dimensionality Gap: A Taxonomy and Survey of 2D Vision Model Adaptation for 3D Analysis

弥合维度差距:2D视觉模型适应3D分析的分类与综述

Akshat Pandya, Bhavuk Jain

机构 * Independent Researcher(独立研究员)

AI总结 本文综述了将2D视觉模型适应3D分析的策略,分类为数据导向、架构导向和混合方法,探讨了计算复杂度、预训练依赖性和几何归纳偏置的权衡。

Comments VISAPP 2026

Journal ref Proceedings of the 21st International Conference on Computer Vision Theory and Applications - Volume 3: VISAPP 2026; ISBN 978-989-758-804-4; ISSN 2184-4321, SciTePress, pages 353-364

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08751 2026-04-07 cs.CV cs.LG

Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning

解耦世界模型:从干扰视频中学习转移语义知识以用于强化学习

Qi Wang, Zhipeng Zhang, Baao Xie, Xin Jin, Yunbo Wang, Shiyu Wang, Liaomo Zheng, Xiaokang Yang, Wenjun Zeng

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(上海交通大学人工智能研究院教育部人工智能重点实验室) Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo, China(东方理工高等研究院宁波数字孪生研究院) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Ningbo, China(宁波市空间智能与数字衍生重点实验室) University of Chinese Academy of Sciences(中国科学院大学) Shenyang Institute of Computing Technology, Chinese Academy of Sciences(中国科学院沈阳计算技术研究所) Shenyang CASNC Technology Co., Ltd(沈阳中科数控技术股份有限公司)

AI总结 本文提出了解耦世界模型,通过离线到在线的潜在蒸馏和灵活解耦约束,从干扰视频中学习语义知识,提升强化学习的样本效率。

Comments Accepted by ICCV 2025. Project page: https://qiwang067.github.io/diswm

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10637 2026-04-02 cs.CV

Processing and acquisition traces in visual encoders: What does CLIP know about your camera?

视觉编码器中的处理与获取轨迹:CLIP对你相机知道什么?

Ryan Ramos, Vladan Stojnić, Giorgos Kordopatis-Zilos, Yuta Nakashima, Giorgos Tolias, Noa Garcia

机构 * The University of Osaka(大阪大学) VRG, FEE, Czech Technical University in Prague(捷克理工大学电气工程学院视觉识别组)

AI总结 研究分析了视觉编码器对图像获取过程参数的编码能力,发现这些参数能影响语义预测,取决于其与语义标签的相关性。

Comments 8 main pages, supplementary attached, ICCV 2025 highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12988 2026-04-02 cs.CV cs.LG

Variance-Based Pruning for Accelerating and Compressing Trained Networks

基于方差的剪枝:用于加速和压缩训练好的网络

Uranik Berisha, Jens Mehnert, Alexandru Paul Condurache

机构 * Automated Driving Research, Robert Bosch GmbH(罗伯特·博世有限公司自动驾驶研究) Institute for Signal Processing, University of Lübeck(吕贝克大学信号处理研究所)

AI总结 本文提出基于方差的剪枝方法,通过激活统计信息选择神经元进行剪枝,减少训练成本并保持性能,实验显示在ImageNet-1k任务中DeiT-Base模型剪枝后性能仍达70%以上,仅需10轮微调即可恢复99%精度,同时降低MACs和模型大小。

Comments Accepted as Oral at ICCV'25 (IEEE/CVF International Conference on Computer Vision)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12690 2026-04-01 cs.CV cs.AI cs.LG

TTA-DAME: Test-Time Adaptation with Domain Augmentation and Model Ensemble for Dynamic Driving Conditions

TTA-DAME: 测试时适应与领域增强和模型集成用于动态驾驶条件

Dongjae Jeon, Taeheon Kim, Seongwon Cho, Minhyuk Seo, Jonghyun Choi

机构 * Yonsei University(延世大学)

AI总结 本文提出TTA-DAME方法,通过领域增强和模型集成应对动态变化的驾驶场景,提升模型在天气变化下的适应能力。

Comments 1st Place in Continual Test-time Adaptation for Object Detection Challenge at VCL Workshop, ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24117 2026-03-26 cs.CV

Combi-CAM: A Novel Multi-Layer Approach for Explainable Image Geolocalization

David Faget, José Luis Lisani, Miguel Colom

机构 * Centre Borelli, ENS Paris-Saclay, Université de Paris, CNRS, INSERM, SSA, France(博雷利中心,巴黎-萨克雷大学,巴黎大学,法国国家科学研究中心,法国国家医学研究院,SSA,法国)

Journal ref 21st International Conference on Computer Vision Theory and Applications, Mar 2026, Marbella, Spain. pp.275-281

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14951 2026-03-26 cs.CV

Morph: A Motion-free Physics Optimization Framework for Human Motion Generation

Morph:一种无需运动的物理优化框架用于人类运动生成

Zhuo Li, Mingshuang Luo, Ruibing Hou, Xin Zhao, Hao Liu, Hong Chang, Zimo Liu, Chen Li

机构 * WeChat, Tencent Inc(微信,腾讯公司) Key Laboratory of Intelligent Information Processing of Chinese Academy of Sciences (CAS), Institute of Computing Technology, CAS, China(中国科学院智能信息处理重点实验室(CAS),计算技术研究所,CAS,中国) Peng Cheng Laboratory, China(鹏城实验室,中国) University of Chinese Academy of Sciences, China(中国科学院大学,中国) MoE Key Laboratory of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能MoE重点实验室,人工智能研究院,上海交通大学)

AI总结 本文提出Morph框架,通过运动生成器和运动物理细化模块提升运动的物理合理性,无需昂贵真实运动数据,实验显示其在文本到运动和音乐到舞蹈生成任务中达到最先进的运动质量。

Comments Accepted by ICCV 2025, 15 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20694 2026-03-24 quant-ph

FALQON-MST: A Fully Quantum Framework for Graph Optimization in Vision Systems

FALQON-MST: 一种用于视觉系统图优化的全量子框架

Guilherme E. L. Pexe, Lucas A. M. Rattighieri, Leandro A. Passos, Douglas Rodrigues, Danilo S. Jodas, João P. Papa, Kelton A. P. da Costa

AI总结 本文提出一种全量子管道计算图的最小生成树,采用反馈量子优化方法FALQON,通过构造哈密顿量形式和不同FALQON策略比较,展示多驱动配置和时间缩放在提升地面态保真度中的作用。

Comments 8 pages. Accepted for publication at the 21st International Conference on Computer Vision Theory and Applications (VISAPP 2026), Marbella, Spain

Journal ref Proceedings of the 21st International Conference on Computer Vision Theory and Applications - Volume 1: VISAPP, 419-426, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20519 2026-03-24 cs.CV

End-to-End Optimization of Polarimetric Measurement and Material Classifier

极化测量与材料分类的端到端优化

Ryota Maeda, Naoki Arikawa, Yutaka No, Shinsaku Hiura

机构 * Graduate School of Engineering, University of Hyogo(大阪大学工学研究院)

AI总结 本文提出端到端优化框架,联合学习材料分类器并确定极化元件的最佳旋转角度配置,以提高极化测量下的材料分类精度。

Comments Presented at VISAPP 2026 (21st International Conference on Computer Vision Theory and Applications)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20930 2026-03-23 cs.CV

Computing a Characteristic Orientation for Rotation-Independent Image Analysis

计算旋转无关图像分析的特征方向

Cristian Valero-Abundio, Emilio Sansano-Sansano, Raúl Montoliu, Marina Martínez García

机构 * Institute of New Imaging Technologies, Universitat Jaume I, 12071 Castellón, Spain(新成像技术研究所,Jaume I大学,西班牙卡斯蒂利亚-拉曼查省12071)

AI总结 本文提出GID预处理方法,通过估计图像全局方向并对其对齐,提升模型在不同旋转下的鲁棒性,实验表明其在旋转MNIST和CIFAR-10数据集上均优于现有旋转不变架构。

Comments Accepted for publication at the 21st International Conference on Computer Vision Theory and Applications (VISAPP 2026). 8 pages

Journal ref Proceedings of the 21st International Conference on Computer Vision Theory and Applications - Volume 2: VISAPP (2026), SciTePress, pp. 644-651

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08723 2026-03-17 cs.LG cs.CV

Is CLIP ideal? No. Can we fix it? Yes!

CLIP是否理想?不。我们能否修复它?是的!

Raphi Kang, Yue Song, Georgia Gkioxari, Pietro Perona

AI总结 本文分析了CLIP潜在空间的几何局限性,提出DCSM方法以解决其根本问题,提升了CLIP-like模型在多个基准上的性能。

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14150 2026-03-17 cs.CV

CIPHER: Culvert Inspection through Pairwise Frame Selection and High-Efficiency Reconstruction

CIPHER: 通过成对帧选择和高效重建进行涵洞检查

Seoyoung Lee, Zhangyang Wang

AI总结 本文提出一种高效的RGB基3D重建管道,用于在视觉重复环境中对涵洞结构进行自动检查,通过成对帧选择和实时重建提升检查效率。

Comments Accepted by ICCV 2026 End-to-End 3D Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13984 2026-03-17 cs.CV cs.AI

CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models

CSD-VAR:视觉自回归模型中的内容-风格分解

Quang-Binh Nguyen, Minh Luu, Quang Nguyen, Anh Tran, Khoi Nguyen

AI总结 本文提出CSD-VAR方法,通过视觉自回归模型的逐尺度生成过程提升内容-风格分离,引入三种创新:尺度感知交替优化策略、基于SVD的校正方法和增强键值记忆以提升内容保真度。

Comments Accepted to ICCV 2025; Project page: https://nqbinhcs.github.io/csd-var-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04862 2026-03-13 cs.CV

Contact-Aware Refinement of Human Pose Pseudo-Ground Truth via Bioimpedance Sensing

基于生物阻抗传感的人体姿态伪真实地面 truth 的接触感知细化

Maria-Paola Forte, Nikos Athanasiou, Giulia Ballardini, Jan Ulrich Bartels, Katherine J. Kuchenbecker, Michael J. Black

AI总结 BioTUCH通过结合视觉姿态估计与生物阻抗传感,利用接触感知优化提升人体姿态重建精度,平均提高11.7%,并提供高效的大规模接触感知训练数据收集方案。

Comments * Equal contribution. Minor figure corrections compared to the ICCV 2025 version

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14351 2026-03-13 cs.CV

SegAnyPET: Universal Promptable Segmentation from Positron Emission Tomography Images

SegAnyPET:从正电子发射断层扫描图像中进行通用可提示分割

Yichi Zhang, Le Xue, Wenbo Zhang, Lanlan Li, Yuchen Liu, Chen Jiang, Yuan Cheng, Yuan Qi

AI总结 SegAnyPET通过构建PETS-5k数据集,提出跨提示自信学习策略,实现从PET图像中高效分割多种器官,提升分割任务的通用性和准确性。

Comments Accept for ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏