arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

International Conference on Computer Vision · 会议 · Computer Vision

共收录 4770
2507.01654 2026-03-09 cs.CV cs.LG

SPoT: Subpixel Placement of Tokens in Vision Transformers

SPoT: 视觉变换器中令牌的子像素位置

Martine Hjelkrem-Tan, Marius Aasan, Gabriel Y. Arteaga, Adín Ramírez Rivera

AI总结 SPoT通过子像素位置策略提升视觉变换器的稀疏性利用,减少令牌数量以提高效率与准确性。

Comments Appeared in Workshop on Efficient Computing under Limited Resources: Visual Computing (ICCV 2025). Code available at https://github.com/dsb-ifi/SPoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09879 2026-03-05 cs.CV

TextMaster: A Unified Framework for Realistic Text Editing via Glyph-Style Dual-Control

TextMaster: 一种通过字形-风格双控实现真实文本编辑的统一框架

Zhenyu Yan, Jian Wang, Aoqiang Wang, Yuhan Li, Wenxiang Shang, Ran Lin

机构 * Taobao & Tmall Group of Alibaba(阿里巴巴淘宝与天猫集团) Shanghai Jiao Tong University(上海交通大学)

AI总结 TextMaster通过字形-风格双控机制,实现了高精度、可控的文本编辑,适用于多种场景和图像区域,提升了文本布局的准确性和风格的可控性。

Comments Accepted to ICCV 2025

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 16112-16121

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23782 2026-03-03 cs.CV

MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion

MonoFusion:通过单目融合实现稀疏视角4D重建

Zihan Wang, Jeff Tan, Tarasha Khurana, Neehar Peri, Deva Ramanan

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 MonoFusion通过单目融合技术,在稀疏视角下实现高质量的动态场景4D重建,优于传统密集多视角方法。

Comments ICCV 2025. Project Page: https://z1hanw.github.io/research/25_DSR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00150 2026-03-03 cs.CV cs.CY

Attention to Neural Plagiarism: Diffusion Models Can Plagiarize Your Copyrighted Images!

对神经剽窃的关注:扩散模型可以剽窃您的受版权保护的图像!

Zihang Zou, Boqing Gong, Liqiang Wang

机构 * University of Central Florida(中央佛罗里达大学) Boston University(波士顿大学)

AI总结 本文提出了一种基于梯度的神经剽窃方法,通过扰动扩散模型的交叉注意机制,使模型能够绕过版权保护措施,揭示了扩散模型在复制受版权图像方面的潜在风险。

Comments Accepted to ICCV 2025. Code available at: https://github.com/zzzucf/Neural-Plagiarism

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22624 2026-02-27 cs.CV cs.AI

Instruction-based Image Editing with Planning, Reasoning, and Generation

基于指令的图像编辑与规划、推理和生成

Liya Ji, Chenyang Qi, Qifeng Chen

机构 * HKUST(香港科技大学)

AI总结 本文提出一种多模态模型,通过链式思考规划、编辑区域推理和编辑,提升基于指令的图像编辑能力,以应对更复杂的真实场景。

Comments 10 pages, 7 figures

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, Page 17506--17515

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21904 2026-02-26 cs.CV cs.RO

UNet-Based Keypoint Regression for 3D Cone Localization in Autonomous Racing

基于UNet的关键点回归用于自动驾驶赛车中3D圆锥定位

Mariia Baidachna, James Carty, Aidan Ferguson, Joseph Agrane, Varad Kulkarni, Aubrey Agub, Michael Baxendale, Aaron David, Rachel Horton, Elliott Atkinson

机构 * School of Computer Science, University of Glasgow(计算机科学学院,格拉斯哥大学) Amazon(亚马逊)

AI总结 本文提出基于UNet的关键点回归方法,用于自动驾驶赛车中3D圆锥定位,通过大规模数据集提升定位精度并实现颜色预测。

Comments 8 pages, 9 figures. Accepted to ICCV End-to-End 3D Learning Workshop 2025 and presented as a poster; not included in the final proceedings due to a conference administrative error

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07525 2026-02-25 cs.CV cs.AI

EHWGesture -- A dataset for multimodal understanding of clinical gestures

EHWGesture -- 一个用于多模态理解临床手势的数据集

Gianluca Amprimo, Alberto Ancilotto, Alessandro Savino, Fabio Quazzolo, Claudia Ferraris, Gabriella Olmo, Elisabetta Farella, Stefano Di Carlo

机构 * Department of Control and Computer Engineering, Politecnico di Torino(控制与计算机工程系,都灵理工大学) Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会) CNR-IEIIT(意大利国家研究委员会-IEIIT)

AI总结 EHWGesture是一个包含五个临床相关手势的多模态视频数据集,用于提升临床手势理解的多模态分析能力。

Comments Accepted at ICCV 2025 Workshop on AI-driven Skilled Activity Understanding, Assessment & Feedback Generation

Journal ref 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16076 2026-02-24 cs.CV cs.GR

Geometry Distributions

几何分布

Biao Zhang, Jing Ren, Peter Wonka

机构 * KAUST(卡斯土尼亚大学) ETH Zurich(苏黎世联邦理工学院)

AI总结 本文提出了一种基于分布的几何数据表示方法,利用扩散模型学习表面点分布,以提升3D几何建模的精度和灵活性。

Comments Accepted to ICCV 2025. For the project site, see https://1zb.github.io/GeomDist/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19318 2026-02-23 cs.RO cs.CV

Perception-to-Pursuit: Track-Centric Temporal Reasoning for Open-World Drone Detection and Autonomous Chasing

感知到追捕:面向开放世界的无人机检测与自主追捕的以轨迹为中心的时序推理

Venkatakrishna Reddy Oruganti

机构 * Sithara Inc.(Sithara公司)

AI总结 P2P通过轨迹为中心的时序推理框架,提升了无人机轨迹预测精度和追捕可行性,实现了准确的预测与可操作的追捕规划。

Comments 7 pages, 2 figures, 3 tables, 15 references. Intended for submission to ICCV 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23592 2026-02-23 cs.LG cs.AI

Toward a Holistic Approach to Continual Model Merging

迈向持续模型融合的总体方法

Hoang Phan, Sungmin Cha, Tung Lam Tran, Qi Lei

机构 * New York University(纽约大学) University of Rochester(罗切斯特大学)

AI总结 本文提出了一种持续模型融合的总体方法,通过预融合、融合过程和后融合三个阶段解决持续学习中的挑战,实现可扩展和高效的灾难性遗忘缓解。

Comments Accepted to Workshop on Continual Learning in Computer Vision, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.12217 2026-02-20 cs.CV

LoLep: Single-View View Synthesis with Locally-Learned Planes and Self-Attention Occlusion Inference

LoLep:单视角视图合成与局部学习平面及自注意力遮挡推断

Cong Wang, Yu-Ping Wang, Dinesh Manocha

机构 * Tsinghua University(清华大学) The Beijing Institute of Technology(北京理工大学) University of Maryland(马里兰大学)

AI总结 LoLep通过局部学习平面和自注意力机制实现单视角视图合成,有效提升遮挡推断和生成质量。

Comments Accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22399 2026-02-18 cs.CV

VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow

VITAL: 通过分布对齐和相关信息流实现更可理解的特征可视化

Ada Gorgun, Bernt Schiele, Jonas Fischer

机构 * Max Planck Institute for Informatics(马克斯·普朗克信息研究所)

AI总结 VITAL通过分布对齐和相关信息流提升特征可视化,实现更清晰的神经网络信息解码。

Comments Accepted at the International Conference on Computer Vision 2025 (ICCV 2025). Code is available at: https://github.com/adagorgun/VITAL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11586 2026-02-17 cs.CV

StrandHead: Text to Hair-Disentangled 3D Head Avatars Using Human-Centric Priors

StrandHead: 基于人类中心先验的文本驱动3D头像生成方法

Xiaokun Sun, Zeyu Cai, Ying Tai, Jian Yang, Zhenyu Zhang

机构 * Nanjing University(南京大学)

AI总结 StrandHead通过人类中心先验生成逼真3D头发并实现解耦头像建模,达到文本到发束生成的最新水平。

Comments Accepted by ICCV 2025; Project page: https://xiaokunsun.github.io/StrandHead.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01102 2026-02-11 cs.CV

Keystep Recognition using Graph Neural Networks

基于图神经网络的按键识别

Julia Lee Romero, Kyle Min, Subarna Tripathi, Morteza Karimzadeh

机构 * University of Colorado Boulder(科罗拉多大学波得罗尔分校) Intel Labs(英特尔实验室)

AI总结 本文提出GLEVR框架,通过构建图结构有效利用第一人称视频的长期依赖关系,实现细粒度按键识别,优于现有方法。

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, 2025, pp. 7624-7633

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10913 2026-02-11 cs.CV cs.CL

Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP

了解“不”:一种数据驱动的方法用于增强CLIP中的否定意识

Junsung Park, Jungbeom Lee, Jongyoon Song, Sangwon Yu, Dahuin Jung, Sungroh Yoon

机构 * Department of Electrical and Computer Engineering, Seoul National University(首尔国立大学电子与计算机工程系) Amazon(亚马逊) Samsung Research(三星研究院) School of Computer Science and Engineering, Soongsil University(顺天大学计算机科学与工程学院) IPAI, AIIS, ASRI, INMC, and ISRC, Seoul National University(首尔国立大学IPAI、AIIS、ASRI、INMC和ISRC)

AI总结 本文提出NegationCLIP,通过生成包含否定的数据增强CLIP的否定意识,同时提出NegRefCOCOg基准用于评估多模态模型的否定理解能力。

Comments Accepted to ICCV 2025

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 2825-2835

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21265 2026-02-10 cs.CV cs.AI

MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation

MedVSR:基于跨状态空间传播的医学视频超分辨率

Xinyu Liu, Guolei Sun, Cheng Wang, Yixuan Yuan, Ender Konukoglu

机构 * The Chinese University of Hong Kong(香港中文大学) ETH Zurich(苏黎世联邦理工学院) Computer Vision Laboratory(计算机视觉实验室)

AI总结 MedVSR通过跨状态空间传播和内状态空间重建模块,提升医学视频超分辨率的重建性能和效率。

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10602 2026-02-10 cs.CV cs.AI cs.CL

TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention

TruthPrInt: 通过潜在真实引导预干预缓解大视觉-语言模型对象幻觉

Jinhao Duan, Fei Kong, Hao Cheng, James Diffenderfer, Bhavya Kailkhura, Lichao Sun, Xiaofeng Zhu, Xiaoshuang Shi, Kaidi Xu

机构 * Drexel University(德雷塞尔大学) University of Electronic Science and Technology of China(电子科技大学) Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) LLNL(劳伦斯利弗莫尔国家实验室) Lehigh University(莱斯大学)

AI总结 TruthPrInt通过学习真相方向并引导解码过程,有效缓解大视觉-语言模型中的对象幻觉问题。

Comments 15 pages, 9 figures, the first two authors contributed equally, Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16001 2026-02-06 cs.CV

Image-to-Image Translation with Diffusion Transformers and CLIP-Based Image Conditioning

基于扩散变换器和CLIP的图像到图像翻译

Qiang Zhu, Kuan Lu, Menghao Huo, Yuxiao Li

机构 * Department of Mechanical and Aerospace Engineering(机械与航空航天工程系) University of Houston(休斯顿大学) School of Engineering(工程学院) Santa Clara University(圣克拉拉大学) School of Electrical and Computer Engineering(电气与计算机工程学院) Cornell University(康奈尔大学) Department of Electrical and Computer Engineering(电气与计算机工程系) Northeastern University(东北大学)

AI总结 本文提出基于扩散变换器和CLIP的图像到图像翻译方法,通过CLIP嵌入引导实现高质量、语义一致的图像转换,为配对图像翻译任务提供新方案。

Comments Published in: 2025 6th International Conference on Computer Vision, Image and Deep Learning (CVIDL)

Journal ref 2025 6th International Conference on Computer Vision, Image and Deep Learning (CVIDL), pp. 626-632,

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19263 2026-02-04 cs.CV

DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning

DWIM: 向工具感知的视觉推理迈进:通过差异感知的工作流生成与指令遮蔽微调

Fucai Ke, Vijay Kumar B G, Xingjian Leng, Zhixi Cai, Zaid Khan, Weiqing Wang, Pari Delir Haghighi, Hamid Rezatofighi, Manmohan Chandraker

AI总结 DWIM通过差异感知的工作流生成与指令遮蔽微调,提升工具感知的视觉推理性能。

Comments ICCV 2025

Journal ref In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16403 2026-02-03 cs.CV

ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering

ReasonVQA: 一个带有结构知识的多跳推理问答基准数据集

Duong T. Tran, Trung-Kien Tran, Manfred Hauswirth, Danh Le Phuoc

机构 * Bosch Center for AI(博世人工智能中心) Technical University of Berlin(柏林技术大学) Fraunhofer FOKUS(弗劳恩霍夫研究所)

AI总结 ReasonVQA是一个结合结构知识的多跳推理问答基准数据集,通过生成复杂问题挑战现有VQA模型,推动领域发展。

Comments Accepted at the IEEE/CVF International Conference on Computer Vision (ICCV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00618 2026-02-03 cs.CV

Tune-Your-Style: Intensity-tunable 3D Style Transfer with Gaussian Splatting

Tune-Your-Style: 可调强度的3D风格迁移与高斯点散布

Yian Zhao, Rushi Ye, Ruochong Zheng, Zesen Cheng, Chaoran Feng, Jiashu Yang, Pengchong Qiao, Chang Liu, Jie Chen

机构 * School of Electronic and Computer Engineering, Peking University, Shenzhen, China(电子与计算机工程学院,北京大学深圳校区) Pengcheng Laboratory, Shenzhen, China(鹏城实验室) AI for Science (AI4S)-Preferred Program, Peking University Shenzhen Graduate School, China(人工智能科学(AI4S)优选计划,北京大学深圳研究生院) Department of Automation and BNRist, Tsinghua University, Beijing, China(自动化系和BNRist,清华大学,北京,中国) Dalian University of Technology, China(大连理工大学)

AI总结 Tune-Your-Style通过可调强度的3D风格迁移方法,实现灵活的内容-风格平衡,提升3D风格迁移的定制性。

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15388 2026-02-03 cs.CV cs.AI cs.CL

LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models

LLaVA-PruMerge: 适应性令牌减少用于高效的大多模态模型

Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, Yan Yan

机构 * UCF(佛罗里达大学) UW-Madison(威斯康星大学麦迪逊分校) USC(南加州大学) UIC(伊利诺伊大学香槟分校)

AI总结 LLaVA-PruMerge通过自适应视觉令牌减少策略,显著降低视觉令牌数量而不影响性能,适用于高效的大多模态模型。

Comments Accepted to ICCV 2025. First Version is released in 2024/03

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00108 2026-02-03 cs.CV cs.AI

SITUATE -- Synthetic Object Counting Dataset for VLM training

SITUATE -- 合成物体计数数据集用于视觉语言模型训练

René Peinl, Vincent Tischler, Patrick Schröder, Christian Groth

AI总结 SITUATE是一个用于训练视觉语言模型计数任务的数据集,通过控制遮挡和空间组成,提升模型对非分布图像的泛化能力。

Comments accepted at 21st International Conference on Computer Vision Theory and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17564 2026-02-02 eess.IV cs.CV cs.LG

ModalTune: Fine-Tuning Slide-Level Foundation Models with Multi-Modal Information for Multi-task Learning in Digital Pathology

ModalTune: 通过多模态信息细调滑片级基础模型以实现数字病理学中的多任务学习

Vishwesh Ramanathan, Tony Xu, Pushpak Pati, Faruk Ahmed, Maged Goubran, Anne L. Martel

机构 * Sunnybrook Research Institute(辛普森布鲁斯研究所在) University of Toronto(多伦多大学) Google Research(谷歌研究)

AI总结 ModalTune通过引入多模态信息和大型语言模型,实现数字病理学中多任务学习的统一细调框架,提升癌症生存和亚型预测性能。

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22349 2026-01-29 cs.CV

GCRayDiffusion: Pose-Free Surface Reconstruction via Geometric Consistent Ray Diffusion

GCRayDiffusion:基于几何一致射线扩散的无姿态表面重建

Li-Heng Chen, Zi-Xin Zou, Chang Liu, Tianjiao Jing, Yan-Pei Cao, Shi-Sheng Huang, Hongbo Fu, Hua Huang

机构 * Beijing Normal University(北京师范大学) VAST(中国科学院软件研究所) Hong Kong University of Science and Technology(香港科技大学)

AI总结 GCRayDiffusion通过几何一致射线扩散实现无姿态表面重建,提升稀疏视角下的3D一致性与精度。

Journal ref Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (ICCV), 2025, pp. 25335-25345

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08005 2026-01-28 cs.CV cs.AI eess.IV

DisCoPatch: Taming Adversarially-driven Batch Statistics for Improved Out-of-Distribution Detection

DisCoPatch: 通过对抗驱动的批量统计来改进分布外检测

Francisco Caetano, Christiaan Viviers, Luis A. Zavala-Mondragón, Peter H. N. de With, Fons van der Sommen

机构 * Eindhoven University of Technology(埃因霍温理工大学)

AI总结 DisCoPatch通过利用对抗驱动的批量统计改进分布外检测,实现高效且准确的检测性能。

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18536 2026-01-27 cs.CV cs.CL cs.IR cs.LG

FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA

FilterRAG: 零样本引导式检索增强生成以缓解视觉问答中的幻觉

Nobin Sarwar

机构 * University of Maryland, Baltimore County(马里兰大学巴尔的摩分校)

AI总结 FilterRAG通过结合BLIP-VQA与检索增强生成,利用外部知识源减少视觉问答中的幻觉问题,提升模型在知识驱动和分布外场景的鲁棒性。

Comments 12 pages, 6 figures and 2 tables; Accepted at ICCV 2025 Workshop on Building Foundation Models You Can Trust (T2FM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25237 2026-01-26 cs.CV

DeepShield: Fortifying Deepfake Video Detection with Local and Global Forgery Analysis

DeepShield: 通过局部和全局伪造分析强化深度伪造视频检测

Yinqi Cai, Jichang Li, Zhaolun Li, Weikai Chen, Rushi Lan, Xi Xie, Xiaonan Luo, Guanbin Li

机构 * Sun Yat-sen University(中山大学) Pengcheng Laboratory(鹏城实验室) Guilin University of Electronic Technology(桂林电子科技大学) Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)

AI总结 DeepShield通过局部和全局伪造分析提升深度伪造检测的鲁棒性,有效应对未知伪造攻击。

Comments ICCV 2025. Code is available at https://github.com/lijichang/DeepShield

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11213 2026-01-23 cs.CV eess.IV

Simulating Dual-Pixel Images From Ray Tracing For Depth Estimation

从光线追踪模拟双像素图像用于深度估计

Fengchen He, Dayang Zhao, Hao Xu, Tingwei Quan, Shaoqun Zeng

机构 * Huazhong University of Science and Technology(华中科技大学)

AI总结 本文提出Sdirt方案,通过光线追踪生成逼真的双像素图像,用于提升深度估计模型对真实双像素数据的泛化能力。

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10860 2026-01-22 cs.CV cs.GR

RI3D: Few-Shot Gaussian Splatting With Repair and Inpainting Diffusion Priors

RI3D: 少样本高斯点云渲染与修复与修复扩散先验

Avinash Paliwal, Xilong Zhou, Wei Ye, Jinhui Xiong, Rakesh Ranjan, Nima Khademi Kalantari

机构 * Texas A&M University(德克萨斯A&M大学) Meta Reality Labs(Meta现实实验室) Max Planck Institute for Informatics(马克斯·普朗克信息研究所)

AI总结 RI3D通过分离视图合成任务并结合修复与修复扩散模型,实现了高质量的少样本3D渲染与缺失区域重建。

Comments ICCV 2025, Project page: https://people.engr.tamu.edu/nimak/Papers/RI3D, Code: https://github.com/avinashpaliwal/RI3D

详情

展开后加载摘要…

URL PDF HTML 收藏