arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Winter Conference on Applications of Computer Vision · 会议 · Computer Vision

共收录 2109
2604.17054 2026-04-21 cs.CV cs.AI

mEOL: Training-Free Instruction-Guided Multimodal Embedder for Vector Graphics and Image Retrieval

mEOL:无需训练的指令引导多模态嵌入器用于矢量图形和图像检索

Kyeong Seon Kim, Baek Seong-Eun, Lee Jung-Mok, Tae-Hyun Oh

机构 * KAIST(韩国科学技术院) POSTECH

AI总结 本文提出无需训练的指令引导多模态嵌入框架,通过多模态大语言模型将文本、位图和SVG代码映射到对齐的嵌入空间,利用模态特定指令和结构化SVG提示实现嵌入方向控制,构建首个文本到SVG检索基准,展示出优于传统基线的性能。

Comments Round 1 early acceptance to WACV 2026, Project page: https://scene-the-ella.github.io/meol

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10073 2026-04-08 cs.CV cs.AI

ReaMIL: Reasoning- and Evidence-Aware Multiple Instance Learning for Whole-Slide Histopathology

ReaMIL:基于推理和证据的多实例学习用于整张滑片病理学

Hyun Do Jung, Jungwon Choi, Hwiyoung Kim

机构 * Yonsei University(延世大学) KAIST(韩国科学技术院)

AI总结 ReaMIL通过引入轻量选择头改进多实例学习,实现紧凑证据集和高准确率,在病理学图像分类中表现优异。

Comments Accepted at LFMBio Workshop, WACV 2026. Oral Presentation

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Workshops, March 2026, pp. 40-45

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01891 2026-04-07 cs.CV

Agentic AI in Remote Sensing: Foundations, Taxonomy, and Emerging Systems

遥感中的代理AI:基础、分类与新兴系统

Niloufar Alipour Talemi, Julia Boone, Fatemeh Afghah

机构 * Clemson University(克莱姆森大学)

AI总结 本文探讨遥感中代理AI的基础、分类及新兴系统,提出统一的分类框架,分析规划机制、检索增强生成和记忆结构,并评估轨迹感知推理的准确性。

Comments Accepted to the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026, GeoCV Workshop

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Workshops, 2026, pp. 786-799

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13855 2026-04-03 cs.CV cs.AI

Improvise, Adapt, Overcome -- Telescopic Adapters for Efficient Fine-tuning of Vision Language Models in Medical Imaging

即兴、适应、克服——用于医疗影像中视觉语言模型高效微调的望远镜适配器

Ujjwal Mishra, Vinita Shukla, Praful Hambarde, Amit Shukla

机构 * Centre for Artificial Intelligence and Robotics, Indian Institute of Technology Mandi(印度理工学院曼迪分校人工智能与机器人中心)

AI总结 本文提出望远镜适配器,通过深度感知缩放提升视觉语言分割模型在医疗影像中的微调效率,仅用613k参数在五个数据集上取得优异性能。

Comments Accepted at the IEEE/CVF winter conference on applications of computer vision (WACV 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04065 2026-04-01 cs.CV cs.LG

Unsupervised Modular Adaptive Region Growing and RegionMix Classification for Wind Turbine Segmentation

无监督模块自适应区域生长与区域混合分类用于风力涡轮机分割

Raül Pérez-Gonzalo, Riccardo Magro, Andreas Espersen, Antonio Agudo

机构 * Politecnico di Milano(米兰理工大学) Wind Power LAB(风力发电实验室)

AI总结 本文提出一种高效的无监督区域分类方法,通过自适应阈值和区域融合技术实现涡轮机叶片分割,提升跨站点的泛化能力。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15244 2026-04-01 cs.CV cs.LG

Image-Specific Adaptation of Transformer Encoders for Compute-Efficient Segmentation

针对计算效率的Transformer编码器特定适应

Manyi Yao, Abhishek Aich, Yumin Suh, Amit Roy-Chowdhury, Christian Shelton, Manmohan Chandraker

机构 * NEC Laboratories, America(美国NEC实验室) University of California, Riverside(加州大学河滨分校) University of California, San Diego(加州大学圣地亚哥分校)

AI总结 本文提出ECO-M2F方法,通过自适应选择编码器层数来提升分割任务的计算效率,同时保持性能并适应不同计算资源。

Comments Accepted at WACV 2026 WVAQ

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09223 2026-03-27 cs.LG cs.AI

Hierarchical Adaptive networks with Task vectors for Test-Time Adaptation

具有任务向量的分层自适应网络用于测试时适应

Sameer Ambekar, Marta Hasny, Laura Daza, Daniel M. Lang, Julia A. Schnabel

机构 * School of Computation, Information and Technology, Technical University of Munich(慕尼黑技术大学计算信息学院) Institute of Machine Learning in Biomedical Imaging, Helmholtz Munich, Germany(生物医学成像中的机器学习研究所,海德堡慕尼黑,德国) School of Biomedical Engineering and Imaging Sciences, King’s College London, UK(国王学院伦敦生物医学工程与成像科学学院,英国) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心(MCML)) relAI – Konrad Zuse School of Excellence in Reliable AI(relAI – 卓越可靠AI Konrad Zuse 学院)

AI总结 本文提出Hi-Vec网络,通过分层结构动态选择适应层,融合权重并采用门控机制提升测试时适应性能,验证了其在复杂分布偏移下的鲁棒性和有效性。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24209 2026-03-26 cs.CV cs.LG

HEART-PFL: Stable Personalized Federated Learning under Heterogeneity with Hierarchical Directional Alignment and Adversarial Knowledge Transfer

HEART-PFL:在异质性下稳定的个性化联邦学习:通过分层方向对齐和对抗性知识转移

Minjun Kim, Minje Kim

机构 * Promedius Inc.(Promedius公司)

AI总结 HEART-PFL通过分层方向对齐和对抗性知识转移,提升在异质分布下的个性化联邦学习效果,实现高准确率和稳定性,适用于CIFAR-100、Flowers-102和Caltech-101数据集。

Comments Accepted at WACV 2026. 8 pages, 7 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01583 2026-03-24 cs.CV

3DSceneEditor: Controllable 3D Scene Editing with Gaussian Splatting

3DSceneEditor: 基于高斯散射的可控3D场景编辑

Ziyang Yan, Yihua Shao, Minwen Liao, Siyu Chen, Nan Wang, Muyuan Lin, Jenq-Neng Hwang, Hao Zhao, Fabio Remondino, Lei Li

机构 * D Optical Metrology unit, Fondazione Bruno Kessler(3D光学测绘单元,布鲁诺·凯斯勒基金会) Department of Information Engineering and Computer Science, University of Trento(信息工程与计算机科学系,特伦托大学)

AI总结 本文提出3DSceneEditor,通过高斯散射实现高效精准的3D场景编辑,整合实例分割模型与CLIP实现语义对齐,支持对象添加、重定位等操作,实验表明其在精度和效率上优于现有方法。

Comments Accepted by WACV 2026, Project Page: https://ziyangyan.github.io/3DSceneEditor

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01781 2026-03-24 cs.CV cs.AI cs.LG

Subimage Overlap Prediction: Task-Aligned Self-Supervised Pretraining For Semantic Segmentation In Remote Sensing Imagery

子图像重叠预测:面向遥感图像语义分割的任务对齐自监督预训练

Lakshay Sharma, Alex Marin

机构 * Instacart New York University(纽约大学) University of Washington(华盛顿大学)

AI总结 本文提出子图像重叠预测任务,通过少量预训练数据提升遥感图像语义分割性能,实现更快收敛和同等或更优的mIoU表现。

Comments Accepted at CV4EO Workshop at WACV 2026

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Workshops, 2026, pp. 1414-1423

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17374 2026-03-17 cs.CV

Revisiting Vision Language Foundations for No-Reference Image Quality Assessment

重新审视用于无参考图像质量评估的视觉语言基础

Ankit Yadav, Ta Duc Huy, Lingqiao Liu

机构 * Australian Institute for Machine Learning(澳大利亚机器学习研究所) The University of Adelaide(阿德莱德大学)

AI总结 本文系统评估了六个预训练模型在无参考图像质量评估中的表现,发现SigLIP2性能优异,且激活函数选择对模型泛化能力影响显著,引入可学习的激活选择机制,达到新状态的SRCC。

Comments 23 pages, 16 figures. Accepted at WACV 2026

Journal ref IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026, pp. 5416-5425

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12369 2026-03-16 cs.CV

Human Knowledge Integrated Multi-modal Learning for Single Source Domain Generalization

融合人类知识的多模态学习用于单源域泛化

Ayan Banerjee, Kuntal Thakur, Sandeep Gupta

机构 * Impact Lab, Arizona State University(影响实验室,亚利桑那州立大学)

AI总结 本文提出GenEval方法,结合基础模型与人类知识提升单源域泛化性能,在糖尿病视网膜病变和癫痫发作区检测中取得优异结果。

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026, pp. 2380-2391

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04102 2026-03-12 cs.CV cs.AI

DMS2F-HAD: A Dual-branch Mamba-based Spatial-Spectral Fusion Network for Hyperspectral Anomaly Detection

DMS2F-HAD: 一种基于双分支Mamba的空谱融合网络用于超光谱异常检测

Aayushma Pant, Lakpa Tamang, Tsz-Kwan Lee, Sunil Aryal

机构 * School of Information Technology, Deakin University(德肯大学信息科技学院)

AI总结 DMS2F-HAD通过双分支Mamba模型实现高效空谱融合,提升超光谱异常检测的精度与效率。

Comments This paper has been accepted in the WACV 2025 conference in algorithm track

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08998 2026-03-11 cs.CV

Diffusion-Based Authentication of Copy Detection Patterns: A Multimodal Framework with Printer Signature Conditioning

基于扩散的复制检测模式认证:一种多模态框架与打印机签名条件化

Bolutife Atoki, Iuliia Tkachenko, Bertrand Kerautret, Carlos Crispim-Junior

机构 * Université Lumière Lyon 2, CNRS, INSA Lyon, Universite Claude Bernard Lyon 1, LIRIS UMR5205(里摩日大学里昂2分校、法国国家科学研究中心、里昂国立应用科学学院、里昂大学克莱尔-贝尔纳分校、LIRIS UMR5205)

AI总结 本文提出基于扩散的多模态认证框架,利用打印机签名条件化技术,实现对复制检测模式的高效认证,优于传统方法和深度学习方法。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08202 2026-03-10 cs.CV cs.AI

MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data

MM-TS: 多模态温度和边距调度用于长尾数据的对比学习

Siarhei Sheludzko, Dhimitrios Duka, Bernt Schiele, Hilde Kuehne, Anna Kukleva

机构 * University of Bonn(波恩大学) MPI for Informatics, SIC(信息研究所) Tuebingen AI Center/University of Tuebingen(图宾根人工智能中心/图宾根大学) MIT-IBM Watson AI Lab(麻省理工-IBM Watson人工智能实验室)

AI总结 MM-TS通过动态温度和边距调度提升多模态对比学习在长尾数据上的性能,实现InfoNCE和最大边距方法的统一。

Comments 18 pages, 11 figures. Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08023 2026-03-10 cs.CV cs.AI cs.GR cs.SD

Not Like Transformers: Drop the Beat Representation for Dance Generation with Mamba-Based Diffusion Model

不同于Transformer:用基于Mamba的扩散模型生成舞蹈,摒弃节奏表示

Sangjune Park, Inhyeok Choi, Donghyeon Soon, Youngwoo Jeon, Kyungdon Joo

机构 * Ulsan National Institute of Science and Technology, South Korea(乌山国立科学技术研究院,韩国) Daegu Gyeongbuk Institute of Science and Technology, South Korea(大邱庆北科学技术研究院,韩国)

AI总结 本文提出基于Mamba的扩散模型,用于生成舞蹈,通过替代Transformer并引入基于高斯的节奏表示,有效生成合理舞蹈动作,保持从短到长舞蹈的一致性。

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07543 2026-03-10 cs.CV cs.MM

CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization

CONSTANT:通过补丁对比增强和风格感知量化实现高质量单次手写生成

Anh-Duy Le, Van-Linh Pham, Thanh-Nam Vo, Xuan Toan Mai, Tuan-Anh Tran

机构 * Viettel Artificial Intelligence and Data Services Center(越南 Viettel 人工智能与数据服务中心) Ho Chi Minh City University of Technology(胡志明市技术大学)

AI总结 CONSTANT通过补丁对比增强和风格感知量化,实现高质量单次手写生成,有效提升图像细节与风格适应性。

Comments Accepted as oral presentation at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14756 2026-03-10 cs.GR cs.CV

SceneEval: Evaluating Semantic Coherence in Text-Conditioned 3D Indoor Scene Synthesis

SceneEval: 评估文本条件3D室内场景合成中的语义连贯性

Hou In Ivan Tam, Hou In Derek Pun, Austin T. Wang, Angel X. Chang, Manolis Savva

机构 * Simon Fraser University(西蒙弗雷泽大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔人工智能研究所)

AI总结 SceneEval通过引入细粒度度量标准评估文本条件3D室内场景合成中的语义连贯性,揭示了现有方法的不足,推动更可控的场景生成研究。

Comments Accepted at WACV 2026 (Oral). Project page: https://3dlg-hcvc.github.io/SceneEval/ . Minor revisions for camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05747 2026-03-09 cs.CV cs.RO

FlyPose: Towards Robust Human Pose Estimation From Aerial Views

FlyPose:迈向从空中视角实现鲁棒的人体姿态估计

Hassaan Farooq, Marvin Brenner, Peter Stütz

机构 * Universität der Bundeswehr Munich(联邦国防军大学慕尼黑)

AI总结 FlyPose通过多数据集训练实现了在航空图像中更准确的人体姿态估计,并在多个测试集上提升了检测和姿态估计的性能,同时具备低延迟和实时部署能力。

Comments 11 pages, 9 figures, IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026, pp. 8617-8627

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08445 2026-03-09 cs.CV cs.LG

Uncertainty-Aware Subset Selection for Robust Visual Explainability under Distribution Shifts

考虑不确定性的子集选择以在分布偏移下实现稳健的视觉可解释性

Madhav Gupta, Vishak Prasad C, Ganesh Ramakrishnan

机构 * Indian Institute of Technology Bombay(印度理工学院孟买学院)

AI总结 本文提出一种结合子模子集选择与层间梯度不确定性估计的方法,以提高在分布偏移下的视觉可解释性鲁棒性和保真度。

Comments Accepted to the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15048 2026-03-09 cs.CV cs.AI cs.LG cs.MM

Make VLM Recognize Visual Hallucination on Cartoon Character Image with Pose Information

使VLM在卡通角色图像上通过姿态信息识别视觉幻觉

Bumsoo Kim, Wonseop Shin, Kyuchul Lee, Yonghoon Jung, Sanghyun Seo

机构 * Chung-Ang University(Chung-Ang 大学) Coupang(韩国Coupang)

AI总结 本文提出基于姿态信息的VLM幻觉检测方法,通过上下文学习提升识别准确率,实验表明在卡通角色图像中幻觉检测效果提升50%-80%。

Comments Accepted at WACV 2025, Project page: https://gh-bumsookim.github.io/Cartoon-Hallucinations-Detection/. (Fixed typos)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04958 2026-03-06 cs.CV cs.GR

Revisiting an Old Perspective Projection for Monocular 3D Morphable Models Regression

重新审视一种旧的透视投影用于单目3D可变形模型回归

Toby Chong, Ryota Nakajima

机构 * TOEI Company(TOEI公司)

AI总结 本文提出一种改进的透视投影模型,用于提升近距离面部视频中单目3D可变形模型回归的准确性与稳定性。

Comments WACV 2026, https://zukunfcs.github.io/RevisitingAnOldPerspective/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04380 2026-03-05 cs.CV cs.CL

TaxonRL: Reinforcement Learning with Intermediate Rewards for Interpretable Fine-Grained Visual Reasoning

TaxonRL: 用于可解释细粒度视觉推理的强化学习

Maximilian von Klinski, Maximilian Schall

机构 * Hasso Plattner Institute University of Potsdam(波茨坦大学)

AI总结 TaxonRL通过强化学习和中间奖励实现结构化层次推理,提升细粒度视觉分类的准确性和可解释性。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17526 2026-03-05 cs.CV

Beyond the Encoder: Joint Encoder-Decoder Contrastive Pre-Training Improves Dense Prediction

超越编码器:联合编码器-解码器对比预训练改进密集预测

Sébastien Quetin, Tapotosh Ghosh, Farhad Maleki

机构 * McGill University(麦吉尔大学) University of Calgary(卡尔加里大学)

AI总结 DeCon通过联合编码器-解码器对比预训练提升密集预测性能,实现多个任务的先进表现。

Journal ref https://openaccess.thecvf.com/content/WACV2026/html/Quetin_Beyond_the_Encoder_Joint_Encoder-Decoder_Contrastive_Pre-Training_Improves_Dense_Prediction_WACV_2026_paper.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19888 2026-03-05 cs.CV cs.LG

FlowCLAS: Enhancing Normalizing Flow Via Contrastive Learning For Anomaly Segmentation

通过对比学习增强归一化流用于异常分割

Chang Won Lee, Selina Leveugle, Svetlana Stolpner, Chris Langley, Paul Grouchy, Jonathan Kelly, Steven L. Waslander

机构 * University of Toronto(多伦多大学) MDA Space(MDA空间)

AI总结 FlowCLAS通过对比学习增强归一化流,有效提升异常分割性能,达到多个机器人异常分割基准的最先进水平。

Comments WACV 2026 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01864 2026-03-03 cs.CV cs.RO

Streaming Real-Time Trajectory Prediction Using Endpoint-Aware Modeling

基于端点意识建模的流式实时轨迹预测

Alexander Prutsch, David Schinagl, Horst Possegger

机构 * Institute of Visual Computing, Graz University of Technology(视觉计算研究所,格拉茨技术大学)

AI总结 本文提出了一种基于端点意识建模的轻量高效流式轨迹预测方法,通过连续时间上下文传播提升预测精度并降低推理延迟,适用于自动驾驶场景。

Comments WACV 2026 Oral. Project Page at https://a-pru.github.io/seam/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00157 2026-03-03 cs.CV

FujiView: Multimodal Late-Fusion for Predicting Scenic Visibility

FujiView: 多模态晚期融合用于预测风景可见性

Bryceton Bible, Shah Md Nehal Hasnaeen, Hairong Qi

机构 * University of Tennessee, Knoxville(田纳西大学,科文克顿)

AI总结 FujiView通过融合摄像头图像与气象数据,实现风景可见性的多模态预测,展示了在短期和长期预测中的不同方法效果。

Comments 9 pages (including references), 8 figures, 2 tables. Accepted to the IEEE/CVF WACV 2026 proceedings. Introduces a large human-labeled Mount Fuji visibility dataset; public release forthcoming

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20816 2026-02-27 cs.CV cs.AI

MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval

具有长度感知DETR的MomentMix增强技术用于时间鲁棒的时刻检索

Seojeong Park, Jiho Choi, Kyungjune Baek, Hyunjung Shim

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院) Sejong University(世宗大学)

AI总结 本文提出Length-Aware Decoder和MomentMix增强技术,通过改进短时刻定位能力,在视频时刻检索任务中取得最佳性能。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22052 2026-02-26 cs.CV

AutoSew: A Geometric Approach to Stitching Prediction with Graph Neural Networks

AutoSew:基于几何方法的缝合预测与图神经网络

Pablo Ríos-Navarro, Elena Garces, Jorge Lopez-Moreno

机构 * Universidad Rey Juan Carlos(雷昂·卡洛斯大学) Adobe Research(Adobe研究)

AI总结 AutoSew通过几何方法和图神经网络实现自动缝合预测,以高精度完成服装组装任务。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13065 2026-02-26 cs.CV

RobustGait: Robustness Analysis for Appearance Based Gait Recognition

RobustGait: 基于外观的步态识别的鲁棒性分析

Reeshoon Sayera, Akash Kumar, Sirshapan Mitra, Prudvi Kamtam, Yogesh S Rawat

机构 * University of Central Florida(中央佛罗里达大学)

AI总结 RobustGait通过评估基于外观的步态识别系统在不同扰动和场景下的鲁棒性,揭示了噪声影响、轮廓提取偏差及架构设计的重要性,并提出改进策略以提升系统部署能力。

Comments IEEE WACV'26 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏