arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

IEEE TPAMI

IEEE Transactions on Pattern Analysis and Machine Intelligence · 期刊 · Computer Vision

共收录 20
2608.13104 2026-08-14 cs.CV 新提交

Online Learning of Correspondences between Images

图像间对应关系的在线学习

Michael Felsberg, Fredrik Larsson, Johan Wiklund, Niclas Wadströmer, Jörgen Ahlberg

AI总结 该研究提出一种基于奈曼卡方散度的在线迭代学习方法,用于解决通用成像几何下图像序列间的点对应问题,算法实时运行且性能优于现有方法。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, ISSN 0162-8828, E-ISSN 1939-3539, Vol. 35, no 1, p. 118-129

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12773 2026-08-14 cs.CV cs.LG eess.IV 新提交

CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

CW-BASS v2:基于基础模型教师的半监督分割中感知饱和度的伪标签选择

Ebenezer Tarubinga

机构 * Ebenworks Systems(埃本沃克斯系统公司)

AI总结 CW-BASS v2是一种感知饱和度的伪标签选择方法,它针对DINOv2教师模型的置信饱和问题,结合预留校准等技术,在6个基准数据集上恢复UniMatch V2操作点并提升性能。

Comments Submitted to IEEE TPAMI. 22 pages, 11 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09986 2026-08-12 cs.AI cs.LG 新提交

MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis

MIDAS:面向不完整多模态情感分析的基于不确定性感知融合的互信息解缠方法

Yuhua Wen, Yingying Zhou, Qifei Li, Yingming Gao, Zhengqi Wen, Jianhua Tao, Ya Li

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Zhongguancun Academy(中关村学院) Tsinghua University(清华大学)

AI总结 本研究针对多模态情感分析中模态不完整的问题,提出MIDAS框架,通过互信息解缠与不确定性感知融合实现鲁棒的多模态表示,在多个数据集上取得优于基线的性能。

Comments Accepted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09438 2026-08-11 cs.CV 新提交

Unveiling the Secret of AdaLN-Zero in Diffusion Transformer

揭示扩散Transformer中AdaLN-Zero的秘密

Jie Zhu, Mingyu Ding, Boqiang Duan, Leye Wang, Jingdong Wang

机构 * Peking University(北京大学) UC Berkeley(加州大学伯克利分校) Baidu(百度)

AI总结 本研究探究扩散Transformer(DiT)中AdaLN-Zero性能优于AdaLN的原因,发现零初始化是关键要素,提出AdaLN-Gaussian初始化策略与SE-adaLN-Zero机制,经多数据集实验验证其有效性与泛化性。

Comments Accept by IEEE TPAMI 2026, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05042 2026-08-06 cs.RO 新提交

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

BridgeVLA++:一种面向三维操作的数据高效、可泛化且内存增强的视觉-语言-动作框架

Peiyan Li, Yuze Zhu, Yixiang Chen, Qisen Ma, Yuan Xu, Jiabing Yang, He Guan, Yan Huang, Hongtao Wu, Xiao Ma, Tao Kong, Liang Wang, Tieniu Tan

机构 * New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室(NLPR)) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) FiveAges ByteDance Seed(字节跳动种子实验室)

AI总结 本研究提出内存增强的三维VLA框架BridgeVLA++,通过新增时空记忆架构,在保留原模型数据效率与泛化能力的同时,提升了记忆相关操作性能,且在多任务与真实平台上验证了其有效性。

Comments This work has been submitted to the IEEE TPAMI for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04106 2026-08-06 cs.CV eess.IV 新提交

LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching

LoRetta:面向全球尺度遥感密集图像匹配的基础模型与大规模数据集

Siwei Yu, Han Guo, Zhenwei Shi, Zhengxia Zou

机构 * Beihang University(北京航空航天大学)

AI总结 针对全球尺度遥感密集图像匹配的挑战,本文提出结合可匹配性感知仿射定位与引导式密集配准的基础模型LoRetta,并构建含原生可匹配性标签的LEVIR-GM基准,实验表明其性能优于现有模型且迁移性良好。

Comments 17 pages, 12 figures, 6 tables. Submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence. Project page: https://siweiyu.com/work/loretta/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25926 2026-07-29 cs.CV cs.AI 新提交

Face De-Identification: A Domain-Centric Survey from Capture to Processing

面部去识别:从采集到处理的以领域为中心的综述

Hui Wei, Hao Yu, Guoying Zhao

机构 * ELLIS Institute Finland(芬兰埃利斯研究所) Center for Machine Vision and Signal Analysis, University of Oulu(奥卢大学机器视觉与信号分析中心)

AI总结 该综述以领域为中心,涵盖面部去识别从采集到处理的完整数据管道,系统分析各阶段方法、进展与挑战,回顾评估协议,确定开放问题与新兴方向,为该领域未来工作提供指导。

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). Github repository: https://github.com/CV-AC/Awesome-FaceDe-ID

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23684 2026-07-28 cs.RO 新提交

Towards Ultrafast Depth Sensing Via Active Event-based Stereo Vision

通过基于主动事件的立体视觉实现超快速深度感知

Jianing Li, Yunjian Zhang, Haiqian Han, Kangyao Huang, Xiangyang Ji

机构 * School of Computer Science, Peking University(北京大学计算机科学学院) Peng Cheng Laboratory(鹏城实验室) Department of Automation, Tsinghua University(清华大学自动化系) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)

AI总结 针对传统主动立体系统在快速运动场景的问题,提出基于主动事件的立体视觉,构建数据集,创新设计ActiveEventNet+网络,性能优于现有方法,降低计算复杂度,能实现高速实时处理,为高速深度感知相机系统设计提供新思路。

Comments Accepted by TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21057 2026-07-24 cs.CV 新提交

Achieving Text-based Person Retrieval with Any Granularity

实现任意粒度的基于文本的行人检索

Jialong Zuo, Hanyu Zhou, Dongyue Wu, Yongtai Deng, Mengdan Tan, Nong Sang, Changxin Gao, Xiang Bai

机构 * National Key Laboratory of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院多光谱信息智能处理技术国家重点实验室) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)

AI总结 该研究针对基于文本的行人检索中查询粒度不确定问题提出新范式,构建多粒度数据集和评估基准,提出CMAM框架,通过多种策略实现粒度感知检索,实验证明其性能优于现有方法,为行人检索系统奠定基础。

Comments TPAMI-2026 Accepted Paper

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026, pp. 1-18

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20866 2026-07-24 cs.CV 新提交

Agentic Designer: Progressive Multi-Agent Collaboration for Structure-Aware Interior Layout Generation

智能设计师:用于结构感知室内布局生成的渐进式多智能体协作

Zhijing Yang, Haocheng Lin, Zhihua Xu, Haojie Li, Keze Wang, Liang Lin, Tianshui Chen

机构 * Guangdong University of Technology(广东工业大学) South China University of Technology(华南理工大学)

AI总结 针对室内布局生成难题,提出智能设计师这一渐进式多智能体框架,将布局生成视为迭代约束验证决策过程,通过渐进共识机制协调三个智能体,经实验验证其显著优于现有方法,提升了结构遵循和功能设计连贯性。

Comments TPAMI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16828 2026-07-21 cs.CV 新提交

UniNDM: A Unified Noise-driven Detection and Mitigation Framework Against Sexual Content in Text-to-Image Generation

UniNDM:文本到图像生成中针对性内容的统一噪声驱动检测与缓解框架

Yao Huang, Yitong Sun, Huanran Chen, Ruochen Zhang, Shouwei Ruan, Ranjie Duan, Maoxun Yuan, Yinpeng Dong, Hui Xue, Xiaochun Cao, Xingxing Wei

机构 * Institute of Artificial Intelligence, State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室人工智能研究院) College of Artificial Intelligence, Tsinghua University(清华大学人工智能学院) Security Department, Alibaba Group(阿里巴巴集团安全部) School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-Sen University(中山大学深圳校区网络科学与技术学院)

AI总结 针对文本到图像生成易受隐式性提示影响的问题,提出UniNDM统一噪声驱动框架。利用早期预测噪声的可分离性开发轻量级检测器,引入噪声增强自适应负引导缓解问题,扩展到扩散变压器架构,实验显示比现有方法有显著改进。

Comments 18 pages, 10 figures, accepted by TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11914 2026-07-15 cs.NE cs.AI 新提交

Burst Spiking Neural Networks

突发脉冲神经网络

Jiahong Zhang, Sijun Shen, Man Yao, Han Xu, Mingqiang Huang, Yonghong Tian, Bo Xu, Guoqi Li

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) State Key Laboratory of Media Convergence and Communication, Communication University of China(中国传媒大学媒体融合与传播国家重点实验室) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) Peng Cheng Laboratory(鹏城实验室) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)

AI总结 研究SNN的准确性 - 鲁棒性问题,提出基于突发增强脉冲神经元和动态权重约束机制的BuSNN,通过理论分析和实验表明其在准确性、鲁棒性及低功耗方面优势显著,推进了SNN在相关应用中的可行性。

Comments 18 pages, 21 figures, 1 supplementary material PDF, submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09741 2026-07-14 cs.RO cs.AI cs.MA 新提交

SWIFT: A Small-World Interaction Framework for Flow-Aware Trajectory Prediction in Autonomous Driving

SWIFT:用于自动驾驶中流量感知轨迹预测的小世界交互框架

Chengyue Wang, Bin Rao, Haicheng Liao, Bonan Wang, Chengzhong Xu, Zhenning Li

机构 * University of Macau(澳门大学) State Key Laboratory of Internet of Things for Smart City, University of Macau(澳门大学智慧城市物联网国家重点实验室) Department of Civil and Environmental Engineering, University of Macau(澳门大学土木与环境工程系)

AI总结 研究自动驾驶中轨迹预测问题,提出SWIFT框架,结合小世界网络与交通流理论,通过特定网络和模块引入结构偏差与增强推理,实验证明其在多方面优于基线,展现出结构感知设计的有效性。

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05783 2026-07-08 cs.CV 新提交

Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective

从车道感知角度评估自动驾驶对环境错觉的鲁棒性:基准测试

Tianyuan Zhang, Xianglong Liu, Aishan Liu, Lu Wang, Yitong Zhang, Peng Yue, Mingchuan Zhang, Siyuan Liang, Dacheng Tao

机构 * SKLCCSE, the School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院软件安全技术与工程北京市重点实验室) the School of Cyber Science and Technology, Sun Yat-sen University(中山大学网络空间科学与技术学院) Henan University of Science and Technology(河南科技大学) the School of Computing, National University of Singapore(新加坡国立大学计算学院) College of Computing & Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

AI总结 研究自动驾驶对环境错觉的鲁棒性,聚焦传统车道检测和视觉语言模型系统。引入基准LanEvil++并评估,发现错觉严重影响性能,阴影干扰最大。提出多模态错觉防御方法MIDA,有效提升了模型在挑战性条件下的鲁棒性。

Comments Accepted by IEEE TPAMI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05464 2026-07-08 cs.LG cs.AI 新提交

Learnable Weighting of Intra-Attribute Distances for Categorical Data Clustering with Nominal and Ordinal Attributes

用于具有标称和有序属性的分类数据聚类的可学习属性内距离加权

Yiqun Zhang, Yiu-ming Cheung

机构 * IEEE

AI总结 研究用于分类数据聚类的距离度量,提出统一测量标称和有序属性内距离的新方法及新聚类算法,将属性内距离权重学习和数据划分整合为单一范式,实验验证了算法有效性。

Comments 16 pages, 11 figures

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04650 2026-07-07 stat.ML cs.LG 新提交

Decomposition for Bayesian Networks: Local and Parallel Inference

贝叶斯网络的分解:局部与并行推理

Pei Heng, Xinyi Hu, Yi Sun

机构 * School of Mathematics and Statistics and KLAS, Northeast Normal University(数学与统计学学院及KLAS,东北师范大学) College of Mathematics and System Sciences, Xinjiang University(数学与系统科学学院,新疆大学) Institute of Statistics and Data Science, Xinjiang University of Finance and Economics(统计与数据科学学院,新疆财经大学)

AI总结 针对高维贝叶斯网络精确推理复杂度随规模指数增长的问题,提出基于有向凸子图的分解框架与最小d-分解树,实现并行推理,效率优于连接树法且保留推理精度。

Comments 13 pages, 5 figures,Code available at https://github.com/Balance-H/Decomposition-for-BNs

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01784 2026-07-03 cs.CV 新提交

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video

SpaceEra++: 面向视频中3D空间推理的统一框架

Weili Guan, Haoyu Zhang, Meng Liu, Qianlong Xiang, Yaowei Wang, Liqiang Nie

机构 * School of Information Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)信息科学与技术学院) Shenzhen Loop Area Institute(深圳河套学院) School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)计算机科学与技术学院) Pengcheng Laboratory(鹏城实验室) School of Computer Science and Technology, Shandong Jianzhu University(山东建筑大学计算机科学与技术学院) Zhongguancun Academy(中关村学院) City University of Hong Kong(香港城市大学)

AI总结 提出SpaceEra++框架,通过ScenePick帧采样策略缓解输入不足,并利用SpaceAlign对齐绝对坐标与相对空间关系增强推理,在多个基准上超越强基线。

Comments Accepted by IEEE TPAMI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21289 2026-06-23 cs.LG 新提交

Reconstructing Randomly Masked Spectra Helps DNNs Identify Discriminant Wavenumbers

重构随机掩蔽光谱有助于深度神经网络识别判别波数

Yingying Wu, Jinchao Liu, Yan Wang, Stuart Gibson, Margarita Osadchy, Yongchun Fang

机构 * Institute of Robotics and Automatic Information System (IRAIS), College of Artificial Intelligence, Nankai University(机器人与自动信息系统研究所(IRAIS),人工智能学院,南开大学) VisionMetric Ltd(VisionMetric有限公司) School of Physics and Astronomy, University of Kent(物理与天文学院,肯特大学)

AI总结 提出任务增强网络TeaNet,通过重构随机掩蔽光谱生成增广样本,同时训练分类模型,在合成和真实数据集上优于CNN,并能更好识别判别波数。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 5, pp. 3845-3861, May 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19316 2026-06-18 cs.CV 新提交

NeuMesh++: Towards Versatile and Efficient Volumetric Editing with Disentangled Neural Mesh-based Implicit Field

NeuMesh++:基于解耦神经网格隐式场的多功能高效体积编辑

Chong Bao, Yuan Li, Bangbang Yang, Yujun Shen, Hujun Bao, Zhaopeng Cui, Yinda Zhang, Guofeng Zhang

机构 * State Key Lab of CAD&CG, College of Computer Science, Zhejiang University(浙江大学计算机科学学院CAD&CG国家重点实验室) Ant Research(蚂蚁研究院) Google(谷歌) ByteDance(字节跳动)

AI总结 提出一种基于网格顶点的解耦神经辐射场表示,实现几何、纹理和语义引导的高效体积编辑,包括网格引导几何编辑、纹理交换填充绘制及语义编辑。

Comments TPAMI 2025; Project Page: https://zju3dv.github.io/neumeshplusplus/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13580 2026-06-12 cs.CV cs.AI 新提交

EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution

EvTexture++: 事件驱动的视频超分辨率纹理增强

Dachun Kai, Jiayao Lu, Yueyi Zhang, Xiaoyan Sun

机构 * MOE Key Laboratory of Brain-Inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学,脑启发智能感知与认知教育部重点实验室) Midea Group(美的集团)

AI总结 提出首个事件驱动的视频超分辨率纹理增强框架EvTexture++,利用事件的高频时空细节逐步恢复纹理,并通过时间纹理对齐模块增强帧间一致性,在多个数据集上达到最优性能。

Comments IEEE TPAMI 2026. Extended version of arXiv:2406.13457 (ICML 2024). Project page: https://dachunkai.github.io/evtexture-project-page/

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 6, pp. 6642-6659, June 2026

详情

展开后加载摘要…

URL PDF HTML 收藏