arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

IEEE TPAMI

IEEE Transactions on Pattern Analysis and Machine Intelligence · 期刊 · Computer Vision

共收录 1562
2408.00001 2026-07-08 cs.CV cs.AI cs.CY 版本更新

Replication in Visual Diffusion Models: A Survey and Outlook

视觉扩散模型中的复制:一项综述与展望

Wenhao Wang, Yifan Sun, Zongxin Yang, Zhengdong Hu, Zhentao Tan, Yi Yang

机构 * Australian Artificial Intelligence Institute, University of Technology Sydney(澳大利亚人工智能研究所,悉尼技术大学) Baidu Inc.(百度公司) College of Computer Science and Technology, Zhejiang University(计算机科学与技术学院,浙江大学)

AI总结 本文综述视觉扩散模型中的复制现象,将现有研究分类为揭示、理解和缓解该现象的方法,还回顾其现实影响,讨论了检测和基准测试复制的挑战及未来方向,助力研究人员和从业者理解AI技术与社会公益交叉点。

Comments Accepted by TPAMI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04650 2026-07-07 stat.ML cs.LG 新提交

Decomposition for Bayesian Networks: Local and Parallel Inference

贝叶斯网络的分解:局部与并行推理

Pei Heng, Xinyi Hu, Yi Sun

机构 * School of Mathematics and Statistics and KLAS, Northeast Normal University(数学与统计学学院及KLAS,东北师范大学) College of Mathematics and System Sciences, Xinjiang University(数学与系统科学学院,新疆大学) Institute of Statistics and Data Science, Xinjiang University of Finance and Economics(统计与数据科学学院,新疆财经大学)

AI总结 针对高维贝叶斯网络精确推理复杂度随规模指数增长的问题,提出基于有向凸子图的分解框架与最小d-分解树,实现并行推理,效率优于连接树法且保留推理精度。

Comments 13 pages, 5 figures,Code available at https://github.com/Balance-H/Decomposition-for-BNs

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02291 2026-07-07 cs.LG cs.AI

FAIR-Pruner: A Flexible Framework for Automatic Layer-Wise Pruning via Tolerance of Difference

FAIR-Pruner: 一种通过差异容忍性实现自动分层剪枝的灵活框架

Chenqing Lin, Mostafa Hussien, Chengyao Yu, Bingyi Jing, Ruixing Ming, Kim Khoa Nguyen, Mohamed Cheriet

机构 * School of Statistics and Mathematics, Zhejiang Gongshang University(浙江工商大学统计与数学学院) École de technologie supérieure (ÉTS), Université du Québec(魁北克大学埃克森技术学院) Southern University of Science and Technology(南方科技大学)

AI总结 本文提出FAIR-Pruner,一种无需搜索的自适应分层结构化剪枝框架,通过引入差异容忍度(ToD)来实现非均匀的分层剪枝深度,从而在多个数据集和模型上实现了良好的准确率-压缩率权衡。

Comments Submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.15277 2026-07-07 cs.CV 版本更新

EPMF: Efficient Perception-aware Multi-sensor Fusion for 3D Semantic Segmentation

EPMF: 高效感知感知多传感器融合用于3D语义分割

Mingkui Tan, Zhuangwei Zhuang, Sitao Chen, Rong Li, Kui Jia, Qicheng Wang, Yuanqing Li

机构 * School of Software Engineering, South China University of Technology(南方科技大学软件工程学院) Pazhou Laboratory, Guangzhou, China(广州帕佐实验室) School of Electronic and Information Engineering, South China University of Technology(南方科技大学电子与信息工程学院) Minieye, Shenzhen, Guangdong, China(深圳Minieye公司)

AI总结 提出一种高效感知多传感器融合方法EPMF,通过透视投影对齐点云与RGB图像,利用残差融合模块和感知损失提升3D语义分割性能,在nuScenes上mIoU超越RangeFormer 0.9%。

Comments 16 pages, 12 figures, 14 tables, IEEE TPAMI 2024, extended version of the ICCV2021 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01784 2026-07-03 cs.CV 新提交

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video

SpaceEra++: 面向视频中3D空间推理的统一框架

Weili Guan, Haoyu Zhang, Meng Liu, Qianlong Xiang, Yaowei Wang, Liqiang Nie

机构 * School of Information Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)信息科学与技术学院) Shenzhen Loop Area Institute(深圳河套学院) School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)计算机科学与技术学院) Pengcheng Laboratory(鹏城实验室) School of Computer Science and Technology, Shandong Jianzhu University(山东建筑大学计算机科学与技术学院) Zhongguancun Academy(中关村学院) City University of Hong Kong(香港城市大学)

AI总结 提出SpaceEra++框架,通过ScenePick帧采样策略缓解输入不足,并利用SpaceAlign对齐绝对坐标与相对空间关系增强推理,在多个基准上超越强基线。

Comments Accepted by IEEE TPAMI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18060 2026-06-30 cs.CV

Semantic Correspondence: Unified Benchmarking and a Strong Baseline

语义对应:统一的基准测试与强大的基线

Kaiyan Zhang, Xinghui Li, Jingyi Lu, Kai Han

机构 * The University of Hong Kong(香港大学)

AI总结 本文首次全面调研语义对应方法,提出分类体系并汇总多基准结果,提出高性能基线,为未来研究奠定基础。

Comments accepted by TPAMI 2025

Journal ref IEEE Trans. Pattern Anal. Mach. Intell. 48, no. 3 (2026) 3911-3930

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00636 2026-06-24 cs.LG cs.SY eess.SY

On the Equilibrium between Feasible Zone and Uncertain Model in Safe Exploration

在安全探索中可行区域与不确定模型之间的平衡

Yujie Yang, Zhilong Zheng, Shengbo Eben Li

机构 * School of Vehicle and Mobility and State Key Lab of Intelligent Green Vehicle and Mobility, Tsinghua University, Beijing, 100084, China(车辆与移动系统学院和智能绿色车辆与移动系统国家重点实验室,清华大学,北京,100084,中国)

AI总结 本文提出安全平衡探索框架SEE,通过交替寻找最大可行区域和最不确定的模型,证明其在安全探索中达到平衡。实验显示算法能零约束违规扩展可行区域。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 48(7), 8344-8360 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21289 2026-06-23 cs.LG 新提交

Reconstructing Randomly Masked Spectra Helps DNNs Identify Discriminant Wavenumbers

重构随机掩蔽光谱有助于深度神经网络识别判别波数

Yingying Wu, Jinchao Liu, Yan Wang, Stuart Gibson, Margarita Osadchy, Yongchun Fang

机构 * Institute of Robotics and Automatic Information System (IRAIS), College of Artificial Intelligence, Nankai University(机器人与自动信息系统研究所(IRAIS),人工智能学院,南开大学) VisionMetric Ltd(VisionMetric有限公司) School of Physics and Astronomy, University of Kent(物理与天文学院,肯特大学)

AI总结 提出任务增强网络TeaNet,通过重构随机掩蔽光谱生成增广样本,同时训练分类模型,在合成和真实数据集上优于CNN,并能更好识别判别波数。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 5, pp. 3845-3861, May 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17015 2026-06-19 cs.AI cs.MA cs.RO 版本更新

UniMM: A Unified Mixture Model Framework for Multi-Agent Simulation

UniMM:一种用于多智能体仿真的统一混合模型框架

Longzhong Lin, Xuewu Lin, Kechun Xu, Haojian Lu, Lichao Huang, Rong Xiong, Yue Wang

机构 * Zhejiang University(浙江大学) Horizon Robotics

AI总结 提出UniMM框架统一回归混合模型与离散NTP模型,通过闭环样本生成缓解分布偏移,并在WOSAC基准上取得最优性能。

Comments Accepted author manuscript. The version of record has been published in IEEE Transactions on Pattern Analysis and Machine Intelligence

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, Early Access, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19316 2026-06-18 cs.CV 新提交

NeuMesh++: Towards Versatile and Efficient Volumetric Editing with Disentangled Neural Mesh-based Implicit Field

NeuMesh++:基于解耦神经网格隐式场的多功能高效体积编辑

Chong Bao, Yuan Li, Bangbang Yang, Yujun Shen, Hujun Bao, Zhaopeng Cui, Yinda Zhang, Guofeng Zhang

机构 * State Key Lab of CAD&CG, College of Computer Science, Zhejiang University(浙江大学计算机科学学院CAD&CG国家重点实验室) Ant Research(蚂蚁研究院) Google(谷歌) ByteDance(字节跳动)

AI总结 提出一种基于网格顶点的解耦神经辐射场表示,实现几何、纹理和语义引导的高效体积编辑,包括网格引导几何编辑、纹理交换填充绘制及语义编辑。

Comments TPAMI 2025; Project Page: https://zju3dv.github.io/neumeshplusplus/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15595 2026-06-18 cs.AI cs.CL cs.LG 版本更新

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

直接偏好优化综述:数据集、理论、变体及应用

Wenyi Xiao, Zechuan Wang, Leilei Gan, Shuai Zhao, Zongrui Li, Ruirui Lei, Wanggui He, Luu Anh Tuan, Long Chen, Hao Jiang, Zhou Zhao, Fei Wu

机构 * Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学) Alibaba Group(阿里巴巴集团)

AI总结 综述直接偏好优化(DPO)在理论、变体、数据集和应用方面的进展,指出其作为RL-free替代方案的潜力与局限,并提出未来研究方向。

Comments Accepted by TPAMI 2026. Project page: https://github.com/Mr-Loevan/DPO-Survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08038 2026-06-18 cs.LG cs.AI cs.CV 版本更新

Generalized Kullback-Leibler Divergence Loss

广义Kullback-Leibler散度损失

Jiequan Cui, Beier Zhu, Qingshan Xu, Zhuotao Tian, Xiaojuan Qi, Bei Yu, Hanwang Zhang, Richang Hong

机构 * Hefei University of Technology(合肥工业大学) University of Science and Technology of China(中国科学技术大学) Nanyang Technological University(南洋理工大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

AI总结 本文提出广义KL散度损失,通过解耦KL损失为加权MSE和交叉熵损失,并引入非对称优化修正和类别全局信息,在对抗训练和知识蒸馏中取得SOTA性能。

Comments TPAMI 2026, extension of our NeurIPS paper "Decoupled Kullback-Leibler Divergence Loss". arXiv admin note: substantial text overlap with arXiv:2305.13948

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09977 2026-06-16 cs.CV 版本更新

A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation

3D高斯泼溅应用综述:分割、编辑与生成

Shuting He, Peilin Ji, Yitong Yang, Changshuo Wang, Jiayi Ji, Yinglin Wang, Henghui Ding

机构 * Shanghai University of Finance and Economics(上海财经大学) University College London(伦敦大学学院) Xiamen University(厦门大学) Fudan University(复旦大学)

AI总结 综述3D高斯泼溅在分割、编辑和生成三大任务中的应用,总结代表性方法、监督策略和学习范式,并分析公共基准上的比较结果。

Comments IEEE TPAMI, GitHub Repo: https://github.com/heshuting555/Awesome-3DGS-Applications

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02288 2026-06-16 cs.CV cs.LG 版本更新

Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization

基于语言引导与表示对齐的提示解缠用于域泛化

De Cheng, Zhipeng Xu, Xinyang Jiang, Dongsheng Li, Nannan Wang, Xinbo Gao

机构 * School of Telecommunications Engineering, the State Key Laboratory of Integrated Services Networks (ISN), Xidian University, Xi’an, China(电信工程学院、集成服务网络国家重点实验室(ISN)、西安电子科技大学) Microsoft Research Asia, Shanghai, China(微软亚洲研究院,上海,中国)

AI总结 提出利用大语言模型自动解缠文本提示,并引入最差显式表示对齐,结合抽象提示增强源域多样性,实现域不变视觉表示学习,在多个基准上超越现有方法。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 6, pp. 6799-6816, June 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13580 2026-06-12 cs.CV cs.AI 新提交

EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution

EvTexture++: 事件驱动的视频超分辨率纹理增强

Dachun Kai, Jiayao Lu, Yueyi Zhang, Xiaoyan Sun

机构 * MOE Key Laboratory of Brain-Inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学,脑启发智能感知与认知教育部重点实验室) Midea Group(美的集团)

AI总结 提出首个事件驱动的视频超分辨率纹理增强框架EvTexture++,利用事件的高频时空细节逐步恢复纹理,并通过时间纹理对齐模块增强帧间一致性,在多个数据集上达到最优性能。

Comments IEEE TPAMI 2026. Extended version of arXiv:2406.13457 (ICML 2024). Project page: https://dachunkai.github.io/evtexture-project-page/

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 6, pp. 6642-6659, June 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05911 2026-06-05 cs.SD cs.LG eess.AS

DBHN-Net: Dual-Branch Hybrid Neural Network For Low-Complexity Monaural Speech Enhancement

DBHN-Net: 低复杂度单声道语音增强的双分支混合神经网络

Cunhang Fan, Enrui Liu, Jing Zhou, Jian Kang, Jie Li, Andong Li, Jian Zhou, Zhao Lv, Xuelong Li

机构 * State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology, (School of Computer Science and Technology), Anhui University(光电信息获取与防护技术国家重点实验室(计算机科学与技术学院),安徽大学) China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd(中国电信人工智能技术(北京)有限公司) Institute of Acoustics, University of Chinese Academy of Sciences(中国科学院声学研究所) Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究所(TeleAI),中国电信,中国)

AI总结 提出一种结合ANN和SNN的双分支混合神经网络,通过BandSplit、TF-Mamba等模块降低计算复杂度,同时利用交互和融合模块保持性能,在三个公共数据集上实现平均7.5倍复杂度降低。

Comments This article has been accepted for publication in IEEE Transactions on Pattern Analysis and Machine Intelligence(TPAMI)

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence(TPAMI2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04797 2026-06-04 cs.CV cs.LG

Crafting Your Evolving Dreams: Concept-Incremental Versatile Customization

打造你不断演变的梦想:概念增量式多功能定制

Jiahua Dong, Wenqi Liang, Hongliu Li, Yang Cong, Duzhen Zhang, Hanbin Zhao, Henghui Ding, Yulun Zhang, Salman Khan, Fahad Shahbaz Khan

机构 * Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed大学人工智能学院) University of Trento(特伦托大学) Department of Civil and Environmental Engineering, The Hong Kong Polytechnic University(香港理工大学土木与环境工程系) South China University of Technology(华南理工大学) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Institute of Big Data, Fudan University(复旦大学大数据研究院) Shanghai Jiao Tong University(上海交通大学)

AI总结 提出持续可定制扩散模型(CCDM),通过属性解耦LoRA模块和相关性引导聚合策略解决灾难性遗忘,并结合可控区域上下文合成策略处理概念忽视,实现概念增量式多功能定制。

Comments Accepted to Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.00204 2026-06-04 cs.LG cs.AI cs.NA math.NA stat.ML

Scalar Quantization as Sparse Least Square Optimization

标量量化作为稀疏最小二乘优化

Chen Wang, Xiaomei Yang, Shaomin Fei, Kai Zhou, Xiaofeng Gong, Miao Du, Ruisen Luo

机构 * College of Electrical Engineering, Sichuan University(四川大学电气工程学院) Department of Computer Science, Rutgers University -- New Brunswick(罗格斯大学新布朗斯维广场分校计算机科学系) Engineering Practice Center, Chengdu University of Information Technology(成都信息科技大学工程实践中心)

AI总结 本文提出了一种基于稀疏最小二乘优化的新方法,用于解决标量量化中的问题,通过引入l1、l1+l2和l0正则化,改进了传统聚类方法的不足,提升了在位宽缩减场景下的性能。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.09879 2026-06-04 math.NA cs.NA eess.IV math.OC

L0TV: A Sparse Optimization Method for Impulse Noise Image Restoration

L0TV:一种用于脉冲噪声图像恢复的稀疏优化方法

Ganzhao Yuan, Bernard Ghanem

AI总结 本文提出了一种新的稀疏优化方法L0TV-PADMM,用于在图像恢复中去除脉冲噪声,通过使用ℓ0范数数据保真度和PADMM算法解决非凸非光滑优化问题,实验表明其优于现有方法。

Comments to appear in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.07140 2026-06-04 math.NA cs.NA

A Fast Frequent Directions Algorithm for Low Rank Approximation

一种用于低秩近似快速频繁方向算法

Dan Teng, Delin Chu

AI总结 本文提出了一种快速频繁方向算法,通过将稀疏子空间嵌入(SpEmb)随机算法植入频繁方向(FD)中,以提高低秩近似问题的计算效率,并在合成和真实数据集上验证了其有效性。

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
1502.02860 2026-06-04 stat.ML cs.LG cs.RO cs.SY eess.SY

Gaussian Processes for Data-Efficient Learning in Robotics and Control

高斯过程在机器人和控制中的数据高效学习

Marc Peter Deisenroth, Dieter Fox, Carl Edward Rasmussen

AI总结 本文提出基于高斯过程的非参数转移模型,通过提取更多数据信息加速学习,减少模型误差影响,实现高效自主学习。

Comments 20 pages, 29 figures; fixed a typo in equation on page 8

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 37, issue no 2, pages 408-423, February 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1605.01278 2026-06-04 stat.ML cs.LG cs.SY eess.SY math.DS math.PR

A Bayesian Approach to Policy Recognition and State Representation Learning

基于贝叶斯方法的策略识别与状态表示学习

Adrian Šošić, Abdelhak M. Zoubir, Heinz Koeppl

AI总结 本文提出一种贝叶斯方法,用于在不假设专家行为最优的情况下,学习任意随机专家策略,并推断专家使用的状态表示复杂度及任务相关的状态空间划分。

Comments 17 pages, 8 figures; ### Version 4 ### to appear in IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01604 2026-06-02 cs.CV

Paving the Way for Point Cloud Video Representation Learning Using A PDE Model

使用PDE模型为点云视频表示学习铺平道路

Zhuoxu Huang, Zhenkun Fan, Jungong Han, Josef Kittler

机构 * Department of Computer Science, Aberystwyth University(阿伯里斯يث大学计算机科学系) Department of Automation, Beijing National Research Center for Information Science and Technology, Tsinghua University(自动化系、北京信息科学与技术国家研究中心、清华大学) Department of Electrical Engineering, Surrey University(Surrey大学电子工程系)

AI总结 提出MotionPDE方法,通过将时空相关性学习建模为可解的偏微分方程(PDE),并利用对比学习结构优化,作为即插即用模块提升点云视频表示学习性能。

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (T-PAMI) in 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00310 2026-06-02 cs.CV cs.AI cs.LG

Beyond Visual Fidelity: Benchmarking Super-Resolution Models for Large-Scale Remote Sensing Imagery via Downstream Task Integration

超越视觉保真度:通过下游任务集成评估大规模遥感影像的超分辨率模型

Zhili Li, Kangyang Chai, Zhihao Wang, Xiaowei Jia, Yanhua Li, Gengchen Mai, Sergii Skakun, Dinesh Manocha, Yiqun Xie

机构 * University of Maryland(马里兰大学) University of Pittsburgh(匹兹堡大学) Worcester Polytechnic Institute(沃思利技术学院) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 针对现有超分辨率评估依赖PSNR/SSIM等保真度指标而忽略下游任务效用的问题,提出GeoSR-Bench基准数据集,集成土地覆盖分割、基础设施映射等下游任务,评估GAN、Transformer等9种SR模型在270种设置下的性能,发现保真度指标与任务性能弱相关甚至负相关。

Comments Under review at IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10219 2026-05-29 cs.AI

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models

认知支点与视觉锚定:揭示并纠正多模态推理模型中的幻觉

Zhe Qian, Yanbiao Ma, Zhuohan Ouyang, Zhonghua Wang, Zhongxing Xu, Fei Luo, Xinyu Liu, Zongyuan Ge, Yike Guo, Jungong Han

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学 龙城人工智能学院) South China Agricultural University(华南农业大学) South China Normal University(华南师范大学) Monash University(莫纳什大学) Jishou University(吉首大学) Hong Kong University of Science and Technology(香港科技大学) Tsinghua University(清华大学)

AI总结 针对多模态大推理模型在长链推理中易产生幻觉的问题,提出V-STAR训练范式,通过分层视觉注意力奖励和强制反思机制,将视觉锚定引入推理过程以减轻幻觉。

Comments TPAMI under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14161 2026-05-29 cs.LG

Promoting Generalization for Exact Solvers via Adversarial Instance Augmentation

通过对抗性实例增强促进精确求解器的泛化能力

Haoyang Liu, Yufei Kuang, Jie Wang, Xijun Li, Yongdong Zhang, Feng Wu

机构 * CAS Key Laboratory of Technology in GIPAS, University of Science and Technology of China(GIPAS技术CAS重点实验室,中国科学技术大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)

AI总结 针对学习型MILP求解器在未见实例上性能下降的问题,提出对抗性实例增强方法AdaSolver,通过将不可微的实例增强建模为上下文赌博机问题并联合对抗训练增强策略与求解器,显著提升基于模仿学习和强化学习的分支定界求解器的泛化能力。

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.09305 2026-05-28 cs.LG stat.ML

Gaussian RBF Centered Kernel Alignment (CKA) in the Large Bandwidth Limit

大带宽极限下的高斯RBF中心核对齐(CKA)

Sergio A. Alvarez

机构 * Boston College(波士顿学院)

AI总结 本文证明基于高斯RBF核的中心核对齐(CKA)在大带宽极限下收敛到线性CKA,并发现收敛起始对特征表示的几何形状敏感,表示偏心率限制了高斯CKA表现非线性的带宽范围。

Comments 11 pages, 3 figures

Journal ref IEEE TPAMI, vol. 45, issue 5, 01 May 2023, pages 6587-6593

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26933 2026-05-27 cs.CV

Leveraging Text-to-Image Diffusion Models for Unsupervised Visual Object Tracking

利用文本到图像扩散模型进行无监督视觉目标跟踪

Zhengbo Zhang, Zhigang Tu, Junsong Yuan, De Wen Soh, Bo Du

机构 * Information Systems Technology and Design Pillar, Singapore University of Technology and Design(新加坡科技设计大学信息系统技术与设计学院) State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(武汉大学测绘遥感信息工程国家重点实验室) Department of Computer Science and Engineering, University at Buffalo, State University of New York(纽约州立大学布法罗分校计算机科学与工程系) School of Computer Science, Wuhan University(武汉大学计算机学院)

AI总结 提出Diff-Tracking方法,利用预训练文本到图像扩散模型的跨注意力机制,通过初始提示学习器和在线提示更新器实现无监督目标跟踪。

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11941 2026-05-26 cs.CV cs.AI

DynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action Recognition

DynaPURLS: 基于骨架的零样本动作识别中部分感知表示的动态细化

Jingmin Zhu, Anqi Zhu, James Bailey, Jun Liu, Hossein Rahmani, Mohammed Bennamoun, Farid Boussaid, Qiuhong Ke

机构 * Monash University(莫纳什大学) Lancaster University(兰卡斯特大学) University of Western Australia(西澳大学)

AI总结 提出DynaPURLS框架,通过多尺度视觉-语义对应和动态细化模块,解决骨架零样本动作识别中的领域偏移问题,在三个基准数据集上取得最优结果。

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22973 2026-05-26 cs.CV

Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method

扩展以占据为中心的驾驶场景生成:数据集与方法

Bohan Li, Xin Jin, Hu Zhu, Hongsi Liu, Ruikai Li, Jiazhe Guo, Kaiwen Cai, Chao Ma, Yueming Jin, Hao Zhao, Xiaokang Yang, Wenjun Zeng

机构 * Shanghai Jiao Tong University(上海交通大学) Eastern Institute of Technology(东部技术研究院) School of Electronic Information and Electrical Engineering(电子信息与电气工程学院) Li Auto(力汽车) National University of Singapore(新加坡国立大学) Tsinghua University(清华大学) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生实验室) Ningbo Institute of Digital Twin(宁波数字孪生研究院)

AI总结 针对占据数据稀缺问题,构建最大语义占据数据集Nuplan-Occ,并提出统一框架联合生成高质量语义占据、多视角视频和LiDAR点云,采用时空解耦架构及高斯泼溅稀疏点图渲染和传感器感知嵌入策略,实现高保真生成。

Comments IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏