arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Huazhong University of Science and Technology(华中科技大学)

共收录 702
2601.19284 2026-01-30 eess.SY cs.LG cs.SY math.OC

Model-Free Output Feedback Stabilization via Policy Gradient Methods

无模型输出反馈稳定化 via 策略梯度方法

Ankang Zhang, Ming Chi, Xiaoling Wang, Lintao Ye

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学) College of Automation, Nanjing University of Posts and Telecommunications(自动化学院,南京邮电大学)

AI总结 本文提出了一种无模型学习方法,用于部分可观测线性动态系统的输出反馈稳定化,通过策略梯度方法解决稳定问题并验证了算法的有效性。

Comments 31 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20325 2026-01-29 cs.CR cs.CV

UnlearnShield: Shielding Forgotten Privacy against Unlearning Inversion

UnlearnShield: 保护被遗忘隐私 against 反向学习

Lulu Xue, Shengshan Hu, Wei Lu, Ziqi Zhou, Yufei Song, Jianhong Cheng, Minghui Li, Yanjun Zhang, Leo Yu Zhang

机构 * Huazhong University of Science and Technology(华中科技大学) Institute of Guizhou Aerospace Measuring and Testing Technology(贵州航天测量测试技术研究院) University of Technology Sydney(悉尼科技大学) Griffith University(格里菲斯大学)

AI总结 UnlearnShield通过引入方向扰动和约束模块,有效降低反向学习反向的风险,同时保持模型准确性和实用性。

Comments This work has been accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20707 2026-01-29 cs.CV

Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models

融合重要性与多样性:KV缓存压缩的联合优化

Xuyang Liu, Xiyan Gui, Yuchao Zhang, Linfeng Zhang

机构 * EPIC Lab, Shanghai Jiao Tong University(上海交通大学) Sichuan University(四川大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 MixKV通过融合重要性与多样性,优化KV缓存压缩,提升多模态模型的存储效率与推理性能。

Comments Accepted by ICLR 2026. Our code is available at https://github.com/xuyang-liu16/MixKV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12605 2026-01-29 cs.CV

WaterFlow: Explicit Physics-Prior Rectified Flow for Underwater Saliency Mask Generation

WaterFlow: 基于显式物理先验的 rectified 流用于水下显著性掩码生成

Runting Li, Shijie Lian, Hua Li, Yutong Li, Wenhui Wu, Sam Kwong

机构 * Hainan University, China(海南大学) Huazhong University of Science and Technology, China(华中科技大学) Shenzhen University, China(深圳大学) Lingnan University, Hong Kong(岭南大学)

AI总结 WaterFlow通过引入显式物理先验和时间维度建模,提升水下显著性目标检测的性能,在USOD10K数据集上取得显著成效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19472 2026-01-28 cs.SD

Dual-Strategy-Enhanced ConBiMamba for Neural Speaker Diarization

双策略增强的ConBiMamba用于神经说话人识别

Zhen Liao, Gaole Dai, Mengqiao Chen, Wenqing Cheng, Wei Xu

机构 * School of Electronic Information(电子信息学院) Hubei Provincial Key Laboratory of Smart Internet Technology(智能互联网技术省重点实验室) Huazhong University of Science and Technology(华中科技大学)

AI总结 本文提出双策略增强的ConBiMamba系统,通过整合Conformer和Mamba的优势,提升说话人识别任务中局部细节和长跨度一致性建模的性能。

Comments Accepted at ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12121 2026-01-28 cs.LG cs.AI

Learning Dynamic Representations via An Optimally-Weighted Maximum Mean Discrepancy Optimization Framework for Continual Learning

通过最优加权最大均值差异优化框架进行动态表示学习以实现持续学习

KaiHui Huang, RunQing Wu, JinHui Sheng, HanYi Zhang, Ling Ge, JinYu Guo, Fei Ye

机构 * School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院) School of Mechanical Engineering, Huazhong University of Science and Technology(华中科技大学机械工程学院) Xihua University(西华大学) School of Computation, Information and Technology, Technische Universität München(慕尼黑工业大学计算、信息与技术学院) China Mobile Communications Group Chongqing Co., Ltd.(中国移动通信集团重庆有限公司)

AI总结 本文提出最优加权最大均值差异优化框架,通过多级特征匹配机制和自适应正则化优化策略,有效解决持续学习中的灾难性遗忘问题,并在实验中取得最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17844 2026-01-27 cs.HC cs.AI cs.LG

RAICL: Retrieval-Augmented In-Context Learning for Vision-Language-Model Based EEG Seizure Detection

RAICL:基于视觉-语言模型的EEG癫痫检测的检索增强上下文学习

Siyang Li, Zhuoya Wang, Xiyan Gui, Xiaoqing Chen, Ziwei Wang, Yaozhi Wen, Dongrui Wu

机构 * Ministry of Education Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(教育部图像处理与智能控制重点实验室,人工智能与自动化学院,华中科技大学) State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences(脑认知与脑启发智能技术国家重点实验室,自动化研究所,中国科学院)

AI总结 RAICL通过利用视觉-语言模型分析EEG波形图,实现了更高效的癫痫检测,无需重新训练,具有广泛临床应用前景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15892 2026-01-26 cs.CL

Stable-DiffCoder: Pushing the Frontier of Code Diffusion Large Language Model

Stable-DiffCoder:推动代码扩散大语言模型的前沿

Chenghao Fan, Wen Heng, Bo Li, Sichen Liu, Yuxuan Song, Jing Su, Xiaoye Qu, Kai Shen, Wei Wei

机构 * Huazhong University of Science and Technology(华中科技大学)

AI总结 Stable-DiffCoder通过基于扩散的训练方法,在代码生成任务中超越了传统自回归模型,提升了代码建模的质量和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11213 2026-01-23 cs.CV eess.IV

Simulating Dual-Pixel Images From Ray Tracing For Depth Estimation

从光线追踪模拟双像素图像用于深度估计

Fengchen He, Dayang Zhao, Hao Xu, Tingwei Quan, Shaoqun Zeng

机构 * Huazhong University of Science and Technology(华中科技大学)

AI总结 本文提出Sdirt方案,通过光线追踪生成逼真的双像素图像,用于提升深度估计模型对真实双像素数据的泛化能力。

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14950 2026-01-22 cs.CV

Erosion Attack for Adversarial Training to Enhance Semantic Segmentation Robustness

侵蚀攻击用于对抗训练以增强语义分割鲁棒性

Yufei Song, Ziqi Zhou, Menghao Deng, Yifan Hu, Shengshan Hu, Minghui Li, Leo Yu Zhang

机构 * Huazhong University of Science and Technology(华中科技大学) National University of Singapore(新加坡国立大学) Griffith University(格里菲斯大学)

AI总结 本文提出EroSeg-AT框架,通过生成具有上下文语义关系的对抗示例,提升语义分割模型在对抗攻击下的鲁棒性。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17160 2026-01-22 cs.CV

Can Synthetic Images Serve as Effective and Efficient Class Prototypes?

合成图像能否作为有效的高效类别原型?

Dianxing Shi, Dingjie Fu, Yuqiao Liu, Jun Wang

机构 * Beijing Research Institute of Uranium Geology(铀矿北京研究机构) Huazhong University of Science and Technology(华中科技大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Great Bay University(大湾大学)

AI总结 LGCLIP通过大语言模型生成类别特定提示,利用扩散模型合成参考图像,以实现高效的零样本分类。

Comments Accepted by IEEE ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13693 2026-01-21 q-bio.BM cs.AI

End-to-End Reverse Screening Identifies Protein Targets of Small Molecules Using HelixFold3

端到端反向筛选利用HelixFold3识别小分子的蛋白质靶标

Shengjie Xu, Xianbin Ye, Mengran Zhu, Xiaonan Zhang, Shanzhuo Zhang, Xiaomin Fang

机构 * PaddleHelix Team(PaddleHelix团队) Baidu Inc.(百度公司) School of Software Engineering(软件工程学院) Huazhong University of Science and Technology(华中科技大学) School of Pharmacy, Tongji Medical College(药学院)

AI总结 利用HelixFold3端到端反向筛选方法,提高小分子与蛋白质靶标识别的精度和结构保真度,支持药物发现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13352 2026-01-21 cs.CL cs.AI cs.MA

LLM-as-RNN: A Recurrent Language Model for Memory Updates and Sequence Prediction

LLM-as-RNN: 一种用于内存更新和序列预测的循环语言模型

Yuxing Lu, J. Ben Tamo, Weichen Zhao, Nan Sun, Yishan Zhong, Wenqi Shi, Jinzhuo Wang, May D. Wang

机构 * Georgia Institute of Technology(佐治亚理工学院) Peking University(北京大学) Shandong University(山东大学) Huazhong University of Science and Technology(华中科技大学) UT Southwestern Medical Center(西南医学中心)

AI总结 LLM-as-RNN通过将冻结的LLM转化为循环预测器,利用自然语言记忆实现在线学习,有效提升序列预测精度并生成可解释的学习轨迹。

Comments 17 pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12366 2026-01-21 cs.CV

DepthCropSeg++: Scaling a Crop Segmentation Foundation Model With Depth-Labeled Data

DepthCropSeg++: 通过深度标注数据扩展作物分割基础模型

Jiafei Zhang, Songliang Cao, Binghui Xu, Yanan Li, Weiwei Jia, Tingting Wu, Hao Lu, Weijuan Hu, Zhiguo Han

机构 * MetaPheno Laboratory(MetaPheno实验室) PhenoTrait Technology Co., Ltd.(PhenoTrait技术有限公司) Wuhan Digital Engineering Institute(武汉数字工程研究院) National Key Laboratory of Multispectral Information Intelligent Processing Technology(国家多光谱信息智能处理技术重点实验室) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) Hubei Key Laboratory of Intelligent Robot(湖北省智能机器人重点实验室) School of Computer Science and Engineering, School of Artificial Intelligence, Wuhan Institute of Technology(武汉理工大学计算机科学与工程学院、人工智能学院)

AI总结 DepthCropSeg++通过深度标注数据扩展作物分割模型,采用两阶段自训练流程提升性能,实现93.11%的mIoU,超越监督基线和SAM模型。

Comments 13 pages, 15 figures and 7 tables

Journal ref IEEE Journal of Selected Topics in Signal Processing, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20291 2026-01-21 cs.LG

Mixture-of-Experts with Gradient Conflict-Driven Subspace Topology Pruning for Emergent Modularity

专家混合模型与基于梯度冲突驱动的子空间拓扑修剪以实现涌现模ularity

Yuxing Gan, Ziyu Lei

机构 * Independent Researcher(独立研究者) Huazhong University of Science and Technology(华中科技大学)

AI总结 CDSP-MoE通过梯度冲突驱动子空间拓扑修剪,实现无指令场景下的稳健内容驱动路由和模块化结构演化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09527 2026-01-21 cs.CV

Generative Diffusion Contrastive Network for Multi-View Clustering

多视图聚类的生成扩散对比网络

Jian Zhu, Xin Zou, Xi Wang, Lei Liu, Chang Tang, Li-Rong Dai

机构 * Zhejiang Lab(浙江实验室) Hong Kong University of Science and Technology(香港科技大学) University of Science and Technology of China(中国科学技术大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 本文提出生成扩散对比网络GDCN,通过多重生成机制解决多视图聚类中的低质量数据问题,实现深度多视图聚类任务的最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11676 2026-01-21 cs.DC cs.AI cs.NI

HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network

HALO:语义感知的分布式LLM推理在失真边缘网络中

Peirong Zheng, Wenchao Xu, Haozhao Wang, Jinyu Chen, Xuemin Shen

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Division of Integrative Systems and Design, The Hong Kong University of Science and Technology(香港理工大学系统与设计学院) School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) Department of Electrical and Computer Engineering, University of Waterloo(滑铁卢大学电气与计算机工程系)

AI总结 HALO通过语义感知预测和负载平衡调度,提升失真边缘网络中LLM推理的效率与性能。

Comments Accepted by IEEE International Conference on Computer Communications (INFOCOM) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06474 2026-01-21 cs.CV cs.AI

SparseOccVLA: Bridging Occupancy and Vision-Language Models via Sparse Queries for Unified 4D Scene Understanding and Planning

SparseOccVLA: 通过稀疏查询弥合占用与视觉语言模型之间的鸿沟,实现统一的4D场景理解和规划

Chenxu Dang, Jie Wang, Guang Li, Zhiwen Hou, Zihan You, Hangjun Ye, Jie Ma, Long Chen, Yan Wang

机构 * Huazhong University of Science and Technology(华中科技大学) Xiaomi EV(小米电动车) Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)

AI总结 SparseOccVLA通过稀疏查询整合视觉语言模型与语义占用,实现统一的4D场景理解和规划,提升自动驾驶性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23045 2026-01-21 cs.AI

A Survey of AI Scientists

AI科学家的综述

Guiyao Tie, Pan Zhou, Lichao Sun

机构 * Huazhong University of Science and Technology(华中科技大学) Lehigh University(莱斯大学)

AI总结 本文综述了AI科学家的发展历程,提出六阶段方法论框架,分析了从基础模块到闭环系统再到可扩展性与人机协作的演进,为未来系统发展提供路线图。

Comments 28 pages, 9 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11522 2026-01-19 cs.CV

UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation

UniX: 统一自回归与扩散以实现胸部X光理解与生成

Ruiheng Zhang, Jingfeng Yao, Huangxuan Zhao, Hao Yan, Xiao He, Lei Chen, Zhou Wei, Yong Luo, Zengmao Wang, Lefei Zhang, Dacheng Tao, Bo Du

机构 * Wuhan University(武汉大学) Huazhong University of Science and Technology(华中科技大学) Nanyang Technological University(南洋理工大学)

AI总结 UniX通过统一自回归与扩散模型,实现了胸部X光的高效理解和生成,显著提升了性能并减少了参数使用。

Comments Codes and models are available at https://github.com/ZrH42/UniX

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11269 2026-01-19 cs.CV cs.AI

X-Distill: Cross-Architecture Vision Distillation for Visuomotor Learning

X-Distill:跨架构视觉蒸馏用于视觉-运动学习

Maanping Shao, Feihong Zhang, Gu Zhang, Baiye Cheng, Zhengrong Xue, Huazhe Xu

机构 * Tsinghua University(清华大学) Institute for Interdisciplinary Information Sciences(交叉信息研究院) Shanghai Qi Zhi Institute(上海启智研究院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Huazhong University of Science and Technology(华中科技大学)

AI总结 X-Distill通过跨架构知识蒸馏结合视觉转换器和紧凑型CNN,实现了在数据有限的机器人操作任务中优于其他方法的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11096 2026-01-19 cs.CV

CoDance: An Unbind-Rebind Paradigm for Robust Multi-Subject Animation

CoDance: 一种用于鲁棒多主体动画的解绑-重新绑定范式

Shuai Tan, Biao Gong, Ke Ma, Yutong Feng, Qiyuan Zhang, Yan Wang, Yujun Shen, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Ant Group(蚂蚁集团) Huazhong University of Science and Technology(华中科技大学) Tsinghua University(清华大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

AI总结 CoDance通过解绑-重新绑定框架实现鲁棒多主体动画,解决传统方法在处理任意主体数量、类型及空间错位时的局限性。

Comments https://lucaria-academy.github.io/CoDance/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10748 2026-01-19 eess.SP cs.AI cs.LG

AnyECG: Evolved ECG Foundation Model for Holistic Health Profiling

AnyECG:用于整体健康概况的进化ECG基础模型

Jun Li, Hongling Zhu, Yujie Xiao, Qinghao Zhao, Yalei Ke, Gongzheng Tang, Guangkun Nie, Deyun Zhang, Jin Li, Canqing Yu, Shenda Hong

机构 * National Institute of Health Data Science, Peking University(北京大学健康数据科学国家研究院) Institute of Medical Technology, Health Science Center of Peking University(北京大学医学技术研究所) Department of Internal Medicine, Tongji Hospital, Tongji Medical College, Huazhong University of Science and Technology(同济大学同济医学院内科部) Department of Cardiology, Peking University People’s Hospital(北京大学人民医院心内科) Department of Epidemiology and Biostatistics, School of Public Health, Peking University(北京大学公共卫生学院流行病与生物统计学系) Peking University Center for Public Health and Epidemic Preparedness & Response(北京大学公共卫生与突发卫生事件应对中心) Key Laboratory of Epidemiology of Major Diseases (Peking University), Ministry of Education(北京大学主要疾病流行病学重点实验室) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) HeartVoice Medical Technology(心声医疗技术)

AI总结 AnyECG通过迁移学习构建基础模型,实现对多种疾病的系统预测和长期风险评估,提升整体健康概况能力。

Comments in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06821 2026-01-15 cs.CL cs.AI cs.SE

Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems

LLMs能否生成可靠的测试用例生成器?对竞赛级编程问题的研究

Yuhan Cao, Zian Chen, Kun Quan, Ziliang Zhang, Yu Wang, Xiaoning Dong, Yeqi Feng, Guanzhong He, Jingcheng Huang, Jianhao Li, Yixuan Tan, Jiafu Tang, Yilin Tang, Junlei Wu, Qianyu Xiao, Can Zheng, Shouchen Zhou, Yuxiang Zhu, Yiming Huang, Tianxing He

机构 * Shanghai Qi Zhi Institute(上海启智研究院) ShanghaiTech University(上海科技大学) Wuhan University(武汉大学) Fuzhou University(福州大学) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学) Nanjing University(南京大学) Beijing University of Posts and Telecommunications(北京邮电大学) Peking University(北京大学)

AI总结 本文研究LLMs在生成竞赛级编程问题测试用例生成器方面的能力,提出TCGBench基准测试,并通过实验发现LLMs在生成针对性测试用例以暴露代码缺陷方面存在不足。

Comments 37 pages, 22 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03683 2026-01-14 cs.LG cs.NE

Rethinking Recurrent Neural Networks for Time Series Forecasting: A Reinforced Recurrent Encoder with Prediction-Oriented Proximal Policy Optimization

重新思考用于时间序列预测的循环神经网络:一种面向预测的强化编码器

Xin Lai, Shiming Deng, Lu Yu, Yumin Lai, Shenghao Qiao, Xinze Zhang

机构 * School of Management, Huazhong University of Science and Technology(华中科技大学管理学院) Center for Applied Mathematics, École des Ponts ParisTech(巴黎高等理工学院应用数学中心) School of Mechanical Engineering, Dalian Jiaotong University(大连交通大学机械工程学院) School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)

AI总结 本文提出RRE-PPO4Pred方法,通过强化编码器和改进的PPO算法提升时间序列预测的准确性和模型能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05096 2026-01-13 cs.CL

AdaSpec: Adaptive Speculative Decoding for Fast, SLO-Aware Large Language Model Serving

AdaSpec: 一种适应性推测解码用于快速、面向SLO的大型语言模型服务

Kaiyu Huang, Hao Wu, Zhubo Shi, Han Zou, Minchen Yu, Qingjiang Shi

机构 * Tongji University(同济大学) Shenzhen Research Institute of Big Data, The Chinese University of Hong Kong, Shenzhen(深圳大数据研究院,香港中文大学(深圳)) Huazhong University of Science and Technology(华中科技大学) School of Data Science, The Chinese University of Hong Kong, Shenzhen(数据科学学院,香港中文大学(深圳))

AI总结 AdaSpec通过动态调整推测策略,提高大型语言模型服务的响应速度和SLO满足度,实验显示性能提升达66%。

Comments This paper is accepted by ACM SoCC 2025

Journal ref In ACM Symposium on Cloud Computing (SoCC '25), November 19-21, 2025, Online, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05789 2026-01-12 cs.HC cs.AI cs.LG

SAFE: Secure and Accurate Federated Learning for Privacy-Preserving Brain-Computer Interfaces

SAFE: 用于隐私保护脑机接口的安全且准确的联邦学习

Tianwang Jia, Xiaoqing Chen, Dongrui Wu

机构 * Key Laboratory of the Ministry of Education for Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(教育部图像处理与智能控制重点实验室,人工智能与自动化学院,华中科技大学) Zhongguancun Academy(中关村学院)

AI总结 SAFE通过联邦学习在保护隐私的同时提升脑机接口的解码准确性和对抗鲁棒性,无需使用目标受试者的校准数据。

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21541 2026-01-12 cs.CV

Video Generation Models Are Good Latent Reward Models

视频生成模型是良好的潜在奖励模型

Xiaoyue Mi, Wenqing Yu, Jiesong Lian, Shibo Jie, Ruizhe Zhong, Zijun Liu, Guozhen Zhang, Zixiang Zhou, Zhiyong Xu, Yuan Zhou, Qinglin Lu, Fan Tang

机构 * University of Chinese Academy of Sciences(中国科学院大学) Tencent Hunyuan(腾讯文元) Huazhong University of Science and Technology(华中科技大学) Peking University(北京大学) Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) Nanjing University(南京大学)

AI总结 本文提出PRFL框架,利用预训练视频生成模型在噪声潜在空间中进行奖励建模,实现高效去噪和降低训练成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05246 2026-01-09 cs.CV

Pixel-Perfect Visual Geometry Estimation

像素级视觉几何估计

Gangwei Xu, Haotong Lin, Hongcheng Luo, Haiyang Sun, Bing Wang, Guang Chen, Sida Peng, Hangjun Ye, Xin Yang

机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Xiaomi EV(小米电动车)

AI总结 本文提出像素级视觉几何模型,通过生成建模技术实现高质量无飞像素点云,提升单目和视频深度估计的性能和准确性。

Comments Code: https://github.com/gangweix/pixel-perfect-depth

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03296 2026-01-09 cs.CL cs.LG

Towards Trustworthy Multimodal Moderation via Policy-Aligned Reasoning and Hierarchical Labeling

通过政策对齐推理和分层标注实现可信的多模态审核

Anqi Li, Wenwei Jin, Jintao Tong, Pengda Qin, Weijia Li, Guo Lu

机构 * Shanghai Jiao Tong University(上海交通大学) Xiaohongshu Inc.(小红书公司) Huazhong University of Science and Technology(华中科技大学)

AI总结 Hi-Guard通过政策对齐推理和分层标注提升多模态审核的准确性、泛化性和可解释性。

Comments Accepted by KDD 2026. Code is available at https://github.com/lianqi1008/Hi-Guard

详情

展开后加载摘要…

URL PDF HTML 收藏