arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高校专区

Huazhong University of Science and Technology(华中科技大学)

2026-03-24 至 2026-03-24 共收录 10
2603.22271 2026-03-24 cs.CV

DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolution

DUO-VSR:双流蒸馏用于一步视频超分辨率

Zhengyao Lv, Menghan Xia, Xintao Wang, Kwan-Yee K. Wong

机构 * The University of Hong Kong(香港大学) Huazhong University of Science and Technology(华中科技大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

AI总结 DUO-VSR提出双流蒸馏策略,通过轨迹保留蒸馏、双流优化和偏好引导细化,解决视频超分辨率中训练不稳定和监督不足的问题,提升视觉质量和效率。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21547 2026-03-24 cs.CV

PROBE: Diagnosing Residual Concept Capacity in Erased Text-to-Video Diffusion Models

PROBE: 诊断 erased 文本到视频扩散模型中的残余概念容量

Yiwei Xie, Zheng Zhang, Ping Liu

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) Department of Computer Science and Engineering, University of Nevada(内华达大学计算机科学与工程系)

AI总结 PROBE 通过多级评估框架揭示文本到视频扩散模型中被擦除概念的残余容量,发现所有测试方法留有可测量的残余能力,且其鲁棒性与干预深度相关,指出当前擦除方法仅实现输出级抑制而非表征移除。

Comments This preprint was posted after submission to IEEE Transactions

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21295 2026-03-24 cs.CV

Text-Image Conditioned 3D Generation

文本-图像条件的3D生成

Jiazhong Cen, Jiemin Fang, Sikuang Li, Guanjun Wu, Chen Yang, Taoran Yi, Zanwei Zhou, Zhikuan Bao, Lingxi Xie, Wei Shen, Qi Tian

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, School of Computer Science, Shanghai Jiao Tong University(人工智能大模型重点实验室、人工智能学院、计算机科学学院、上海交通大学) Huawei Inc.(华为公司) Huazhong University of Science and Technology(华中科技大学)

AI总结 本文提出结合文本和图像条件的3D生成方法,通过跨模态互补性提升生成质量,引入TIGON模型实现高效融合。

Comments CVPR 2026. Project page: https://jumpat.github.io/tigon-page Code: https://github.com/Jumpat/tigon

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21229 2026-03-24 cs.CV

Plant Taxonomy Meets Plant Counting: A Fine-Grained, Taxonomic Dataset for Counting Hundreds of Plant Species

植物分类与植物计数:一个细粒度、分类学数据集,用于计数数百种植物物种

Jinyu Xu, Tianqi Hu, Xiaonan Hu, Letian Zhou, Songliang Cao, Meng Zhang, Hao Lu

机构 * Huazhong University of Science and Technology(华中科技大学)

AI总结 本文提出TPC-268数据集,用于细粒度、分类学-aware的植物计数,包含10,000张图像和678,050个点注释,涵盖268种可计数的植物类别,推动了细粒度类无关计数的发展。

Comments Accepted by CVPR 2026. Project page: https://github.com/tiny-smart/TPC-268

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21134 2026-03-24 cs.RO cs.CV

Anatomical Prior-Driven Framework for Autonomous Robotic Cardiac Ultrasound Standard View Acquisition

基于解剖先验的自主机器人心脏超声标准视图采集框架

Zhiyan Cao, Zhengxi Wu, Yiwei Wang, Pei-Hsuan Lin, Li Zhang, Zhen Xie, Huan Zhao, Han Ding

机构 * State Key Laboratory of Intelligent Manufacturing Equipment and Technology, Huazhong University of Science and Technology(华中科技大学智能制造装备与技术国家重点实验室) School of Biomedical Engineering, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)生物医学工程学院) Information Intelligence Lab, Department of Electrical Engineering, National Chung Hsing University(中原大学电子工程系信息智能实验室) Institute of Medical Equipment Science and Engineering, Huazhong University of Science and Technology(华中科技大学医学装备科学与工程研究院) Department of Ultrasound Medicine, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology(华中科技大学同济医学院附属同济医院超声医学科) Institute of Systems Science (ISS), National University of Singapore (NUS)(新加坡国立大学系统科学研究所)

AI总结 本文提出结合心脏结构分割与自主探头调整的框架,通过解剖先验引导提升超声标准视图采集的自动化水平,实验表明其在分割精度和探头调整成功率上均有显著提升。

Comments Accepted for publication at the IEEE ICRA 2026. 8 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20708 2026-03-24 cs.CV

High-Quality and Efficient Turbulence Mitigation with Events

高质高效湍流抑制方法

Xiaoran Zhang, Jian Ding, Yuxing Duan, Haoyue Liu, Gang Chen, Yi Chang, Luxin Yan

机构 * State Key Laboratory of Multispectral Information Intelligent Processing Technology(多谱信息智能处理技术国家重点实验室) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)

AI总结 本文提出EHETM方法,利用事件相机的特性,通过事件极性变化和事件管约束实现高效湍流抑制,提升恢复质量并减少数据开销和系统延迟。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17655 2026-03-24 cs.CV cs.AI

Interpretable Cross-Domain Few-Shot Learning with Rectified Target-Domain Local Alignment

可解释的跨领域少样本学习与修正的目标域局部对齐

Yaze Zhao, Yixiong Zou, Yuhua Li, Ruixuan Li

机构 * School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)

AI总结 本文提出CC-CDFSL方法,通过循环一致性解决CLIP-based CDFSL中的局部对齐问题,提升局部视觉语言对齐和可解释性,实现SOTA性能。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23306 2026-03-24 cs.CV

ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding

ThinkOmni: 通过指导解码将文本推理提升到多模态场景

Yiran Guan, Sifan Tu, Dingkang Liang, Linghao Zhu, Jianzhong Ju, Zhenbo Luo, Jian Luan, Yuliang Liu, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) MiLM Plus, Xiaomi Inc.(小米公司)

AI总结 ThinkOmni提出一种无需训练和数据的框架,通过指导解码将文本推理扩展到多模态场景,实验显示在多个多模态推理基准上取得显著提升。

Comments Accept by ICLR 2026, Code: https://github.com/1ranGuan/thinkomni

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00596 2026-03-24 cs.IT cs.CR cs.DC cs.LG math.IT

Information-Theoretic Decentralized Secure Aggregation with Passive Collusion Resilience

信息论视角下的去中心化安全聚合与被动合谋鲁棒性

Xiang Zhang, Zhou Li, Shuangyang Li, Kai Wan, Derrick Wing Kwan Ng, Giuseppe Caire

机构 * Department of Electrical Engineering and Computer Science, Technical University of Berlin(技术大学柏林电气工程与计算机科学系) Guangxi Key Laboratory of Multimedia Communications and Network Technology, Guangxi University(广西多媒体通信与网络技术重点实验室,广西大学) School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院) School of Electrical Engineering and Telecommunications, University of New South Wales(新南威尔士大学电子工程与电信学院)

AI总结 研究去中心化安全聚合的理论极限,提出在合谋情况下保证输入总和安全计算的通信与密钥使用最优界限。

Comments Accepted by IEEE JSAC

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11289 2026-03-24 cs.CV cs.MM

UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer

UniAnimate-DiT:基于大规模视频扩散变换器的人像图像动画

Xiang Wang, Shiwei Zhang, Longxiang Tang, Yingya Zhang, Changxin Gao, Yuehuan Wang, Nong Sang

机构 * Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学图像处理与智能控制重点实验室,人工智能与自动化学院) Alibaba Group(阿里巴巴集团) Tsinghua University(清华大学)

AI总结 本文提出UniAnimate-DiT,利用Wan2.1模型实现一致的人像动画,通过LoRA技术优化参数,设计轻量姿态编码器并整合参考外观,实验表明其能生成高质量且时间一致的动画,支持从480p到720P的超分辨率。

Comments The training and inference code (based on Wan2.1) is available at https://github.com/ali-vilab/UniAnimate-DiT

详情

展开后加载摘要…

URL PDF HTML 收藏