arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

The Chinese University of Hong Kong(香港中文大学)

共收录 2407
2510.22213 2025-12-02 cs.CV

DynamicTree: Interactive Real Tree Animation via Sparse Voxel Spectrum

DynamicTree: 通过稀疏体素光谱实现交互式真实树动画

Yaokun Li, Lihe Ding, Xiao Chen, Guang Tan, Tianfan Xue

机构 * Sun Yat-sen University(中山大学) CUHK MMLab(香港中文大学多媒体实验室) CPII under InnoHK(创新科技署CPII)

AI总结 DynamicTree通过稀疏体素光谱实现交互式真实树动画,首次在3DGS重建中生成长期动态,提升视觉质量和效率。

Comments Project Page: https://dynamictree-dev.github.io/DynamicTree.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08594 2025-12-02 cs.CV cs.LG

PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction

PointNSP: 通过下一尺度细节预测实现自回归3D点云生成

Ziqiao Meng, Qichao Wang, Zhiyang Dou, Zixing Song, Zhipeng Zhou, Irwin King, Peilin Zhao

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) University of Hong Kong(香港大学) University of Cambridge(剑桥大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 PointNSP通过下一尺度细节预测实现自回归3D点云生成,首次在自回归范式中达到最先进的生成质量,并在参数、训练和推理效率上超越扩散基线。

Comments 24 pages; Previously this version appeared as arXiv:2510.05613 which was submitted as a new work by accident

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05808 2025-12-02 cs.CV cs.MM

SizeGS: Size-aware Compression of 3D Gaussian Splatting via Mixed Integer Programming

SizeGS: 通过混合整数规划实现3D高斯点云的尺寸感知压缩

Shuzhao Xie, Jiahang Liu, Weixiang Zhang, Shijia Ge, Sicheng Pan, Chen Tang, Yunpeng Bai, Cong Zhang, Xiaoyi Fan, Zhi Wang

机构 * SIGS, Tsinghua University(清华大学SIGS实验室) Harbin Institute of Technology(哈尔滨工业大学) MMLab, The Chinese University of Hong Kong(香港中文大学MMLab) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Jiangxing Intelligence Inc.(江行智能有限公司)

AI总结 SizeGS通过混合整数规划优化3DGS的超参数,实现高效尺寸感知压缩,提升压缩效率和视觉质量。

Comments Automatically compressing 3DGS into the desired file size while maximizing the visual quality

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00352 2025-12-02 cs.LG stat.ML

Sample-Efficient Tabular Self-Play for Offline Robust Reinforcement Learning

样本高效的目标导向自博弈用于离线鲁棒强化学习

Na Li, Zewu Zheng, Wei Ni, Hangguan Shan, Wenjie Zhang, Xinyu Li

机构 * Zhejiang University(浙江大学) The Chinese University of Hong Kong(香港中文大学) Edith Cowan University(埃迪斯·科温大学) University of New South Wales(新南威尔士大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 本文提出RTZ-VI-LCB算法,通过乐观鲁棒值迭代和数据驱动的伯恩斯坦惩罚项,实现离线鲁棒双人零和马尔可夫游戏的最优样本复杂性保证。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00308 2025-12-02 cs.CV

Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation

基于最优传输的分布几何对齐优化用于生成数据集蒸馏

Xiao Cui, Yulei Qin, Wengang Zhou, Hongsheng Li, Houqiang Li

机构 * University of Science and Technology of China(中国科学技术大学) CUHK MMLab(香港中文大学多媒体实验室) Independent Researcher(独立研究者)

AI总结 本文提出基于最优传输的分布几何对齐方法,通过优化传输距离提升数据集蒸馏效果,实验表明在ImageNet-1K上准确率提升至少4%。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16177 2025-12-02 cs.GR cs.CV

OccluGaussian: Occlusion-Aware Gaussian Splatting for Large Scene Reconstruction and Rendering

OccluGaussian:面向大场景重建与渲染的遮挡感知高斯点云技术

Shiyong Liu, Xiao Tang, Zhihao Li, Yingfan He, Chongjie Ye, Jianzhuang Liu, Binxiao Huang, Shunbo Zhou, Xiaofei Wu

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室) The Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳)) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) The University of Hong Kong(香港大学) Huawei Embodied Intelligence Lab(华为具身智能实验室)

AI总结 OccluGaussian通过遮挡感知的场景划分和区域渲染技术,提升大规模场景重建与渲染的质量和效率。

Comments Accepted to ICCV 2025. Project website: https://occlugaussian.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10125 2025-12-02 cs.CV cs.MM

Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation

代理调优:为以主题驱动的图像生成定制多模态自回归模型

Yi Wu, Shengju Qian, Lingting Zhu, Lei Liu, Wandi Qiao, Ziqiang Li, Lequan Yu, Bin Li

机构 * University of Science and Technology of China(中国科学技术大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Nanjing University of Information Science and Technology(南京信息工程大学)

AI总结 本文提出代理调优方法,通过扩散模型增强AR模型在主题驱动图像生成中的能力,揭示了弱到强泛化现象,提升了多主题组合和上下文理解的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23476 2025-12-01 cs.AI

Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction

通过多轮交互构建高效的LLM世界模型推理:WMAct

Bao Shu, Yan Cai, Jianjian Sun, Chunrui Han, En Yu, Liang Zhao, Jingcheng Hu, Yinmin Zhang, Haoran Lv, Yuang Peng, Zheng Ge, Xiangyu Zhang, Daxin Jiang, Xiangyu Yue

机构 * CUHK MMLab(香港中文大学MMLab) Peking University(北京大学) StepFun Tsinghua University(清华大学)

AI总结 WMAct通过高效交互和主动推理机制,使LLM在复杂环境中实现高效世界模型推理,提升任务解决能力和迁移性能。

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23282 2025-12-01 cs.LG cs.DC cs.IT math.IT

Closing the Generalization Gap in Parameter-efficient Federated Edge Learning

缩小参数高效联邦边缘学习中的泛化差距

Xinnong Du, Zhonghao Lyu, Xiaowen Cao, Chunyang Wen, Shuguang Cui, Jie Xu

机构 * School of Science and Engineering (SSE)(科学与工程学院) Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen)(深圳未来网络智能研究所) Guangdong Provincial Key Laboratory of Future Networks of Intelligence(广东省未来网络智能重点实验室) The Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳)) School of Electrical Engineering and Computer Science(电气工程与计算机科学学院) KTH Royal Institute of Technology(皇家理工学院) College of Electronic and Information Engineering(电子与信息工程学院) University of Science and Technology of China (USTC)(中国科学技术大学)

AI总结 本文提出了一种参数高效的联邦边缘学习框架,通过联合模型剪枝和客户端选择,解决本地数据异质性和资源受限问题,提升模型泛化能力和学习性能。

Comments 13 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15200 2025-12-01 cs.RO

VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation

VIRAL:大规模视觉仿真到现实用于双足机器人类地交互

Tairan He, Zi Wang, Haoru Xue, Qingwei Ben, Zhengyi Luo, Wenli Xiao, Ye Yuan, Xingye Da, Fernando Castañeda, Shankar Sastry, Changliu Liu, Guanya Shi, Linxi Fan, Yuke Zhu

机构 * NVIDIA CMU(卡内基梅隆大学) UC Berkeley(加州大学伯克利分校) CUHK(香港中文大学)

AI总结 VIRAL通过大规模仿真训练并零样本部署,实现双足机器人在现实中的自主移动- manipulation能力,无需现实微调。

Comments Project website: https://viral-humanoid.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24693 2025-12-01 cs.SD cs.CL eess.AS

STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence

STAR-Bench:探测深度时空推理作为音频4D智能

Zihan Liu, Zhikang Niu, Qiuyang Xiao, Zhisheng Zheng, Ruoqi Yuan, Yuhang Zang, Yuhang Cao, Xiaoyi Dong, Jianze Liang, Xie Chen, Leilei Sun, Dahua Lin, Jiaqi Wang

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Shanghai Innovation Institute(上海创新研究院) Beihang University(北京航空航天大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 STAR-Bench通过测试音频4D智能,揭示了模型在细粒度感知和推理上的不足,为未来模型发展提供方向。

Comments Homepage: https://internlm.github.io/StarBench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22677 2025-12-01 cs.CV

Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield

解耦DMD:CFG增强作为矛,分布匹配作为盾

Dongyang Liu, Peng Gao, David Liu, Ruoyi Du, Zhen Li, Qilong Wu, Xin Jin, Sihan Cao, Shifeng Zhang, Hongsheng Li, Steven Hoi

机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团) The Chinese University of Hong Kong(香港中文大学)

AI总结 本文提出解耦DMD方法,发现CFG增强是蒸馏核心驱动因素,分布匹配作为正则化器,通过解耦噪声调度提升生成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22134 2025-12-01 cs.CV cs.RO

DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action

DualVLA: 通过推理与行动部分解耦构建通用具身代理

Zhen Fang, Zhuoyang Liu, Jiaming Liu, Hao Chen, Yu Zeng, Shiting Huang, Zehui Chen, Lin Chen, Shanghang Zhang, Feng Zhao

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知国家重点实验室,中国科学技术大学) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,北京大学计算机学院) CUHK(香港大学)

AI总结 DualVLA通过推理与行动部分解耦,提升通用具身代理的行动与多模态理解平衡能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22103 2025-12-01 cs.CV

MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding

MoE3D:专家混合方法与多模态3D理解

Yu Li, Yuenan Hou, Yingmei Wei, Xinge Zhu, Yuexin Ma, Wenqi Shao, Yanming Guo

机构 * National University of Defense Technology(国防科技大学) Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) ShanghaiTech University(上海科技大学)

AI总结 MoE3D通过整合专家混合方法,提升多模态3D理解的性能,尤其在Multi3DRefer任务中表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22043 2025-12-01 cs.RO cs.SY eess.SY

SwordRiding: A Unified Navigation Framework for Quadrotors in Unknown Complex Environments via Online Guiding Vector Fields

SwordRiding:一种基于在线引导矢量场的四旋翼无人机在未知复杂环境中的统一导航框架

Xuchen Liu, Ruocheng Li, Bin Xin, Weijia Yao, Qigeng Duan, Jinqiang Cui, Ben M. Chen, Jie Chen

机构 * Pengcheng Laboratory, Shenzhen, Guangdong, China(鹏城实验室,深圳,广东,中国) School of Automation, Beijing Institute of Technology, Beijing, China(自动化学院,北京理工大学,北京,中国) School of Artificial Intelligence and Robotics, Hunan University(人工智能与机器人学院,湖南大学) Department of Control Science and Engineering, Harbin Institute of Technology(控制科学与工程学院,哈尔滨工业大学) National Key Laboratory of Autonomous Intelligent Unmanned Systems(自主智能无人系统国家重点实验室) Department of Mechanical and Automation Engineering, the Chinese University of Hong Kong, Hong Kong, China(机械与自动化工程系,香港中文大学,香港,中国)

AI总结 SwordRiding提出了一种基于在线引导矢量场的四旋翼无人机统一导航框架,通过闭环导航提升在未知复杂环境中的实时适应性和鲁棒性。

Comments For an experimental demo, see https://www.youtube.com/watch?v=tKYCg266c4o. For the lemma proof, see https://github.com/SmartGroupSystems/GVF_close_loop_planning/blob/main/proofs.md

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21688 2025-12-01 cs.CV cs.AI cs.CL

G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning

G$^2$VLM:基于几何的视觉语言模型,实现统一的3D重建与空间推理

Wenbo Hu, Jingli Lin, Yilin Long, Yunlong Ran, Lihan Jiang, Yifan Wang, Chenming Zhu, Runsen Xu, Tai Wang, Jiangmiao Pang

机构 * Shanghai AI Lab(上海人工智能实验室) UCLA(加州大学洛杉矶分校) SJTU(上海交通大学) FDU(福建师范大学) ZJU(浙江大学) USTC(中国科学技术大学) HKU(香港大学) CUHK(香港中文大学)

AI总结 G$^2$VLM通过统一3D重建与空间推理,提升视觉语言模型在空间智能任务中的性能。

Comments code are released at https://github.com/InternRobotics/G2VLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07769 2025-12-01 cs.CV

Dream4D: Lifting Camera-Controlled I2V towards Spatiotemporally Consistent 4D Generation

Dream4D: 通过可控视频生成与神经4D重建提升I2V向时空一致4D生成

Xiaoyan Liu, Kangrui Li, Yuehao Song, Jiaxin Liu

机构 * The Chinese University of Hong Kong(香港中文大学) The Hong Kong Polytechnic University(香港理工大学) Huazhong University of Science and Technology(华中科技大学) The University of New South Wales(新南威尔士大学)

AI总结 Dream4D通过可控视频生成与神经4D重建的协同,实现高质量的时空一致4D内容生成。

Comments Project Page: https://wanderer7-sk.github.io/Dream4D.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17511 2025-12-01 cs.CV

Accelerating Parallel Diffusion Model Serving with Residual Compression

通过残差压缩加速并行扩散模型服务

Jiajun Luo, Yicheng Xiao, Jianru Xu, Yangxiu You, Rongwei Lu, Chen Tang, Jingyan Jiang, Zhi Wang

机构 * Southern University of Science and Technology(南方科技大学) Jiangnan University(江南大学) Shenzhen Technology University(深圳技术大学) The Chinese University of Hong Kong(香港中文大学)

AI总结 CompactFusion通过残差压缩技术提升并行扩散模型服务效率,实现3倍加速并显著提高生成质量。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15327 2025-12-01 cs.RO cs.LG

Advancing Embodied Intelligence in Robotic-Assisted Endovascular Procedures: A Systematic Review of AI Solutions

推进机器人辅助血管内手术中的具身智能:人工智能解决方案的系统综述

Tianliang Yao, Bo Lu, Markus Kowarschik, Yixuan Yuan, Hubin Zhao, Sebastien Ourselin, Kaspar Althoefer, Junbo Ge, Peng Qi

机构 * Department of Control Science and Engineering, College of Electronics and Information Engineering, and Shanghai Institute of Intelligent Science and Technology, Tongji University(控制科学与工程系,电子信息工程学院,上海智能科学技术研究院,同济大学) Department of Electronic Engineering, Faculty of Engineering, The Chinese University of Hong Kong(电子工程系,工程学院,香港中文大学) Robotics and Microsystems Center, School of Mechanical and Electrical Engineering, Soochow University(机器人与微系统中心,机械与电气工程学院,苏州大学) Siemens Healthineers Advanced Therapies (AT), Forchheim, Bavaria(西门子医疗先进治疗(AT), Forchheim, 巴伐利亚) HUB of Intelligent Neuro-Engineering (HUBIN), CREATe, Division of Surgery & Interventional Science, University College London(智能神经工程中心(HUBIN),CREATE,外科与介入科学系,伦敦大学学院) School of Biomedical Engineering & Imaging Sciences, King’s College London(生物医学工程与成像科学学院,伦敦国王学院)

AI总结 本文系统综述了人工智能在机器人辅助血管内手术中具身智能的应用,探讨了其挑战与未来发展方向。

Comments 20 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11561 2025-12-01 cs.CV

Teaching Large Language Models to Regress Accurate Image Quality Scores using Score Distribution

利用大规模语言模型回归准确的图像质量评分使用分数分布

Zhiyuan You, Xin Cai, Jinjin Gu, Tianfan Xue, Chao Dong

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Multimedia Laboratory, The Chinese University of Hong Kong(香港中文大学多媒体实验室) Shanghai AI Laboratory(上海人工智能实验室) Shenzhen University of Advanced Technology(深圳先进技术大学) CPII under InnoHK(创新香港下的CPII)

AI总结 本研究提出基于分布的DeQA-Score模型,通过离散化评分分布为软标签,提升图像质量评分的准确性和一致性。

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18842 2025-12-01 cs.CV

Enhancing Descriptive Image Quality Assessment with A Large-scale Multi-modal Dataset

通过大规模多模态数据集增强描述性图像质量评估

Zhiyuan You, Jinjin Gu, Xin Cai, Zheyuan Li, Kaiwen Zhu, Chao Dong, Tianfan Xue

机构 * The Chinese University of Hong Kong(香港中文大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) Sofia University(索菲亚大学) University of Macau(澳门大学) Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shenzhen University of Advanced Technology(深圳先进技术大学)

AI总结 本研究提出DepictQA-Wild模型,通过构建大规模多模态数据集提升图像质量评估的准确性和实用性。

Comments Accepted by TIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21557 2025-11-27 cs.RO cs.AI

VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation

VacuumVLA: 通过统一的吸力和抓取工具提升VLA能力以实现复杂机器人操作

Hui Zhou, Siyuan Huang, Minxing Li, Hao Zhang, Lue Fan, Shaoshuai Shi

机构 * The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) DiDi Global(滴滴出行)

AI总结 VacuumVLA通过整合吸力与抓取功能的末端执行器,提升VLA在复杂机器人操作中的能力,实现更多任务的可行性。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21139 2025-11-27 cs.CV

Referring Video Object Segmentation with Cross-Modality Proxy Queries

基于跨模态代理查询的指引用视频目标分割

Baoli Sun, Xinzhu Ma, Ning Wang, Zhihui Wang, Zhiyong Wang

机构 * DUT-RU International School of Information Science & Engineering, Dalian University of Technology, China(大连理工大学国际信息科学与工程学院,中国) Chinese University of Hong Kong, China(香港中文大学,中国) University of Sydney, Australia(悉尼大学,澳大利亚)

AI总结 ProxyFormer通过引入代理查询和联合语义一致性策略,提升指引用视频目标分割的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18780 2025-11-27 cs.CV cs.AI

ConceptGuard: Proactive Safety in Text-and-Image-to-Video Generation through Multimodal Risk Detection

ConceptGuard:通过多模态风险检测实现文本-图像到视频生成的主动安全

Ruize Ma, Minghong Cai, Yilei Jiang, Jiaming Han, Yi Feng, Yingshui Tan, Xiaoyong Zhu, Bo Zhang, Bo Zheng, Xiangyu Yue

机构 * CUHK MMLab(香港中文大学多模态实验室) Future Lab, Alibaba Group(阿里巴巴集团未来实验室) Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室)

AI总结 ConceptGuard通过多模态风险检测实现文本-图像到视频生成的主动安全,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15703 2025-11-27 cs.CV cs.AI cs.CL

Think Visually, Reason Textually: Vision-Language Synergy in ARC

视觉优先,文本推理:ARC中的视觉-语言协同

Beichen Zhang, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, Jiaqi Wang

机构 * The Chinese University of Hong Kong(香港中文大学) Shanghai AI Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院)

AI总结 本文提出视觉-语言协同推理和模态切换自我纠正策略,通过结合视觉抽象与语言推理提升ARC-AGI任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17729 2025-11-27 cs.LG math.ST stat.ME stat.TH

A Conditional Distribution Equality Testing Framework using Deep Generative Learning

基于深度生成学习的条件分布相等性检验框架

Siming Zheng, Tong Wang, Meifang Lan, Yuanyuan Lin

机构 * Department of Industrial Systems Engineering and Management, Nationalal University of Singapore(工业系统工程与管理系,新加坡国立大学) Department of Statistics and Data Science, The Chinese University of Hong Kong(统计与数据科学系,香港中文大学)

AI总结 本文提出基于深度生成学习的条件分布相等性检验框架,通过转换条件检验问题为无条件问题,结合生成分类准确度方法,验证了在协变量偏移和因果发现中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01759 2025-11-27 cs.LG cs.AR

TinyFormer: Efficient Transformer Design and Deployment on Tiny Devices

TinyFormer: 在微型设备上高效设计和部署变换器

Jianlei Yang, Jiacheng Liao, Fanding Lei, Meichen Liu, Lingkun Long, Junyi Chen, Han Wan, Bei Yu, Weisheng Zhao

机构 * School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) Qingdao Research Institute, Beihang University(北京航空航天大学青岛研究所) Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学(深圳)计算机科学与工程系) Fert Beijing Research Institute, School of Integrated Circuit Science and Engineering, Beihang University(北京航空航天大学集成电路科学与工程学院)

AI总结 TinyFormer是一种专为微型设备设计的高效变换器框架,通过SuperNAS、SparseNAS和SparseEngine实现资源高效的模型设计与部署,显著提升推理速度。

Comments This paper is accepted by IEEE Transactions on Circuits and Systems I: Regular Papers

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12029 2025-11-27 cs.LG cs.AI

Dual-Balancing for Multi-Task Learning

多任务学习中的双平衡

Baijiong Lin, Weisen Jiang, Feiyang Ye, Yu Zhang, Pengguang Chen, Ying-Cong Chen, Shu Liu, Ivor W. Tsang, James T. Kwok

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) HKUST(GZ) - SmartMore Joint Lab(HKUST(广州)- SmartMore联合实验室) The Chinese University of Hong Kong(香港中文大学) Southern University of Science and Technology(南方科技大学) Centre for Frontier AI Research, A$^*$STAR(前沿人工智能研究中心,A$^*$STAR)

AI总结 本文提出DB-MTL方法,通过损失和梯度双角度平衡提升多任务学习性能。

Comments Accepted by Neural Networks

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20646 2025-11-26 cs.CV

3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding

具有跨视角相关性的3D感知多任务学习

Xiaoye Wang, Chen Tang, Xiangyu Yue, Wei-Hong Li

机构 * University of Cambridge(剑桥大学) The Chinese University of Hong Kong(香港中文大学) University of Bristol(布里斯托大学)

AI总结 本文提出一种具有跨视角相关性的3D感知多任务学习方法,通过整合成本体积作为几何一致性来提升密集场景理解的性能。

Comments 3D-aware Multi-task Learning, Cross-view Correlations, Code will be available at https://github.com/WeiHongLee/CrossView3DMTL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20624 2025-11-26 cs.CV

ShapeGen: Towards High-Quality 3D Shape Synthesis

ShapeGen:迈向高质量3D形状合成

Yangguang Li, Xianglong He, Zi-Xin Zou, Zexiang Liu, Wanli Ouyang, Ding Liang, Yan-Pei Cao

机构 * The Chinese University of Hong Kong, VAST(香港中文大学,VAST) Tsinghua University, VAST(清华大学,VAST) The Chinese University of Hong Kong, Shanghai Artificial Intelligence Laboratory(香港中文大学上海人工智能实验室)

AI总结 ShapeGen通过改进3D表示、分辨率扩展和线性变换器,实现了高质量的图像到3D形状生成,提升了生成资产的细节和结构完整性。

Comments Accepted to SIGGRAPH Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏