arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

The Hong Kong University of Science and Technology(香港科技大学)

共收录 2801
2507.16559 2026-02-03 cs.CV

Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge

内镜手术阶段识别、器械关键点估计和器械实例分割的比较验证:PhaKIR 2024挑战赛结果

Tobias Rueckert, David Rauber, Raphaela Maerkl, Leonard Klausmann, Suemeyye R. Yildiran, Max Gutbrod, Danilo Weber Nunes, Alvaro Fernandez Moreno, Imanol Luengo, Danail Stoyanov, Nicolas Toussaint, Enki Cho, Hyeon Bae Kim, Oh Sung Choo, Ka Young Kim, Seong Tae Kim, Gonçalo Arantes, Kehan Song, Jianjun Zhu, Junchen Xiong, Tingyi Lin, Shunsuke Kikuchi, Hiroki Matsuzaki, Atsushi Kouno, João Renato Ribeiro Manesco, João Paulo Papa, Tae-Min Choi, Tae Kyeong Jeong, Juyoun Park, Oluwatosin Alabi, Meng Wei, Tom Vercauteren, Runzhi Wu, Mengya Xu, An Wang, Long Bai, Hongliang Ren, Amine Yamlahi, Jakob Hennighausen, Lena Maier-Hein, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Shu Yang, Yihui Wang, Hao Chen, Santiago Rodríguez, Nicolás Aparicio, Leonardo Manrique, Juan Camilo Lyons, Olivia Hosie, Nicolás Ayobi, Pablo Arbeláez, Yiping Li, Yasmina Al Khalil, Sahar Nasirihaghighi, Stefanie Speidel, Daniel Rueckert, Hubertus Feussner, Dirk Wilhelm, Christoph Palm

机构 * Regensburg Medical Image Computing (ReMIC), OTH Regensburg(雷根萨大学医学影像计算中心) Research Group MITI, TUM University Hospital, School of Medicine and Health(技术大学慕尼黑大学医院MITI研究组) Regensburg Center of Biomedical Engineering (RCBE), OTH Regensburg(雷根萨生物医学工程研究中心) Regensburg University(雷根萨大学) Regensburg Center of Health Sciences and Technology (RCHST), OTH Regensburg(雷根萨健康科学与技术研究中心) AI Centre of Excellence, Medtronic Ltd.(医学影像人工智能卓越中心) Engineering Sciences, University College London(伦敦大学学院工程科学系) Augmented Intelligence Lab, Kyung Hee University(庆熙大学增强智能实验室) University of Minho, Braga(明霍大学) Jmees Inc.(Jmees公司) School of Sciences, São Paulo State University (UNESP), Bauru(圣保罗州立大学科学学院) KIST HARILAB, Center for Humanoid Research, Artificial Intelligence and Robot Institute, Korea Institute of Science and Technology (KIST)(韩国科学技术院HARILAB中心) King's College London(伦敦国王学院) The Chinese University of Hong Kong(香港中文大学) Division of Intelligent Medical Systems, German Cancer Research Center (DKFZ)(德国癌症研究中心智能医疗系统部门) Muroran Institute of Technology, Hokkaido(北海道Muroran技术学院) Niigata University of Health and Welfare(Niigata医疗福利大学) Konica Minolta, Inc.(东宝株式会社) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系) Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology(香港科技大学化学与生物工程系) HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, Shenzhen(香港科技大学深圳-香港协同创新研究院) Center for Research and Formation in Artificial Intelligence (CinfonIA), Los Andes University, Bogota(安第斯大学人工智能研究与培养中心) Department of Biomedical Engineering, Medical Image Analysis Group, Eindhoven University of Technology(埃因霍温理工大学生物医学工程系) Institute of Information Technology (ITEC), Klagenfurt University(克雷格夫大学信息技术研究所) Center for Tactile Internet with Human-in-the-loop (CeTI), TU Dresden(德累斯顿技术大学触觉互联网中心)

AI总结 PhaKIR 2024挑战赛通过多中心数据集验证内镜手术阶段识别、关键点估计和实例分割的性能,推动RAMIS领域的时间感知和情境驱动方法发展。

Comments A challenge report pre-print accepted by the journal Medical Image Analysis (MedIA), containing 37 pages, 15 figures, and 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24061 2026-02-03 cs.LG

Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learning

测量梯度,而非激活!增强深度强化学习中的神经元活动

Jiashun Liu, Zihao Wu, Johan Obando-Ceron, Pablo Samuel Castro, Aaron Courville, Ling Pan

机构 * Hong Kong University of Science and Technology(香港科技大学) Mila - Québec AI Institute(魁北克AI研究所) Université de Montréal(蒙特利尔大学)

AI总结 本文提出GraMa用于衡量深度强化学习中神经元的学习能力,通过梯度而非激活来增强神经元活动,提升学习性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16530 2026-02-03 cs.CR cs.AI cs.CL

DuFFin: A Dual-Level Fingerprinting Framework for LLMs IP Protection

DuFFin:一种用于LLM知识产权保护的双层指纹框架

Yuliang Yan, Haochun Tang, Shuo Yan, Enyan Dai

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Jilin University(吉林大学)

AI总结 DuFFin通过双层指纹框架在黑盒环境下实现LLM知识产权验证,准确识别模型来源并达到高验证精度。

Comments Accepted by EACL 2026, Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06914 2026-02-03 q-bio.QM cs.AI cs.LG

UniZyme: A Unified Protein Cleavage Site Predictor Enhanced with Enzyme Active-Site Knowledge

UniZyme:一种结合酶活性位点知识的统一蛋白质裂解位点预测器

Chenao Li, Shuo Yan, Enyan Dai

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 UniZyme通过结合活性位点知识的统一模型,实现了对多种酶裂解位点的高精度预测。

Comments 22 pages,9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13448 2026-02-03 cs.MA cs.AI cs.ET cs.LG

BMG-Q: Localized Bipartite Match Graph Attention Q-Learning for Ride-Pooling Order Dispatch

BMG-Q:局部双图匹配图注意力Q学习用于拼车订单调度

Yulong Hu, Siyuan Feng, Sen Li

机构 * Department of Civil and Environmental Engineering, The Hong Kong University of Science and Technology(香港科学与技术大学土木与环境工程系) Intelligent Transportation Thrust, Systems Hub, The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)智能交通 thrust,系统中心) Department of Aeronautical and Aviation Engineering, The Hong Kong Polytechnic University(香港理工大学航空与航空工程系)

AI总结 BMG-Q通过局部双图匹配图注意力Q学习提升拼车订单调度的决策效率与鲁棒性。

Journal ref IEEE Transactions on Intelligent Transportation Systems ( Volume: 26, Issue: 10, October 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19420 2026-02-03 cs.LG

UniGAP: A Universal and Adaptive Graph Upsampling Approach to Mitigate Over-Smoothing in Node Classification Tasks

UniGAP: 一种通用且自适应的图上采样方法以缓解节点分类任务中的过平滑问题

Xiaotang Wang, Yun Zhu, Haizhou Shi, Yongchao Liu, Yongqi Zhang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Rutgers University(罗格斯大学)

AI总结 UniGAP通过自适应图上采样缓解节点分类任务中的过平滑问题,提升模型性能并启发进一步研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04120 2026-02-03 cs.CL

RIDE: Difficulty Evolving Perturbation with Item Response Theory for Mathematical Reasoning

RIDE: 通过项目反应理论进行数学推理的难度演变扰动

Xinyuan Li, Murong Xu, Wenbiao Tao, Hanlun Zhu, Yike Zhao, Jipeng Zhang, Yunshi Lan

机构 * East China Normal University(东华大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

AI总结 RIDE通过项目反应理论生成更具挑战性的数学问题,评估大型语言模型的数学推理能力,实验显示其性能下降21.73%,验证了评估方法的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.23000 2026-02-02 cs.LG cs.AI

Mano: Restriking Manifold Optimization for LLM Training

Mano:重新审视面向大语言模型训练的流形优化

Yufei Gu, Zeke Xie

机构 * xLeaF Lab, The Hong Kong University of Science and Technology (Guangzhou)(xLeaF实验室,香港科学与技术大学(广州))

AI总结 Mano是一种新颖高效的流形优化器,通过创新投影和约束方法,在训练大语言模型时显著优于AdamW和Muon。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14133 2026-02-02 cs.RO cs.CV

TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers

TwinBrainVLA: 通过非对称Transformer混合释放通用视觉语言模型在具身任务中的潜力

Bin Yu, Shijie Lian, Xiaopeng Lin, Yuliang Wei, Zhaolong Shen, Changti Wu, Yuzhuo Miao, Xinming Wang, Bailing Wang, Cong Huang, Kai Chen

机构 * HIT(哈尔滨工业大学) ZGCA(中钢集团人工智能研究院) ZGCI(中钢集团智能计算研究院) HUST(华中科技大学) HKUST(GZ)(香港科技大学(广州)) BUAA(北京航空航天大学) ECNU(华东师范大学) CASIA(中国科学院自动化研究所) DeepCybo

AI总结 TwinBrainVLA通过非对称Transformer混合机制,利用冻结的通用视觉语言模型和可训练的专家路径,实现机器人任务中的具身智能提升。

Comments GitHub: https://github.com/ZGC-EmbodyAI/TwinBrainVLA

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10460 2026-02-02 cs.SE cs.AI

Understanding and Bridging the Planner-Coder Gap: A Systematic Study on the Robustness of Multi-Agent Systems for Code Generation

理解并弥合规划者-生成者之间的差距:对多智能体系统在代码生成中鲁棒性的系统研究

Zongyi Lyu, Songqiang Chen, Zhenlan Ji, Liwen Wang, Shuai Wang, Daoyuan Wu, Wenxuan Wang, Shing-Chi Cheung

机构 * The Hong Kong University of Science and Technology(香港科技大学) Lingnan University(岭南大学) Renmin University of China(中国人民大学)

AI总结 本文研究多智能体系统在代码生成中的鲁棒性问题,发现规划者与生成者之间的信息丢失是主要缺陷,并提出修复方法以提升系统鲁棒性。

Comments 18pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17621 2026-02-02 cs.LG

Navigate the Unknown: Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration

穿越未知:通过内在动机引导的探索增强大语言模型推理

Jingtong Gao, Ling Pan, Yejing Wang, Rui Zhong, Chi Lu, Maolin Wang, Qingpeng Cai, Peng Jiang, Xiangyu Zhao

机构 * City University of Hong Kong(香港城市大学) Hong Kong University of Science and Technology(香港科学与技术大学)

AI总结 IMAGINE通过内在动机引导的探索增强大语言模型推理能力,提升复杂任务性能22.23%

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22686 2026-02-02 cs.RO

FlyAware: Inertia-Aware Aerial Manipulation via Vision-Based Estimation and Post-Grasp Adaptation

FlyAware: 基于视觉估计与抓取后适应的惯性感知空中 manipulation

Biyu Ye, Na Fan, Zhengping Fan, Weiliang Deng, Hongming Chen, Qifeng Chen, Ximin Lyu

机构 * School of Intelligent Systems Engineering, Sun Yat-sen University(中山大学智能系统工程学院) Visual Intelligence Lab, Cheng Kar-Shun Robotics Institute, The Hong Kong University of Science and Technology(香港科技大学成卡顺机器人研究所视觉智能实验室) Differential Robotics Technology Company, Ltd.(差分机器人技术有限公司)

AI总结 FlyAware 通过基于视觉的预抓取惯性估计与抓取后适应机制,实现空中 manipulators 的稳健操控,验证了其在复杂惯性参数下的有效性。

Comments 8 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22517 2026-02-02 cs.RO

RoboStriker: Hierarchical Decision-Making for Autonomous Humanoid Boxing

RoboStriker: 为自主人形 boxing 的分层决策

Kangning Yin, Zhe Cao, Wentao Dong, Weishuai Zeng, Tianyi Zhang, Qiang Zhang, Jingbo Wang, Jiangmiao Pang, Ming Zhou, Weinan Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) Peking University(北京大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

AI总结 RoboStriker通过分层三阶段框架实现自主人形拳击,结合动作跟踪学习与潜在空间正则化,提升竞争性能和现实迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11391 2026-02-02 cs.LG

Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization

通过空域约束策略优化缓解安全对齐税

Yifan Niu, Han Xiao, Dongyi Liu, Nuo Chen, Jia Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 通过空域约束策略优化缓解安全对齐税,提出NSPO框架在保留LLM核心能力的同时提升安全性能。

Comments accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06013 2026-02-02 cs.CV cs.RO

VAT: Vision Action Transformer by Unlocking Full Representation of ViT

通过解锁ViT的完整表示构建视觉动作变换器

Wenhao Li, Chengwei Ma, Weixin Mao

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 VAT通过解锁ViT的完整表示,提出了一种新的视觉动作变换器架构,实现了在模拟操作任务中98.15%的成功率,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04769 2026-02-02 cs.CV

Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges

视觉-语言-动作(VLA)模型:概念、进展、应用与挑战

Ranjan Sapkota, Yang Cao, Konstantinos I. Roumeliotis, Manoj Karkee

机构 * Cornell University(康奈尔大学) The Hong Kong University of Science and Technology(香港科学与技术大学) University of the Peloponnese(希腊皮洛斯大学)

AI总结 本文综述了视觉-语言-动作(VLA)模型的概念、进展、应用与挑战,探讨了其在自动驾驶、医疗机器人等领域的应用及未来发展方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22094 2026-01-30 cs.CV

RefAny3D: 3D Asset-Referenced Diffusion Models for Image Generation

RefAny3D: 3D资产参考扩散模型用于图像生成

Hanzhuo Huang, Qingyang Bao, Zekai Gu, Zhongshuo Du, Cheng Lin, Yuan Liu, Sibei Yang

机构 * ShanghaiTech University(上海科技大学) Sun Yat-sen University(中山大学) University of Toronto(多伦多大学) The Hong Kong University of Science and Technology(香港科学与技术大学) SynWorld Macau University of Science and Technology(澳门科学理工学院)

AI总结 RefAny3D通过整合3D资产,提出一种双分支扩散模型,实现2D图像与3D资产的协同生成,提升图像生成的精确性和多样性。

Comments ICLR 2026. Project page: https://judgementh.github.io/RefAny3D Codes: https://github.com/JudgementH/RefAny3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21950 2026-01-30 cs.LG

Embracing Aleatoric Uncertainty in Medical Multimodal Learning with Missing Modalities

在医疗多模态学习中拥抱概率不确定性与缺失模态

Linxiao Gong, Yang Liu, Lianlong Sun, Yulai Bi, Jing Liu, Xiaoguang Zhu

机构 * HKUST (GZ)(香港科技大学) Tongji University(同济大学) University of Rochester(罗切斯特大学) Meta Fudan University(复旦大学) The University of British Columbia(不列颠哥伦比亚大学) University of California, Davis(加州大学戴维斯分校)

AI总结 本文提出AUM框架,通过建模单模态概率不确定性来应对医疗多模态学习中的缺失模态问题,在死亡预测任务中取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21804 2026-01-30 cs.CL

Distribution-Aware Reward Estimation for Test-Time Reinforcement Learning

面向分布的奖励估计用于测试时强化学习

Bodong Du, Xuanqi Huang, Xiaomeng Li

机构 * Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology, Hong Kong, China(电子与计算机工程系,香港科学与技术大学)

AI总结 DARE通过利用回放分布而非单一多数结果来改进测试时强化学习的奖励估计,提高了优化稳定性和最终性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21454 2026-01-30 cs.RO cs.CV

4D-CAAL: 4D Radar-Camera Calibration and Auto-Labeling for Autonomous Driving

4D-CAAL:面向自动驾驶的4D雷达-相机校准与自动标注

Shanliang Yao, Zhuoxiao Li, Runwei Guan, Kebin Cao, Meng Xia, Fuping Hu, Sen Xu, Yong Yue, Xiaohui Zhu, Weiping Ding, Ryan Wen Liu

机构 * School of Information Engineering, Yancheng Institute of Technology(信息工程学院,盐城职业技术学院) School of Navigation, Wuhan University of Technology(导航学院,武汉理工大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) School of Information Engineering, Yancheng Institute Technology(信息工程学院,盐城职业技术学院) School of Advanced Technology, Xi’an Jiaotong-Liverpool University(先进技术学院,西安交通大学利物浦大学) School of Information Science and Technology, Nantong University(信息科学与技术学院,南通大学) State Key Laboratory of Maritime Technology and Safety(船舶技术与安全国家重点实验室)

AI总结 4D-CAAL提出了一种统一框架,通过双用途校准目标和自动标注流程,实现4D雷达与相机的高精度校准,减少人工标注工作量,加速自动驾驶多模态感知系统开发。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21301 2026-01-30 cs.LG stat.ML

Achieving $\varepsilon^{-2}$ Dependence for Average-Reward Q-Learning with a New Contraction Principle

实现平均回报Q学习的ε⁻²依赖性:一种新的收缩原理

Zijun Chen, Zaiwei Chen, Nian Si, Shengbo Wang

机构 * Department of Computer Science and Engineering, HKUST(香港科技大学计算机科学与工程系) Edwardson School of Industrial Engineering, Purdue University(普渡大学工业工程学院) Department of Industrial Engineering and Decision Analytics, HKUST(香港科技大学工业工程与决策分析系) Daniel J. Epstein Department of Industrial and Systems Engineering, USC(美国南加州大学工业与系统工程系)

AI总结 本文提出了一种新的收缩原理,通过可达性假设实现了平均回报Q学习的最优ε⁻²样本复杂性保证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21192 2026-01-30 cs.AI cs.CL

Do Reasoning Models Enhance Embedding Models?

推理模型能增强嵌入模型吗?

Wun Yu Chan, Shaojin Chen, Huihao Jing, Kwun Hang Lau, Elton Chun-Chai Li, Zihao Wang, Haoran Li, Yangqiu Song

机构 * CSE, HKUST(香港科技大学计算机科学与工程系)

AI总结 本文研究了推理模型是否能提升嵌入模型的性能,发现其初始化对对比学习效果无显著影响,并提出HRSA框架揭示流形重新对齐现象。

Comments 10 main pages, 18 appendix pages, 13 figures, 11 tables, 4 prompts

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18874 2026-01-30 cs.HC cs.AI cs.CR cs.CY

When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs

当广告成为资料:利用LLMs揭示大规模网络广告中的隐形风险

Baiyu Chen, Benjamin Tag, Hao Xue, Daniel Angus, Flora Salim

机构 * The University of New South Wales(新南威尔士大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Queensland University of Technology(昆士兰理工大学)

AI总结 利用LLMs揭示广告流中的隐私信息泄露风险,展示其在大规模数据中的高精度推断能力。

Comments The ACM Web Conference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14250 2026-01-30 cs.LG

Towards Anomaly-Aware Pre-Training and Fine-Tuning for Graph Anomaly Detection

面向图异常检测的异常感知预训练与微调

Yunhui Liu, Jiashun Cheng, Yiqing Lin, Qizhuo Xie, Jia Li, Fugee Tsung, Hongzhi Yin, Tao Zheng, Jianhua Zhao, Tieke He

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) HKUST(GZ)(香港科技大学(广州)) Tsinghua University(清华大学) The University of Queensland(昆士兰大学)

AI总结 本文提出APF框架,通过预训练和微调提升图异常检测的异常感知能力,结合谱多项式滤波器和门控融合机制,实现更有效的异常识别。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02844 2026-01-30 cs.CL cs.IR cs.LG

GORAG: Graph-based Online Retrieval Augmented Generation for Dynamic Few-shot Social Media Text Classification

基于图的在线检索增强生成用于动态少样本社交媒体文本分类

Yubo Wang, Haoyang Li, Fei Teng, Lei Chen

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) The Hong Kong Polytechnic University(香港理工大学) Guangzhou HKUST Fok Ying Tung Research Institute(广州HKUST福ying顿研究 institute)

AI总结 GORAG提出一种基于图的在线检索增强生成框架,用于动态少样本社交媒体文本分类,通过构建关键词和标签的加权图并动态检索上下文以提升分类性能。

Comments Accepted by WWW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20575 2026-01-29 eess.IV cs.CV

SegRap2025: A Benchmark of Gross Tumor Volume and Lymph Node Clinical Target Volume Segmentation for Radiotherapy Planning of Nasopharyngeal Carcinoma

SegRap2025:鼻咽癌放疗计划中GTV和淋巴结临床靶区体积分割的基准测试

Jia Fu, Litingyu Wang, He Li, Zihao Luo, Huamin Wang, Chenyuan Bian, Zijun Gao, Chunbin Gu, Xin Weng, Jianghao Wu, Yicheng Wu, Jin Ye, Linhao Li, Yiwen Ye, Yong Xia, Elias Tappeiner, Fei He, Abdul qayyum, Moona Mazher, Steven A Niederer, Junqiang Chen, Chuanyi Huang, Lisheng Wang, Zhaohu Xing, Hongqiu Wang, Lei Zhu, Shichuan Zhang, Shaoting Zhang, Wenjun Liao, Guotai Wang

机构 * School of Mechanical and Electrical Engineering, University of Electronic Science and Technology of China(电子科技大学机械与电子工程学院) Department of Radiation Oncology, Sichuan Cancer Center, Radiation Oncology Key Laboratory of Sichuan Province, Sichuan Clinical Research Center for Cancer, Sichuan Cancer Hospital and Institute, University of Electronic Science and Technology of China(四川省肿瘤医院放射肿瘤科) Shanghai AI Lab, Shanghai, China(上海人工智能实验室) The Affiliated Hospital of Qingdao University, Qingdao, China(青岛大学附属医院) Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong, China(香港中文大学计算机科学与工程系) Faculty of Information Technology, Monash University, Melbourne, Australia(墨尔本大学信息科技学院) Department of Computing and Department of Brain Sciences, Imperial College London, United Kingdom(伦敦帝国理工学院计算机系和脑科学系) School of Computer Science and Engineering, Northwestern Polytechnical University, Xi'an, China(西北工业大学计算机科学与工程学院) UMIT Tirol - Private University for Health Sciences and Health Technology, Austria(蒂罗尔大学(私人健康科学与健康技术大学)) School of Information and Communication Engineering, University of Electronic Science and Technology of China(电子科技大学信息与通信工程学院) National Heart and Lung Institute, Imperial College London, London, United Kingdom(伦敦帝国理工学院国家心脏和肺研究所) Hawkes Institute, Department of Computer Science, University College London, London, United Kingdom(伦敦大学学院霍克斯研究所) Shanghai MediWorks Precision Instruments Co., Ltd., China(上海 MediWorks 精密仪器有限公司) School of Automation and Intelligent Sensing, Shanghai Jiao Tong University, Shanghai, China(上海交通大学自动化与智能感知学院) Department of Systems Hub, Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China(香港科技大学(广州)系统中心)

AI总结 SegRap2025旨在通过多中心、多模态的基准测试提升鼻咽癌放疗计划中GTV和LN CTV分割的通用性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20301 2026-01-29 cs.CV cs.AI

Towards Compact and Robust DNNs via Compression-aware Sharpness Minimization

通过压缩感知锐度最小化实现紧凑且鲁棒的DNN

Jialuo He, Huangxun Chen

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 本文提出C-SAM框架,通过压缩感知锐度最小化提升模型紧凑性与鲁棒性,实验显示鲁棒性提升达42%

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18714 2026-01-29 cs.CV

PLANA3R: Zero-shot Metric Planar 3D Reconstruction via Feed-Forward Planar Splatting

PLANA3R: 通过前馈平面撒点实现零样本度量平面三维重建

Changkun Liu, Bin Tan, Zeran Ke, Shangzhan Zhang, Jiachen Liu, Ming Qian, Nan Xue, Yujun Shen, Tristan Braud

机构 * The Hong Kong University of Science and Technology(香港科技大学) Ant Group(蚂蚁集团) Wuhan University(武汉大学) Zhejiang University(浙江大学) The Pennsylvania State University(宾夕法尼亚州立大学)

AI总结 PLANA3R通过前馈平面撒点实现零样本度量平面三维重建,无需显式平面监督,适用于大规模立体数据集。

Comments Camera-ready version of a paper in 39th Conference on Neural Information Processing Systems (NeurIPS 2025). The project page is available at: https://lck666666.github.io/plana3r

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16656 2026-01-29 cs.AI

NUMINA: A Natural Understanding Benchmark for Multi-dimensional Intelligence and Numerical Reasoning Abilities

NUMINA:多维智能与数值推理能力的自然理解基准

Changyu Zeng, Yifan Wang, Zimu Wang, Wei Wang, Zhengni Yang, Muyi Bao, Jiming Xiao, Anh Nguyen, Yutao Yue

机构 * School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学先进技术学院) Department of Computer Science, University of Liverpool(利物浦大学计算机科学系) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Institute of Deep Perception Technology, JITRI(深度感知技术研究所)

AI总结 NUMINA是一个用于多维智能和数值推理能力的自然理解基准,通过多尺度注释和自动化注释流程提升多模态室内感知理解,揭示当前LLM在三维空间计算中的不足。

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, pages 22575--22590

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11497 2026-01-29 cs.CV

QVGen: Pushing the Limit of Quantized Video Generative Models

QVGen:推动量化视频生成模型的极限

Yushi Huang, Ruihao Gong, Jing Liu, Yifu Ding, Chengtao Lv, Haotong Qin, Jun Zhang

机构 * Hong Kong University of Science and Technology(香港理工大学) Beihang University(北京航空航天大学) SenseTime Research(商汤科技研究院) Monash University(墨尔本大学) Nanyang Technological University(南洋理工大学) ETH Zürich(苏黎世联邦理工学院)

AI总结 QVGen通过量化感知训练框架在极低比特下实现高性能视频生成模型,首次达到全精度质量并优于现有方法。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏