arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高校专区

The Hong Kong University of Science and Technology(香港科技大学)

2026-03-24 至 2026-03-24 共收录 20
2603.22280 2026-03-24 cs.CV cs.RO

DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models

DualCoT-VLA: 通过并行推理实现视觉-语言链式思维的视觉-语言-动作模型

Zhide Zhong, Junfeng Li, Junjie He, Haodong Yan, Xin Gong, Guanyi Zhao, Yingjie Cai, Jiantao Gao, Xu Yan, Bingbing Liu, Yingcong Chen, Liuqing Yang, Haoang Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Huawei Foundation Model Department(华为基础模型部门)

AI总结 DualCoT-VLA通过并行推理机制,结合视觉和语言链式思维,解决传统VLA模型在复杂多步骤任务和精细空间感知中的不足,实现更高效的视觉-语言-动作处理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21928 2026-03-24 cs.CV cs.LG

The Golden Subspace: Where Efficiency Meets Generalization in Continual Test-Time Adaptation

黄金子空间:在持续测试时适应中效率与泛化的平衡

Guannan Lai, Da-Wei Zhou, Zhenguo Li, Han-Jia Ye

机构 * School of Artificial Intelligence, Nanjing University(南京大学人工智能学院) National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Hong Kong University of Science and Technology(香港科技大学) Frontier Robotics(前沿机器人)

AI总结 本文提出GOLD方法,通过轻量级适配器将特征投影到黄金子空间,并动态更新子空间以提升效率和稳定性,实验证明其在分类和分割任务中表现优异。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21492 2026-03-24 cs.LG math.OC

Multinoulli Extension: A Lossless Continuous Relaxation for Partition-Constrained Subset Selection

多努利扩展:一种无损的连续松弛方法用于分区约束子集选择

Qixin Zhang, Wei Huang, Yan Sun, Yao Shu, Yi Yu, Dacheng Tao

机构 * college of computing and data science, Nanyang Technological University, Singapore(南洋理工大学计算与数据科学学院,新加坡) Hong Kong University of Science and Technology(香港科学大学)

AI总结 本文提出Multinoulli-SCG算法,通过多努利扩展框架在满足分区约束下实现子集选择的无损连续松弛,提升效率并保证近似保证。

Comments 45 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21475 2026-03-24 cs.AI

Unified-MAS: Universally Generating Domain-Specific Nodes for Empowering Automatic Multi-Agent Systems

统一-MAS:通用领域节点生成以增强自动多智能体系统

Hehai Lin, Yu Yan, Zixuan Wang, Bo Xu, Sudong Wang, Weiquan Huang, Ruochen Zhao, Minzhi Li, Chengwei Qin

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学) Institute for Infocomm Research (I 2 R), A*STAR(信息通信研究院(I2R),A*STAR)

AI总结 本文提出统一-MAS,通过离线节点合成解耦细粒度节点实现与拓扑编排,提升多智能体系统在知识密集型领域的性能与成本效益。

Comments Code is available at https://github.com/linhh29/Unified-MAS

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19961 2026-03-24 cs.CL cs.IR

Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval

解锁多模态文档智能:从当前成就到视觉文档检索的未来前沿

Yibo Yan, Jiahao Huo, Guanbo Feng, Mingdong Ou, Yi Cao, Xin Zou, Shuliang Liu, Yuanhuiyi Lyu, Yu Huang, Jungang Li, Kening Zheng, Xu Zheng, Philip S. Yu, James Kwok, Xuming Hu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Alibaba Cloud Computing(阿里云计算) Hong Kong University of Science and Technology(香港科技大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

AI总结 本文综述了视觉文档检索领域,探讨了多模态大语言模型时代下的方法演进与挑战,提出未来发展方向。

Comments Under review. This version updates the relevant works released before 15 March, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21565 2026-03-24 cs.CV

UAVLight: A Benchmark for Illumination-Robust 3D Reconstruction in Unmanned Aerial Vehicle (UAV) Scenes

UAVLight:用于无人机(UAV)场景中抗光照干扰的3D重建基准

Kang Du, Xue Liao, Junpeng Xia, Chaozheng Guo, Yi Gu, Yirui Guan, Duotun Wang, Sheng Huang, Zeyu Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Meituan UAV(美团无人机) Beijing University of Chemical Technology(北京化工大学)

AI总结 UAVLight基准通过可控且真实的飞行路径捕获多时段多光照条件下的场景,提供一致几何和校准下的自然光照变化,用于评估抗光照干扰的3D重建方法。

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04559 2026-03-24 cs.CV

Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning

面向可扩展多模态推理的推理对齐感知解耦

Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Xin Jin, Zhenguo Li, James T. Kwok, Yu Zhang

机构 * Southern University of Science and Technology(南方科技大学) The Hong Kong University of Science and Technology(香港科技大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Huawei Cloud Project(华为云项目)

AI总结 本文提出RAPID方法,通过解耦多模态模型的感知与推理模块,利用强化学习优化感知输出,使模型能与任意强大文本推理模型配合,实现无需重新训练的性能提升。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.04317 2026-03-24 cs.CV cs.CL

WiFi-GEN: High-Resolution Indoor Imaging from WiFi Signals Using Generative AI

WiFi-GEN:利用生成式人工智能进行高分辨率室内成像

Jianyang Shi, Bowen Zhang, Amartansh Dubey, Ross Murch, Liwen Jing

机构 * Pengcheng Laboratory(鹏城实验室) College of Big Data and Internet(大数据与互联网学院) Shenzhen Technology University(深圳技术大学) Department of Electrical Engineering(电子工程系) Indian Institute of Technology Delhi(印度理工学院德里分校) Department of Electronic and Computer Engineering(电子与计算机工程系) HKUST(香港科技大学) School of Computer Science and Technology(计算机科学与技术学院) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

AI总结 本文提出WiFi-GEN,通过生成式人工智能将测量到的WiFi功率转换为高分辨率室内图像,其形状重建精度比物理模型方法高275%,且Frechet Inception Distance得分降低82%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21280 2026-03-24 cs.CY cs.AI

WARBENCH: A Comprehensive Benchmark for Evaluating LLMs in Military Decision-Making

WARBENCH:评估大语言模型在军事决策中的综合基准

Zongjie Li, Chaozheng Wang, Yuchong Xie, Pingchuan Ma, Shuai Wang

机构 * Hong Kong University of Science and Technology(香港科学与技术大学) Chinese University of Hong Kong(香港中文大学) Zhejiang University of Technology(浙江工业大学)

AI总结 本文提出WARBENCH基准,揭示现有大语言模型在军事决策中存在结构性缺陷,包括战术推理崩溃、法律违规率高、量化降维导致性能下降等问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21269 2026-03-24 cs.RO

DyGeoVLN: Infusing Dynamic Geometry Foundation Model into Vision-Language Navigation

DyGeoVLN: 将动态几何基础模型注入视觉-语言导航

Xiangchen Liu, Hanghan Zheng, Jeil Jeong, Minsung Yoon, Lin Zhao, Zhide Zhong, Haoang Li, Sung-Eui Yoon

机构 * KAIST(韩国科学技术院) HKUST(GZ)(香港科技大学(广州)) JD Explore Academy(JD探索学院)

AI总结 本文提出DyGeoVLN框架,通过跨分支特征融合将动态几何基础模型注入视觉-语言导航,实现显式3D空间表示和视觉-语义推理,并引入新的无姿态自适应分辨率令牌剪枝策略以提升效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21169 2026-03-24 cs.LG

Model Evolution Under Zeroth-Order Optimization: A Neural Tangent Kernel Perspective

零阶优化下的模型演化:神经切线核视角

Chen Zhang, Yuxin Cheng, Chenchen Ding, Shuqi Wang, Jingreng Lei, Runsheng Yu, Yik-Chung WU, Ngai Wong

机构 * The University of Hong Kong(香港大学) The Hong Kong University of Science and Technology(香港科技大学)

AI总结 本文从神经切线核视角探讨零阶优化下的模型演化,证明线性模型中预期NZK恒定,提出通过NZK解释零阶更新的核梯度下降,实验证实理论结果并展示加速效果。

Comments ICLR 2026 Workshop on Scientific Methods for Understanding Deep Learning (20 pages, 18 figures)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21073 2026-03-24 eess.AS cs.CL cs.SD

SqueezeComposer: Temporal Speed-up is A Simple Trick for Long-form Music Composing

SqueezeComposer:长篇音乐创作中的时间压缩是一种简单技巧

Jianyi Chen, Rongxiu Zhong, Shilei Zhang, Kun Qian, Jinglei Liu, Yike Guo, Wei Xue

机构 * The Hong Kong University of Science and Technology(香港科技大学) JIUTIAN Research of China Mobile(中国移动JIUTIAN研究所) Beijing Institute of Technology(北京理工大学) China Mobile (Hong Kong) Innovation Research Institute(中国移动(香港)创新研究院)

AI总结 本文提出通过时间压缩技术简化长篇音乐生成,利用加速生成与恢复的方法降低资源消耗,实现高效高质量的音乐创作。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10448 2026-03-24 cs.RO

DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control

DiT4DiT:联合建模视频动态与动作以实现通用机器人控制

Teli Ma, Jia Zheng, Zifan Wang, Chunli Jiang, Andy Cui, Junwei Liang, Shuo Yang

机构 * Mondo Robotics(Mondo机器人公司) HKUST(GZ)(香港科技大学(广州)) HKUST(香港科技大学)

AI总结 本文提出DiT4DiT模型,通过结合视频扩散变换器与动作扩散变换器,实现视频动态与动作的联合建模,提升机器人控制的泛化能力与样本效率。

Comments https://dit4dit.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11026 2026-03-24 cs.CV

GIR-Bench: Versatile Benchmark for Generating Images with Reasoning

GIR-Bench:用于生成图像的多功能基准

Hongxiang Li, Yaowei Li, Bin Lin, Yuwei Niu, Yuhang Yang, Xiaoshuang Huang, Jiayin Cai, Xiaolong Jiang, Yao Hu, Long Chen

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Peking University(北京大学) University of Science and Technology of China(中国科学技术大学) Xiaohongshu Inc.(小红书公司)

AI总结 本文提出GIR-Bench,从理解-生成一致性、文本到图像生成和多步骤编辑三个角度评估统一模型,揭示其在复杂视觉任务中的表现与不足。

Comments ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24817 2026-03-24 cs.CV

UP2You: Fast Reconstruction of Yourself from Unconstrained Photo Collections

UP2You:从无约束照片集合中快速重建自己

Zeyu Cai, Ziyang Li, Xiaoben Li, Boqian Li, Zeyu Wang, Zhenyu Zhang, Yuliang Xiu

机构 * Westlake University(西湖大学) Nanjing University(南京大学) The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州))

AI总结 UP2You通过数据校正方法高效处理无结构照片,实现高保真的3D人物重建,无需预设模板,提升几何精度和纹理保真度。

Comments Page: https://zcai0612.github.io/UP2You Code: https://github.com/zcai0612/UP2You

Journal ref International Conference on Learning Representations (ICLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15340 2026-03-24 cs.LG

SSR: Speculative Parallel Scaling Reasoning in Test-time

SSR:测试时的推测并行扩展推理

Yuanlin Chu, Bo Wang, Xiang Liu, Hong Chen, Aiwei Liu, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Tsinghua University(清华大学)

AI总结 SSR提出一种无需训练的框架,通过引入分步推测解码加速推理而不牺牲正确性,提升数学推理任务的效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08947 2026-03-24 cs.LG cs.AI

Meta-Transfer Learning Powered Temporal Graph Networks for Cross-City Real Estate Appraisal

元迁移学习驱动的时序图网络用于跨城市房地产评估

Weijia Zhang, Jindong Han, Hao Liu, Wei Fan, Hao Wang, Hui Xiong

机构 * The Hong Kong University of Science and Technology(香港科技大学) Shandong University(山东大学) Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心)

AI总结 本文提出MetaTransfer,通过元迁移学习将多个数据丰富的城市知识迁移到数据稀缺城市,提升房地产估值性能,采用时序图网络和多任务学习模块实现跨城市知识迁移。

Comments Accepted by TIST 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01749 2026-03-24 cs.CY cs.AI cs.LG

Towards Urban General Intelligence: A Review and Outlook of Urban Foundation Models

迈向城市通用智能:城市基础模型的综述与展望

Weijia Zhang, Jindong Han, Zhao Xu, Hang Ni, Tengfei Lyu, Hao Liu, Hui Xiong

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州)) Shandong University(山东大学) Shandong University China(山东大学中国)

AI总结 本文综述了城市基础模型的发展,探讨了其定义、挑战及应用前景,提出了一种数据导向的分类框架,并汇总了相关基准和数据集,以推动城市通用智能的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.06135 2026-03-24 cs.GT cs.IR cs.LG

No-Regret Bayesian Recommendation to Homogeneous Users

无遗憾的贝叶斯推荐给同质用户

Yiding Feng, Wei Tang, Haifeng Xu

机构 * Hong Kong University of Science and Technology(香港科技大学) Chinese University of Hong Kong(香港中文大学) University of Chicago(芝加哥大学)

AI总结 研究在线贝叶斯推荐问题,设计无Stackelberg遗憾的推荐策略,实现双对数级遗憾依赖于回合数,并通过线性规划优化展示多项式依赖于状态数的推荐策略。

Comments Accepted by OR'26, conference version in EC'22

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10371 2026-03-24 cs.LG

Geometric Imbalance in Semi-Supervised Node Classification

半监督节点分类中的几何失衡

Liang Yan, Shengzhong Zhang, Bisheng Li, Menglin Yang, Chen Yang, Min Zhou, Weiyang Ding, Yutong Xie, Zengfeng Huang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) MBZUAI Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Logs AI Project(Logs AI项目)

AI总结 本文提出几何失衡概念,通过伪标签对齐、节点重排和模糊过滤缓解类别不平衡问题,实验表明在严重类别不平衡下性能优于现有方法。

Comments Accepted by NeurIPS 2025

Journal ref Proceedings of the Thirty-ninth Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏