arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

The University of Hong Kong(香港大学)

2026-06-11 至 2026-06-11 共收录 6
2606.12217 2026-06-11 cs.CV cs.AI cs.RO 新提交

Making Foresight Actionable: Repurposing Representation Alignment in World Action Models

使远见可操作:在世界动作模型中重新利用表示对齐

Lu Qiu, Yizhuo Li, Yi Chen, Yuying Ge, Yixiao Ge, Xihui Liu

机构 * The University of Hong Kong(香港大学) XPENG Robotics(小鹏机器人)

AI总结 针对世界动作模型中视觉预测与动作提取不匹配的问题,提出AGRA方法,通过对齐视频扩散特征与语义表示,提升动作解码器对任务相关区域的关注,从而改善操作任务的性能与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12065 2026-06-11 cs.AI cs.MA 新提交

Automating Geometry-Intensive Compliance Checking in BIM: Graph-Based Semantic Reasoning Framework

BIM中几何密集型合规检查自动化:基于图的语义推理框架

Zixuan Xiao, Pei Troh Koh, Jun Ma, Jack C. P. Cheng

机构 * Department of Urban Planning and Design, The University of Hong Kong(香港大学城市规划与设计系) Department of Civil and Environmental Engineering, The Hong Kong University of Science and Technology(香港科学与技术大学土木与环境工程系)

AI总结 针对BIM中几何密集型法规自动检查的语义鸿沟问题,提出SGR-BIM图驱动推理框架,通过跨模态知识图谱实现可解释推理,在679个消防规范查询上达到84.3%准确率,较基线提升8.6%。

Journal ref Automation in Construction 189 (2026) 107038

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11602 2026-06-11 cs.CV 新提交

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning

面向音视频广义零样本学习的层次化标准化嵌入对齐

Zihan Zhang, Jie Hong, Siyuan Fan, Yanghao Zhou, Pengfei Fang

机构 * Southeast University(东南大学) The University of Hong Kong(香港大学) Beijing Institute of Technology(北京理工大学) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education(新一代人工智能技术及其跨学科应用重点实验室(东南大学),教育部) School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)

AI总结 提出AHSE方法,通过Z-score标准化和层次化对齐策略(语义、类别、批次三级)解决音视频与文本模态间的分布与结构差异,在三个基准数据集上取得竞争性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11260 2026-06-11 cs.SD cs.AI 新提交

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark

RAIL: 基于CHC框架重新思考大型音频语言模型中的听觉智能

Hongyu Jin, Siyi Wang, Yang Xiao, Jiaheng Dong, Shihong Tan, Kaiyuan peng, Georgiana Juravle, Shanquan Chen, Gongping Huang, Hong Jia, Eun-Jung Holden, James Bailey, Ting Dang

机构 * School of Computing and Information Systems, The University of Melbourne(墨尔本大学计算与信息系统学院) Faculty of Psychology and Educational Sciences, Alexandru Ioan Cuza University of Iași(亚历山德鲁伊万库扎大学心理学与教育科学学院) School of Electronic Information, Wuhan University(武汉大学电子信息学院) School of Public Health, The University of Hong Kong(香港大学公共卫生学院) School of Computer Science, The University of Auckland(奥克兰大学计算机科学学院) Department of Data Science and Artificial Intelligence, Monash University(莫纳什大学数据科学与人工智能系)

AI总结 提出RAIL基准,基于CHC认知框架将听觉智能分解为五种核心能力,构建结构化评估任务,系统评测大型音频语言模型的认知行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10040 2026-06-11 cs.RO 版本更新

Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination

Efficient-WAM: 一种具有低成本未来想象能力的10亿参数世界-动作模型

Jiajun Li, Tiecheng Guo, Yifan Ye, Rongyu Zhang, Xiaowei Chi, Qianpu Sun, Ying Li, Yunfan Lou, Yan Huang, Zhihe Lu, Meng Guo, Shanghang Zhang

机构 * The University of Hong Kong(香港大学) Peking University(北京大学) Muka Robotics(Muka机器人) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Nanjing University(南京大学)

AI总结 提出Efficient-WAM,通过紧凑视频专家、稀疏视频潜变量和非对称去噪降低未来想象成本,在保持控制性能的同时实现30倍推理加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08211 2026-06-11 cs.LG 版本更新

MobileFineTuner: A Mobile-Native Framework for On-Device LLM Fine-Tuning in Real-World Embedded AI Applications

MobileFineTuner:面向真实世界嵌入式AI应用中设备端大语言模型微调的移动原生框架

Jiaxiang Geng, Lunyu Zhao, Yiyi Lu, Bing Luo

机构 * Duke Kunshan University(Duke昆山大学) The University of Hong Kong(香港大学)

AI总结 提出移动原生框架MobileFineTuner,通过C++实现资源感知训练运行时(内存高效注意力、激活检查点等),在商用手机上实现端到端LLM微调,显著降低内存压力并提升可执行性。

Comments 26 pages, 25 figures

详情

展开后加载摘要…

URL PDF HTML 收藏