arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

The Hong Kong University of Science and Technology(香港科技大学)

共收录 2801
2508.10963 2025-12-09 cs.CV

EVCtrl: Efficient Control Adapter for Visual Generation

EVCtrl: 用于视觉生成的高效控制适配器

Zixiang Yang, Yue Ma, Yinhan Zhang, Shanhui Mo, Dongrui Liu, Linfeng Zhang

机构 * UESTC(电子科技大学) HKUST(香港科技大学) SJTU(上海交通大学)

AI总结 EVCtrl通过轻量级控制适配器提升视觉生成效率,实现高效可控生成,无需重新训练模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18886 2025-12-09 cs.LG cs.AI

A Survey on Diffusion Models for Time Series and Spatio-Temporal Data

时间序列和时空数据扩散模型综述

Yiyuan Yang, Ming Jin, Haomin Wen, Chaoli Zhang, Yuxuan Liang, Lintao Ma, Yi Wang, Chenghao Liu, Bin Yang, Zenglin Xu, Shirui Pan, Qingsong Wen

机构 * University of Oxford(牛津大学) Griffith University(格里菲斯大学) Carnegie Mellon University(卡内基梅隆大学) Zhejiang Normal University(浙江师范大学) Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The University of Hong Kong(香港大学) Salesforce Research(Salesforce研究) East China Normal University(东华大学) Fudan University(复旦大学)

AI总结 本文综述了扩散模型在时间序列和时空数据中的应用,系统梳理了模型类别、任务类型及实际应用场景,为后续研究提供基础。

Comments Accepted by ACM Computing Surveys; 37 pages; Github Repo: https://github.com/yyysjz1997/Awesome-TimeSeries-SpatioTemporal-Diffusion-Model

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05409 2025-12-08 cs.CL

SQ-format: A Unified Sparse-Quantized Hardware-friendly Data Format for LLMs

SQ-format: 一种统一的稀疏量化硬件友好数据格式用于大语言模型

Ruixuan Huang, Hao Zeng, Hantao Huang, Jinyuan Shi, Minghui Yu, Ian En-Hsu Yen, Shuai Wang

机构 * HKUST(香港科技大学) Moffett AI ByteDance Seed(字节跳动种子)

AI总结 本文提出SQ-format,一种统一的稀疏量化数据格式,旨在提升大语言模型在精度与效率之间的平衡,通过硬件友好设计实现性能与吞吐量的帕累托改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07978 2025-12-08 cs.CV

Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions

火星世界模型:可控视频合成与物理准确的3D重建

Longfei Li, Zhiwen Fan, Wenyan Cong, Xinhang Liu, Yuyang Yin, Matt Foutter, Panwang Pan, Chenyu You, Yue Wang, Zhangyang Wang, Yao Zhao, Marco Pavone, Yunchao Wei

机构 * BJTU(北京工业大学) UT Austin(德克萨斯大学奥斯汀分校) HKUST(香港科技大学) Stanford University(斯坦福大学) XMU(厦门大学) SBU(雪城大学) USC(南加州大学) NVIDIA(英伟达)

AI总结 本文提出M3arsSynth和MarsGen,通过物理准确的3D重建生成逼真的火星视频,提升任务模拟与机器人训练的可视化效果。

Comments Project Page: https://marsgenai.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05109 2025-12-08 cs.DB cs.AI

A Survey of Text-to-SQL in the Era of LLMs: Where are we, and where are we going?

大语言模型时代下的文本到SQL调研:我们目前在哪里,未来将去向何处?

Xinyu Liu, Shuyu Shen, Boyan Li, Peixian Ma, Runzhi Jiang, Yuxin Zhang, Ju Fan, Guoliang Li, Nan Tang, Yuyu Luo

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Renmin University of China(中国人民大学) Tsinghua University(清华大学)

AI总结 本文综述了大语言模型时代下文本到SQL技术的发展现状、核心方法及未来挑战。

Comments 20 pages, 11 figures, 3 tables

Journal ref TKDE July 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04864 2025-12-05 cs.AI

Are Your Agents Upward Deceivers?

您的代理是向上欺骗者吗?

Dadi Guo, Qingyu Liu, Dongrui Liu, Qihan Ren, Shuai Shao, Tianyi Qiu, Haoran Li, Yi R. Fung, Zhongjie Ba, Juntao Dai, Jiaming Ji, Zhikai Chen, Jialing Tao, Yaodong Yang, Jing Shao, Xia Hu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Hong Kong University of Science and Technology(香港科学与技术大学) Zhejiang University(浙江大学) Shanghai Jiao Tong University(上海交通大学) Peking University(北京大学) Alibaba Group(阿里巴巴集团)

AI总结 研究发现基于LLM的代理可能通过欺骗行为隐瞒失败,需加强缓解策略以确保安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04738 2025-12-05 cs.CL cs.AI cs.DB

OsmT: Bridging OpenStreetMap Queries and Natural Language with Open-source Tag-aware Language Models

OsmT: 通过开源标签感知语言模型连接OpenStreetMap查询与自然语言

Zhuoyue Wan, Wentao Hu, Chen Jason Zhang, Yuanfeng Song, Shuaimin Li, Ruiqiang Xiao, Xiao-Yong Wei, Raymond Chi-Wing Wong

机构 * The Hong Kong Polytechnic University(香港理工大学) WeBank Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) The Hong Kong University of Science and Technology(香港科技大学)

AI总结 OsmT通过开源标签感知语言模型连接自然语言与OpenStreetMap查询,提升查询生成和解释的准确性与可访问性。

Comments 42nd IEEE International Conference on Data Engineering (ICDE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04488 2025-12-05 cs.CV cs.AI cs.HC cs.MM

"I Can See Forever!": Evaluating Real-time VideoLLMs for Assisting Individuals with Visual Impairments

我能看到永远!

Ziyi Zhang, Zhen Sun, Zongmin Zhang, Zifan Peng, Yuemeng Zhao, Zichun Wang, Zeren Luo, Ruiting Zuo, Xinlei He

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 本文评估了实时视频LLMs在帮助视障者日常生活中的有效性,发现GPT-4o在任务成功率方面最高,并提出SafeVid数据集提升风险识别准确率至76%

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05661 2025-12-04 cs.CV

Language-Driven Object-Oriented Two-Stage Method for Scene Graph Anticipation

基于语言的面向对象两阶段方法用于场景图预见

Xiaomeng Zhu, Changwei Wang, Haozhe Wang, Xinyu Liu, Fangzhen Lin

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科学与技术大学) Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center, Qilu University of Technology(计算能力网络与信息安全重点实验室,教育部,山东计算机科学中心,齐鲁大学) Academy of Interdisciplinary Studies, The Hong Kong University of Science and Technology(跨学科研究学院,香港科学与技术大学)

AI总结 本文提出基于语言的面向对象两阶段方法,通过时间一致性正则化预测对象集动态和关系轨迹,显著提升视频场景图预见性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15644 2025-12-04 cs.CV cs.AI cs.CR

Can VLMs Detect and Localize Fine-Grained AI-Edited Images?

视觉语言模型能否检测并定位细粒度的人工智能编辑图像?

Zhen Sun, Ziyi Zhang, Zeren Luo, Zhiyuan Zhong, Zeyang Sha, Tianshuo Cong, Zheng Li, Shiwen Cui, Weiqiang Wang, Jiaheng Wei, Xinlei He, Qi Li, Qian Wang

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Ant Group(蚂蚁集团) Tsinghua University(清华大学) Shandong University(山东大学) Wuhan University(武汉大学)

AI总结 本文提出FragFake基准,首次系统研究VLMs在编辑图像分类和定位中的应用,发现微调模型在准确性上表现优异,同时探索了基于GRPO的RLVR训练方法。

Comments 14pages,19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03510 2025-12-04 cs.CV cs.RO

CSMapping: Scalable Crowdsourced Semantic Mapping and Topology Inference for Autonomous Driving

CSMapping: 可扩展的众包语义制图与拓扑推断用于自动驾驶

Zhijian Qiao, Zehuan Yu, Tong Li, Chih-Chung Chou, Wenchao Ding, Shaojie Shen

机构 * the Department of Electronic and Computer Engineering, the Hong Kong University of Science and Technology, Hong Kong, China(电子与计算机工程系,香港科技大学,香港,中国) the Academy for Engineering and Technology, Fudan University, Shanghai, China(工程与技术学院,复旦大学,上海,中国) Zhuoyu Technology Co., Ltd., Shenzhen, China(珠海优创科技有限公司,深圳,中国)

AI总结 CSMapping通过潜在扩散模型和拓扑优化方法,实现高质量的语义地图和道路中心线生成,适用于自动驾驶场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03080 2025-12-04 physics.chem-ph cs.AI cs.CE

AtomDisc: An Atom-level Tokenizer that Boosts Molecular LLMs and Reveals Structure--Property Associations

AtomDisc: 一种提升分子大语言模型并揭示结构-性质关联的原子级分词器

Mingxu Zhang, Dazhong Shen, Ying Sun

机构 * The Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(人工智能动力部,香港科学与技术大学(广州)) The College of Computer Science and Technology, The Nanjing University of Aeronautics and Astronautics(计算机科学与技术学院,南京航空航天大学)

AI总结 AtomDisc通过原子级分词提升分子LLM性能,揭示结构-性质关联。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19304 2025-12-04 cs.AI cs.CL cs.LG

AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning

AutoEnv:用于测量跨环境智能体学习的自动化环境

Jiayi Zhang, Yiran Peng, Fanqi Kong, Cheng Yang, Yifan Wu, Zhaoyang Yu, Jinyu Xiang, Jianhao Ruan, Jinlin Wang, Maojia Song, HongZhang Liu, Xiangru Tang, Bang Liu, Chenglin Wu, Yuyu Luo

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) DeepWisdom Peking University(北京大学) Singapore University of Technology and Design(新加坡科技设计大学) Sydney University(悉尼大学) Yale University(耶鲁大学) Université de Montréal & Mila(蒙特利尔大学及Mila)

AI总结 AutoEnv通过自动化生成异质环境和组件驱动的学习方法,探讨了跨环境智能体学习的挑战与扩展性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18212 2025-12-04 cs.AI cs.LG

A Definition of AGI

AGI 的定义

Dan Hendrycks, Dawn Song, Christian Szegedy, Honglak Lee, Yarin Gal, Erik Brynjolfsson, Sharon Li, Andy Zou, Lionel Levine, Bo Han, Jie Fu, Ziwei Liu, Jinwoo Shin, Kimin Lee, Mantas Mazeika, Long Phan, George Ingebretsen, Adam Khoja, Cihang Xie, Olawale Salaudeen, Matthias Hein, Kevin Zhao, Alexander Pan, David Duvenaud, Bo Li, Steve Omohundro, Gabriel Alfour, Max Tegmark, Kevin McGrew, Gary Marcus, Jaan Tallinn, Eric Schmidt, Yoshua Bengio

机构 * Center for AI Safety(AI安全中心) University of California, Berkeley(加州大学伯克利分校) Virtue AI Morph Labs(Morph实验室) University of Michigan(密歇根大学) LG AI Research(LG人工智能研究) University of Oxford(牛津大学) Stanford University(斯坦福大学) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Gray Swan AI Carnegie Mellon University(卡内基梅隆大学) Cornell University(康奈尔大学) Hong Kong Baptist University(香港 Baptist大学) HKUST(香港科技大学) Nanyang Technological University(南洋理工大学) KAIST(韩国科学技术院) University of California, Santa Cruz(加州大学圣克鲁兹分校) Massachusetts Institute of Technology(麻省理工学院) University of Tübingen(图宾根大学) University of Washington(华盛顿大学) University of Toronto(多伦多大学) Vector Institute(向量研究所) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Beneficial AI Research(有益AI研究) Conjecture Institute for Applied Psychometrics(应用心理测量研究所) New York University(纽约大学) CSER Université de Montréal(蒙特利尔大学) LawZero

AI总结 本文提出了一种基于卡特尔-霍恩-卡罗尔理论的可量化框架,定义AGI为与受过良好教育的成年人认知能力相匹配,并通过心理测量电池评估AI系统,揭示当前AI在基础认知机制上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02475 2025-12-04 cs.LG cs.CL

Comba: Improving Bilinear RNNs with Closed-loop Control

Comba:通过闭环控制改进双线性RNNs

Jiaxi Hu, Yongqi Pan, Jusen Du, Disen Lan, Xiaqiang Tang, Qingsong Wen, Yuxuan Liang, Weigao Sun

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Shanghai AI Laboratory(上海人工智能实验室) Squirrel Ai Learning, USA(Squirrel Ai Learning)

AI总结 Comba通过闭环控制理论改进双线性RNNs,采用标量加低秩状态转移实现高效语言和视觉建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04968 2025-12-04 cs.MM cs.AI

The Dream Within Huang Long Cave: AI-Driven Interactive Narrative for Family Storytelling and Emotional Reflection

黑龙洞之梦:由AI驱动的互动叙事用于家庭故事讲述与情感反思

Jiayang Huang, Lingjie Li, Kang Zhang, David Yip

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 《黑龙洞之梦》通过AI驱动的互动叙事,探索家庭关系中的情感连接与大他者不存在的概念。

Comments 8 pages,8 figures, International Symposium on Electronic/Emerging Art (ISEA)

Journal ref Proceedings of the International Symposium on Electronic/Emerging Art (ISEA 2025), Seoul, Republic of Korea, 2025, pp.247-254

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03046 2025-12-03 cs.CV

MagicQuillV2: Precise and Interactive Image Editing with Layered Visual Cues

MagicQuillV2: 基于分层视觉提示的精确交互式图像编辑

Zichen Liu, Yue Yu, Hao Ouyang, Qiuyu Wang, Shuailei Ma, Ka Leong Cheng, Wen Wang, Qingyan Bai, Yuxuan Zhang, Yanhong Zeng, Yixuan Li, Xing Zhu, Yujun Shen, Qifeng Chen

机构 * HKUST(香港科技大学) Ant Group(蚂蚁集团) NEU(南京大学) ZJU(浙江大学) CUHK(香港中文大学)

AI总结 MagicQuillV2通过分层视觉提示实现精确交互式图像编辑,解决传统生成模型与图形软件间的控制与语义平衡问题。

Comments Code and demo available at https://magicquill.art/v2/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02834 2025-12-03 cs.RO cs.AI

Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach

引导视觉-语言-动作模型作为反探索:一种测试时间缩放方法

Siyuan Yang, Yang Zhang, Haoran He, Ling Pan, Xiu Li, Chenjia Bai, Xuelong Li

机构 * Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院) University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

AI总结 本文提出TACO框架,通过测试时间缩放方法在视觉-语言-动作模型中引入反探索机制,提升推理稳定性和任务成功率。

Comments The first two authors contributed equally. Yang Zhang leads the whole project

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02807 2025-12-03 cs.CL

SR-GRPO: Stable Rank as an Intrinsic Geometric Reward for Large Language Model Alignment

SR-GRPO:稳定秩作为大语言模型对齐的内在几何奖励

Yixuan Tang, Yi Yang

机构 * The Hong Kong University of Science and Technology(香港理工大学)

AI总结 SR-GRPO通过稳定秩作为内在几何奖励信号,无需外部监督提升大语言模型对齐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09481 2025-12-03 cs.SE cs.AI cs.CL

Evaluating LLMs on Sequential API Call Through Automated Test Generation

通过自动化测试生成评估LLMs的顺序API调用

Yuheng Huang, Jiayang Song, Da Song, Zhenlan Ji, Wenhan Wang, Shuai Wang, Lei Ma

机构 * The University of Tokyo(东京大学) Macau University of Science and Technology(澳门科技大学) Shandong University(山东大学) Hong Kong University of Science and Technology(香港科技大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) University of Alberta(阿尔伯塔大学)

AI总结 本文提出StateGen框架,通过自动化测试生成评估LLMs在顺序API调用中的性能,构建了包含120个测试用例的StateEval基准测试,揭示了当前LLM在API整合方面的改进方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02536 2025-12-03 cs.CV

WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens

WeMMU: 通过噪声查询标记增强视觉-语言模型与扩散模型的桥梁

Jian Yang, Dacheng Yin, Xiaoxuan He, Yong Li, Fengyun Rao, Jing Lyu, Wei Zhai, Yang Cao, Zheng-Jun Zha

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知联合实验室,中国科学技术大学) ZheJiang University(浙江大学) The Hong Kong University of Science and Technology(香港科技大学)

AI总结 WeMMU通过噪声查询标记和VAE分支,提升视觉-语言模型与扩散模型的连接效率,缓解泛化崩溃问题,实现稳定持续学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01960 2025-12-02 cs.CV cs.HC

SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation

SpriteHand: 实时多样的手-物体交互与自回归视频生成

Zisu Li, Hengye Lyu, Jiaxin Shi, Yufeng Zeng, Mingming Fan, Hanwang Zhang, Chen Liang

机构 * HKUST (Guangzhou)(香港科技大学(广州)) XMax.AI Ltd.(XMax人工智能有限公司) Nanyang Technological University(南洋理工大学)

AI总结 SpriteHand通过自回归视频生成框架实现实时高质量手-物体交互视频合成,支持多种物体和运动模式,具有高视觉真实性和物理合理性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01677 2025-12-02 cs.CV

Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation

基于结构和接触感知表示的开放世界手-物体交互视频生成

Haodong Yan, Hang Yu, Zhide Zhong, Weilin Yuan, Xin Gong, Zehang Luo, Chengxi Heyu, Junfeng Li, Wenxuan Song, Shunbo Zhou, Haoang Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 本文提出一种结构和接触感知表示,用于生成逼真的手-物体交互视频,通过联合生成范式提升开放世界场景下的交互物理建模能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23127 2025-12-02 cs.CV

DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation

DualCamCtrl: 基于双分支扩散模型的几何感知摄像机控制视频生成

Hongfei Zhang, Kanghao Chen, Zixin Zhang, Harold Haodong Chen, Yuanhuiyi Lyu, Yuqi Zhang, Shuai Yang, Kun Zhou, Yingcong Chen

机构 * HKUST (GZ)(香港科技大学) HKUST(香港科技大学) Fudan University(复旦大学) Shenzhen University(深圳大学)

AI总结 DualCamCtrl通过双分支扩散模型实现几何感知的摄像机控制视频生成,有效提升视频生成的一致性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21309 2025-12-02 cs.CV

CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture Generation

CaliTex: 用于视图一致3D纹理生成的几何校准注意力

Chenyu Liu, Hongze Chen, Jingzhi Bao, Lingting Zhu, Runze Zhang, Weikai Chen, Zeyu Hu, Yingda Yin, Keyang Luo, Xin Wang

机构 * PKU(北京大学) HKUST(香港科技大学) CUHK(SZ)(香港中文大学(深圳)) HKU(香港大学) LIGHTSPEED

AI总结 CaliTex通过几何校准注意力机制,解决3D纹理生成中的跨视图不一致问题,提升纹理的视图一致性与生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11569 2025-12-02 cs.LG cs.IT eess.SP math.IT

Pre-Training and Personalized Fine-Tuning via Over-the-Air Federated Meta-Learning: Convergence-Generalization Trade-Offs

通过空中联邦元学习进行预训练和个性化微调:泛化与收敛的权衡

Haifeng Wen, Hong Xing, Osvaldo Simeone

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Department of ECE, The Hong Kong University of Science and Technology(香港科学与技术大学电子工程系) Department of Engineering, King’s College London(伦敦国王学院工程系)

AI总结 本文研究了无线环境下基于元学习的个性化联邦学习在泛化与收敛之间的权衡,通过空中计算分析了信道损伤对泛化和收敛的影响。

Comments 40 pages, 10 figures, to appear in IEEE Trans. Cogn. Commun. Netw

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00877 2025-12-02 cs.CV

Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling

前馈3D高斯散射压缩与长上下文建模

Zhening Liu, Rui Song, Yushi Huang, Yingdong Hu, Xinjie Zhang, Jiawei Shao, Zehong Lin, Jun Zhang

机构 * Hong Kong University of Science and Technology(香港科技大学) Institute of Artificial Intelligence (TeleAI)(人工智能研究所) China Telecom(中国电信)

AI总结 本文提出了一种前馈3DGS压缩方法,通过大规模上下文结构和自回归熵模型实现长距离依赖建模,达到20倍压缩比和先进性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00410 2025-12-02 cs.RO cs.AI

Balancing Efficiency and Fairness: An Iterative Exchange Framework for Multi-UAV Cooperative Path Planning

平衡效率与公平:多无人机协作路径规划的迭代交换框架

Hongzong Li, Luwei Liao, Xiangguang Dai, Yuming Feng, Rong Feng, Shiqin Tang

机构 * The Hong Kong University of Science and Technology(香港科技大学) College of Automation Engineering, Nanjing University of Aeronautics and Astronautics(南京航空航天大学自动化工程学院) College of Computer Science and Engineering, Chongqing Three Gorges University(重庆三峡大学计算机科学与工程学院) Chinese Academy of Sciences, New Territories, Hong Kong(中国科学院香港新 Territories 分院) City University of Hong Kong, Kowloon, Hong Kong(香港城市大学)

AI总结 本文提出一种迭代交换框架,用于多无人机协作路径规划,通过任务交换和路径细化平衡效率与公平性,实验表明其在总距离和makespan之间取得更优权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00395 2025-12-02 cs.CV

Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction

更优、更强、更快:在基于多模态大语言模型的分割中解决三重困境

Jiazhen Liu, Mingkuan Feng, Long Chen

机构 * The Hong Kong University of Science and Technology(香港科技大学)

AI总结 STAMP通过同时预测文本和分割掩码,解决了多模态大语言模型在保持对话能力、高分割性能和快速推理间的三重困境。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17442 2025-12-02 cs.CL cs.AI

Confident RAG: Enhancing the Performance of LLMs for Mathematics Question Answering through Multi-Embedding and Confidence Scoring

Confident RAG: 通过多嵌入和置信度评分提升LLM在数学问题回答中的性能

Shiting Chen, Zijian Zhao, Jinsong Chen

机构 * Faculty of Education, The University of Hong Kong(香港大学教育学院) Department of Civil and Environmental Engineering, The Hong Kong University of Science and Technology(香港科学与技术大学土木与环境工程系)

AI总结 Confident RAG通过多嵌入和置信度评分提升LLM在数学问题回答中的性能,实验表明其在准确率上比普通LLMs和RAG分别提升10%和5%。

详情

展开后加载摘要…

URL PDF HTML 收藏