arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

The Hong Kong University of Science and Technology(香港科技大学)

共收录 2801
2509.14981 2026-01-16 cs.CV

SPATIALGEN: Layout-guided 3D Indoor Scene Generation

SPATIALGEN: 布局引导的3D室内场景生成

Chuan Fang, Heng Li, Yixun Liang, Jia Zheng, Yongsen Mao, Yuan Liu, Rui Tang, Zihan Zhou, Ping Tan

机构 * Hong Kong University of Science and Technology(香港科技大学) Manycore Tech Inc(Manycore科技公司)

AI总结 SPATIALGEN通过布局引导的多视角多模态扩散模型生成高质量3D室内场景,解决现有方法在视觉质量、多样性及语义一致性方面的不足。

Comments 3D scene generation; diffusion model; Scene reconstruction and understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10124 2026-01-16 cs.CV

VQ-Seg: Vector-Quantized Token Perturbation for Semi-Supervised Medical Image Segmentation

VQ-Seg: 基于向量量化令牌扰动的半监督医学图像分割

Sicheng Yang, Zhaohu Xing, Lei Zhu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 VQ-Seg通过向量量化和量化扰动模块提升半监督医学图像分割性能,有效解决传统dropout正则化问题。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22972 2026-01-16 cs.CV eess.SP

Wavelet-based Multi-View Fusion of 4D Radar Tensor and Camera for Robust 3D Object Detection

基于小波的4D雷达张量与相机多视图融合用于鲁棒3D目标检测

Runwei Guan, Jianan Liu, Shaofeng Liang, Fangqiang Ding, Shanliang Yao, Xiaokai Bai, Daizong Liu, Tao Huang, Guoqiang Mao, Hui Xiong

机构 * Thrust of Artificial Intelligence, Hong Kong University of Science and Technology (Guangzhou)(人工智能 thrust,香港科技大学(广州)) Momoniai AI Department of Mechanical Engineering, Massachusetts Institute of Technology(机械工程系,麻省理工学院) School of Information Engineering, Yancheng Institute of Technology(信息工程学院,盐城科技学院) College of Information Science and Electronic Engineering, Zhejiang University(信息科学与电子工程学院,浙江大学) Institute for Math & AI, Wuhan University(数学与人工智能研究所,武汉大学) College of Science and Engineering and the Centre for AI and Data Science Innovation, James Cook University(科学与工程学院及人工智能与数据科学创新中心,詹姆斯库克大学) School of Transportation, Southeast University(交通运输学院,东南大学)

AI总结 WRCFormer通过小波注意力模块和几何引导渐进融合机制,高效融合4D雷达张量与相机图像,提升3D目标检测在恶劣天气下的鲁棒性。

Comments 10 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02064 2026-01-16 cs.CV

RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video

RTV-Bench: 通过实时视频对多模态大语言模型的连续感知、理解和推理进行基准测试

Shuhang Xun, Sicheng Tao, Jungang Li, Yibo Shi, Zhixin Lin, Zhanhui Zhu, Yibo Yan, Hanqian Li, Linghao Zhang, Shikang Wang, Yixin Liu, Hanbo Zhang, Ying Ma, Xuming Hu

机构 * HIT(哈尔滨工业大学) HKUST (GZ)(香港科技大学(广州)) HKUST(香港科技大学) XJTU(西安交通大学) SDU(山东大学) CityU(城市大学) HUST(华中科技大学)

AI总结 RTV-Bench通过实时视频对多模态大语言模型的连续感知、理解和推理能力进行细粒度基准测试,揭示了实时模型在长时段视频处理中的性能优势与局限。

Comments Accepted by NeurIPS 2025 Datasets and Benchmarks Track;

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07115 2026-01-16 cs.LG cs.AI math.OC

Online Scheduling for LLM Inference with KV Cache Constraints

在线调度用于具有KV缓存约束的LLM推理

Patrick Jaillet, Jiashuo Jiang, Konstantina Mellou, Marco Molinaro, Chara Podimata, Zijie Zhou

机构 * Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology(麻省理工学院电子工程与计算机科学系) HKUST(香港科技大学) Microsoft Research(微软研究院) Sloan School of Management, Massachusetts Institute of Technology(斯隆管理学院,麻省理工学院) Operations Research Center, Massachusetts Institute of Technology(运营研究中心,麻省理工学院)

AI总结 本文提出了一种在线调度算法,通过理论建模和实证验证,在管理KV缓存内存的同时最小化LLM推理延迟,显著优于现有基准算法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09385 2026-01-15 cs.SD cs.CL cs.MM

SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing

SLAM-LLM: 一种模块化、开源的多模态大语言模型框架及语音、语言、音频和音乐处理的最佳实践

Ziyang Ma, Guanrou Yang, Wenxi Chen, Zhifu Gao, Yexing Du, Xiquan Li, Zhisheng Zheng, Haina Zhu, Jianheng Zhuo, Zheshu Song, Ruiyang Xu, Tiranrui Wang, Yifan Yang, Yanqiao Zhu, Zhikang Niu, Liumeng Xue, Yinghao Ma, Ruibin Yuan, Shiliang Zhang, Kai Yu, Eng Siong Chng, Xie Chen

机构 * X-LANCE Lab, School of Computer Science, MoE Key Lab of Artificial Intelligence Shanghai Jiao Tong University(X-LANCE实验室,计算机科学学院,人工智能教育部重点实验室,上海交通大学) Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团) Peng Cheng Laboratory(鹏城实验室) University of Texas at Austin(德克萨斯大学奥斯汀分校) Tianjin University(天津大学) Hong Kong University of Science and Technology(香港科学大学) Queen Mary University of London(伦敦玛丽女王大学) Nanyang Technological University(南洋理工大学) Shanghai Innovation Institute(上海创新研究院)

AI总结 SLAM-LLM是一种开源多模态大语言模型框架,专注于语音、语言、音频和音乐处理,提供模块化配置和高性能检查点以加速研究开发。

Comments Published in IEEE Journal of Selected Topics in Signal Processing (JSTSP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09264 2026-01-15 cs.AI

Coordinated Pandemic Control with Large Language Model Agents as Policymaking Assistants

利用大语言模型代理进行协调的流行病防控

Ziyi Shi, Xusen Guo, Hongliang Lu, Mingxing Peng, Haotian Wang, Zheng Zhu, Zhenning Li, Yuxuan Liang, Xinhu Zheng, Hai Yang

机构 * The Hong Kong University of Science and Technology, Hong Kong(香港科学与技术大学) The Hong Kong University of Science and Technology (Guangzhou), China(香港科学与技术大学(广州)) Zhejiang University, China(浙江大学) University of Macau, Macau(澳门大学)

AI总结 本文提出利用大语言模型多代理系统进行协调流行病防控,通过模拟和闭环过程减少感染和死亡率。

Comments 20pages, 6 figures, a 60-page supporting material pdf file

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09231 2026-01-15 cs.RO

Online Trajectory Optimization for Arbitrary-Shaped Mobile Robots via Polynomial Separating Hypersurfaces

通过多项式分离超曲面实现任意形状移动机器人的在线轨迹优化

Shuoye Li, Zhiyuan Song, Yulin Li, Zhihai Bi, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 本文提出通过多项式分离超曲面实现任意形状移动机器人的非凸轨迹优化,解决传统方法对凸近似的保守假设问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09213 2026-01-15 cs.CV cs.AI

SpikeVAEDiff: Neural Spike-based Natural Visual Scene Reconstruction via VD-VAE and Versatile Diffusion

SpikeVAEDiff: 通过VD-VAE和多功能扩散实现神经尖峰基于的自然视觉场景重建

Jialu Li, Taiyan Zhou

机构 * HKUST Clear Water Bay(香港科技大学清水湾分校)

AI总结 SpikeVAEDiff通过VD-VAE和多功能扩散模型,利用神经尖峰数据实现高分辨率自然视觉场景重建,提升时间与空间分辨率,验证了特定脑区对重建质量的关键作用。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23675 2026-01-15 cs.CR cs.AI

QueryIPI: Query-agnostic Indirect Prompt Injection on Coding Agents

QueryIPI: 针对编码代理的查询无关间接提示注入

Yuchong Xie, Zesen Liu, Mingyu Luo, Zhixiang Zhang, Kaikai Zhang, Yuanyuan Yuan, Zongjie Li, Ping Chen, Shuai Wang, Dongdong She

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Fudan University(复旦大学) Tsinghua University(清华大学)

AI总结 QueryIPI通过利用系统不变性与工具描述,实现对编码代理的查询无关间接提示注入,实验显示其在模拟代理中达到87%的成功率,揭示了现实中的安全风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08519 2026-01-14 cs.CV cs.AI

CD^2: Constrained Dataset Distillation for Few-Shot Class-Incremental Learning

CD²:约束数据集蒸馏用于少样本类增量学习

Kexin Bao, Daichi Zhang, Hansong Zhang, Yong Li, Yutao Yue, Shiming Ge

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 CD²通过约束数据集蒸馏方法,在少样本类增量学习中有效缓解灾难性遗忘问题,提升模型对先前知识的保留能力。

Journal ref International Joint Conferences on Artificial Intelligence (IJCAI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08440 2026-01-14 cs.CV

Incentivizing Cardiologist-Like Reasoning in MLLMs for Interpretable Echocardiographic Diagnosis

激励MLLMs实现类似心内科医生的推理以实现可解释的超声心动图诊断

Yi Qin, Lehan Wang, Chenxu Zhao, Alex P. W. Lee, Xiaomeng Li

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) The Chinese University of Hong Kong(香港中文大学)

AI总结 本文提出CRT和CardiacMind,通过引入心内科医生的思维模式,提升MLLMs在超声心动图诊断中的推理能力,实现48%的诊断性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08107 2026-01-14 cs.LG cs.AI cs.SY eess.SY

STO-RL: Offline RL under Sparse Rewards via LLM-Guided Subgoal Temporal Order

STO-RL:通过LLM引导的子目标时间顺序实现稀疏奖励下的离线RL

Chengyang Gu, Yuxin Pan, Hui Xiong, Yize Chen

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州)) University of Alberta(阿尔伯塔大学)

AI总结 STO-RL通过LLM生成子目标时间顺序,结合潜在奖励塑造,提升稀疏奖励下离线RL的性能与稳定性。

Comments Accepted at International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07582 2026-01-14 cs.CL cs.AI

ES-Mem: Event Segmentation-Based Memory for Long-Term Dialogue Agents

基于事件分割的记忆:长期对话代理中的记忆

Huhai Zou, Tianhao Sun, Chuanjiang He, Yu Tian, Zhenyang Li, Li Jin, Nayu Liu, Jiang Zhong, Kaiwen Wei

机构 * College of Computer Science, Chongqing University(重庆大学计算机科学学院) Tsinghua University(清华大学) Hong Kong Generative AI Research & Development Center, HKUST(香港科技大学生成式人工智能研究与开发中心) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空航天信息研究所) School of Computer Science and Technology, Tiangong University(天津理工大学计算机科学与技术学院)

AI总结 ES-Mem通过动态事件分割和分层记忆架构,提升长期对话代理的记忆效率和上下文定位能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06789 2026-01-14 cs.SE cs.AI

MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences

MemGovern:通过学习受控的人类经验增强代码代理

Qihao Wang, Ziming Cheng, Shuo Zhang, Fan Liu, Rui Xu, Heng Lian, Kunyi Wang, Xiaoming Yu, Jianghao Yin, Sen Hu, Yue Hu, Shaolei Zhang, Yanbing Liu, Ronghao Chen, Huacan Wang

机构 * UCAS(中国科学院大学) NUS(新加坡国立大学) ECNU(华东师范大学) FDU(福建师范大学) XDU(西安电子科技大学) UBC(不列颠哥伦比亚大学) HKUST(GZ)(香港科技大学) PKU(北京大学) RUC(中国人民大学) QuantaAlpha(量子阿尔法)

AI总结 MemGovern通过学习受控的人类经验,提升代码代理的修复效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14507 2026-01-14 cs.AI cs.CL

DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extraction

DeKeyNLU:通过任务分解和关键词提取增强自然语言到SQL生成

Jian Chen, Zhenyan Chen, Xuming Hu, Peilin Zhou, Yining Hua, Han Fang, Cissy Hing Yee Choy, Xinmei Ke, Jingfeng Luo, Zixuan Yuan

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) HSBC(汇丰银行) South China University of Technology(华南理工大学) Harvard University(哈佛大学) Chicago University(芝加哥大学)

AI总结 DeKeyNLU通过任务分解和关键词提取提升NL2SQL生成精度,提出DeKeySQL流程并验证其在BIRD和Spider数据集上的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19804 2026-01-14 cs.CL

Compliance-to-Code: Enhancing Financial Compliance Checking via Code Generation

合规至代码:通过代码生成增强财务合规检查

Siyuan Li, Jian Chen, Rui Yao, Xuming Hu, Peilin Zhou, Weihua Qiu, Simin Zhang, Chucheng Dong, Zhiyao Li, Qipeng Xie, Zixuan Yuan

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Sun Yat-Sen University(孙中山大学) University of California, Riverside(加州大学河滨分校)

AI总结 Compliance-to-Code通过构建大规模中文金融监管数据集,提升金融合规检查的自动化水平,采用代码生成技术实现法规结构化和合规逻辑验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00365 2026-01-14 cs.LG

ROSS: RObust decentralized Stochastic learning based on Shapley values

ROSS:基于Shapley值的鲁棒去中心化随机学习

Lina Wang, Yunsheng Yuan, Feng Li, Lingjie Duan

机构 * School of Computer Science and Technology, Shandong University(计算机科学与技术学院,山东大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 本文提出基于Shapley值的ROSS算法,通过加权导数提升去中心化学习的鲁棒性和收敛效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09838 2026-01-14 cs.CV cs.AI

ClimateIQA: A New Dataset and Benchmark to Advance Vision-Language Models in Meteorology Anomalies Analysis

ClimateIQA: 一种新的数据集和基准,以推进气象异常分析中的视觉-语言模型

Jian Chen, Peilin Zhou, Yining Hua, Dading Chong, Meng Cao, Yaowei Li, Wei Chen, Bing Zhu, Junwei Liang, Zixuan Yuan

机构 * Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(人工智能研究所,香港科学与技术大学(广州)) Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)(数据科学与分析研究所,香港科学与技术大学(广州)) Harvard University(哈佛大学) Peking University(北京大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

AI总结 ClimateIQA通过引入SPOT算法和新型数据集,提升视觉-语言模型在气象异常分析中的准确性和表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.13620 2026-01-14 eess.IV cs.CV

Generative Adversarial Networks for Image Super-Resolution: A Survey

生成对抗网络用于图像超分辨率:综述

Ziang Wu, Xuanyu Zhang, Yinbo Yu, Qi Zhu, Jerry Chun-Wei Lin, Chunwei Tian

机构 * School of Engineering, The Hong Kong University of Science and Technology(香港科技大学工程学院) School of Software, Northwestern Polytechnical University(西北工业大学软件学院) College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院) Department of Distributed Systems and IT Devices, Silesian University of Technology(桑特大学技术学院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

AI总结 本文综述了生成对抗网络在图像超分辨率中的应用,分析了不同GAN变种的性能,并探讨了SISR中的挑战与未来研究方向。

Comments 32 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07779 2026-01-13 cs.MA cs.AI cs.CL cs.CV cs.HC

OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent

OS-Symphony: 一个用于鲁棒且通用计算机使用代理的综合框架

Bowen Yang, Kaiming Jin, Zhenyu Wu, Zhaoyang Liu, Qiushi Sun, Zehao Li, JingJing Xie, Zhoumianze Liu, Fangzhi Xu, Kanzhi Cheng, Qingyun Li, Yian Wang, Yu Qiao, Zun Wang, Zichen Ding

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai AI Laboratory(上海人工智能实验室) National University of Singapore(新加坡国立大学) The Hong Kong University of Science and Technology(香港科学与技术大学) The University of Hong Kong(香港大学) CUHK MMLab(香港中文大学MMLab) Xi’an Jiaotong University(西安交通大学) Nanjing University(南京大学) Harbin Institute of Technology(哈尔滨工业大学)

AI总结 OS-Symphony通过反思记忆代理和多功能工具代理,提升计算机使用代理在长周期任务和新领域中的鲁棒性和泛化能力。

Comments 31 pages, 11 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07558 2026-01-13 cs.RO

FlyCo: Foundation Model-Empowered Drones for Autonomous 3D Structure Scanning in Open-World Environments

FlyCo:基于基础模型的无人机自主3D结构扫描系统

Chen Feng, Guiyong Zheng, Tengkai Zhuang, Yongqian Wu, Fangzhan He, Haojia Li, Juepeng Zheng, Shaojie Shen, Boyu Zhou

机构 * Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(香港科技大学电子与计算机工程系) School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院) Department of Mechanical and Energy Engineering, Southern University of Science and Technology(南方科技大学机械与能源工程系) Differential Robotics, Hangzhou, China(杭州差分机器人)

AI总结 FlyCo通过整合基础模型实现无人机自主3D扫描,提升开放世界环境下的目标定位与预测效率。

Comments 34 pages, 24 figures, 9 tables. Video: https://www.youtube.com/playlist?list=PLqjZjnqsCyl40rw3y15Yzc7Mdo-z1y2j8

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07518 2026-01-13 cs.CV cs.AI

Mon3tr: Monocular 3D Telepresence with Pre-built Gaussian Avatars as Amortization

Mon3tr: 单目3D远程存在与预构建高斯人偶作为记忆化

Fangyu Lin, Yingdong Hu, Zhening Liu, Yufan Zhuang, Zehong Lin, Jun Zhang

机构 * Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(电子与计算机工程系,香港科技大学)

AI总结 Mon3tr通过单目3DGS技术实现远程存在,利用预构建的高斯人偶降低系统复杂性,实现实时高精度3D可视化与低带宽传输。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07347 2026-01-13 cs.CL

DiffER: Diffusion Entity-Relation Modeling for Reversal Curse in Diffusion Large Language Models

DiffER: 用于扩散大语言模型中反转诅咒的扩散实体-关系建模

Shaokai He, Kaiwen Wei, Xinyi Zeng, Xiang Chen, Xue Yang, Zhenyang Li, Jiang Zhong, Yu Tian

机构 * Chongqing University(重庆大学) Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Hong Kong University of Science and Technology(香港理工大学)

AI总结 DiffER通过实体感知训练和平衡数据构建,解决扩散大语言模型中的反转诅咒问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07296 2026-01-13 cs.AI cs.CL

LRAS: Advanced Legal Reasoning with Agentic Search

LRAS: 基于代理搜索的先进法律推理

Yujin Zhou, Chuxue Cao, Jinluan Yang, Lijun Wu, Conghui He, Sirui Han, Yike Guo

机构 * Hong Kong University of Science and Technology(香港科技大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

AI总结 LRAS通过整合反思模仿学习和难度感知强化学习,使大推理模型能够识别知识边界并处理法律推理的复杂性,从而在法律领域取得显著性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07023 2026-01-13 cs.AI

CloneMem: Benchmarking Long-Term Memory for AI Clones

CloneMem:用于AI克隆的长期记忆基准测试

Sen Hu, Zhiyu Zhang, Yuxiang Wei, Xueran Han, Zhenheng Tang, Huacan Wang, Ronghao Chen

机构 * Peking University(北京大学) UC Davis(加州大学戴维斯分校) Georgia Tech(佐治亚理工学院) MBZUAI(穆桑人工智能研究所) HKUST(香港科技大学) UCAS(中国科学技术大学)

AI总结 CloneMem是一个用于评估AI克隆长期记忆能力的基准,通过非对话数字痕迹数据,评估代理对个人状态演变的跟踪能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05635 2026-01-13 cs.CR cs.CL

Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs

在加密合成数据上进行持续预训练以实现隐私保护的大语言模型

Honghao Liu, Xuhui Jiang, Chengjin Xu, Cehao Yang, Yiran Cheng, Lionel Ni, Jian Guo

机构 * The PII shown in the figure is synthetic and not real(合成的PII) International Digital Economy Academy(国际数字经济学院) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学) DataArc Tech Ltd(DataArc科技有限公司)

AI总结 本文提出了一种基于实体的加密合成数据预训练框架,通过加密保护PII,实现隐私保护的大语言模型预训练,同时保持模型的指令遵循能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05680 2026-01-13 cs.LG cs.AI

Learning Design-Score Manifold to Guide Diffusion Models for Offline Optimization

学习设计-分数流形以指导扩散模型进行离线优化

Tailin Zhou, Zhilin Chen, Wenlong Lyu, Zhitang Chen, Danny H. K. Tsang, Jun Zhang

机构 * The Hong Kong University of Science and Technology, Hong Kong, China(香港科学与技术大学(香港)) The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China(香港科学与技术大学(广州)) Huawei Technologies Co., Ltd.(华为技术有限公司)

AI总结 本文提出ManGO,一种基于扩散模型的框架,通过学习设计-分数流形指导离线优化,实现超越训练数据的泛化能力,并在多个领域中优于现有方法。

Comments This manuscript was accepted by npj AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13772 2026-01-13 cs.SD cs.AI cs.LG cs.MM eess.AS

Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models

Jailbreak-AudioBench: 对大型音频语言模型中 jailbreak 威胁的深入评估与分析

Hao Cheng, Erjia Xiao, Jing Shao, Yichi Wang, Le Yang, Chao Shen, Philip Torr, Jindong Gu, Renjing Xu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Oxford(牛津大学) Xi’an Jiaotong University(西安交通大学) Hong Kong University of Science and Technology(香港科技大学) Northeastern University(东北大学) Beijing University of Technology(北京理工大学)

AI总结 Jailbreak-AudioBench 通过构建工具箱、数据集和基准,深入评估大型音频语言模型中 jailbreak 威胁,并促进安全防护机制的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06781 2026-01-13 cs.HC cs.AI cs.CV

AutoTour: Automatic Photo Tour Guide with Smartphones and LLMs

AutoTour:基于智能手机和LLMs的自动照片导览系统

Huatao Xu, Zihe Liu, Zilin Zeng, Baichuan Li, Mo Li

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

AI总结 AutoTour利用智能手机和LLMs自动为用户照片生成细粒度地标注释和描述,实现可扩展且上下文感知的交互式探索体验。

Comments 21

详情

展开后加载摘要…

URL PDF HTML 收藏