LongVideoAgent: Multi-Agent Reasoning with Long Videos
LongVideoAgent: 多智能体推理与长视频
机构 * Hong Kong University of Science and Technology(香港理工大学)
AI总结 LongVideoAgent通过多智能体框架实现长视频的多模态推理,利用强化学习提升多智能体协作效率,优于现有基线方法。
高校专区
LongVideoAgent: 多智能体推理与长视频
机构 * Hong Kong University of Science and Technology(香港理工大学)
AI总结 LongVideoAgent通过多智能体框架实现长视频的多模态推理,利用强化学习提升多智能体协作效率,优于现有基线方法。
通过闭环世界建模实现视频虚拟角色的主动智能
机构 * The Hong Kong University of Science and Technology(香港科技大学) ; Meituan(美团) ; University of Science and Technology of China(中国科学技术大学)
AI总结 本文提出ORCA框架,通过闭环世界建模和双系统架构,实现视频虚拟角色的主动智能与目标导向行为。
Comments Project Page: https://xuanhuahe.github.io/ORCA/
Bohrium + SciMaster: 构建大规模智能科学的基础设施和生态系统
机构 * DP Technology Beijing China(北京DP技术有限公司) ; AI for Science Institute Beijing China(北京AI for Science研究院) ; Shanghai Jiao Tong University Shanghai China(上海交通大学) ; Beihang University Beijing China(北京航空航天大学) ; Peking University Beijing China(北京大学) ; Institute of Theoretical Physics Chinese Academy of Sciences Beijing China(中国科学院理论物理研究所) ; Shanghai Innovation Institute Shanghai China(上海创新研究院) ; East China Normal University Shanghai China(华东师范大学) ; Zhongguancun Academy Beijing China(中关村学院) ; University of Science and Technology of China Hefei China(中国科学技术大学) ; Tongji University Shanghai China(同济大学) ; The Hong Kong University of Science and Technology Hong Kong China(香港科技大学)
AI总结 Bohrium+SciMaster通过构建基础设施和生态系统,实现大规模智能科学的高效工作流编排与执行,显著提升科学产出效率和可追溯性。
Memory-T1: 基于强化学习的多会话代理时间推理
机构 * The Chinese University of Hong Kong(香港中文大学) ; Huawei Technologies Co.,Ltd(华为技术有限公司) ; HKUST(香港科技大学) ; The University of Edinburgh(爱丁堡大学)
AI总结 Memory-T1通过强化学习提升多会话代理的时间推理能力,实现7B模型在Time-Dialog基准上的67.0%高分,优于基线模型。
迈向灾害响应的生成性位置感知:一种概率跨视角地理定位方法
机构 * Department of Geography, National University of Singapore(新加坡国立大学地理系) ; Institute of Distributed Intelligent Systems, University of the Bundeswehr Munich(联邦国防军 Munich 分布式智能系统研究所) ; Professorship of Big Geospatial Data Management, School of Engineering and Design, Technical University of Munich(慕尼黑技术大学工程与设计学院大数据管理教授职位) ; School of Environment and Spatial Informatics, China University of Mining and Technology(中国矿业大学环境与空间信息学院) ; Interdisciplinary Center for Scientific Computing, Heidelberg University(海德堡大学跨学科科学计算中心) ; Heidelberg Institute for Geoinformation Technology, Heidelberg University(海德堡大学地理信息科技研究所) ; Urban Governance and Design Thrust, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)城市治理与设计研究组) ; Department of Architecture, National University of Singapore(新加坡国立大学建筑系) ; Department of Real Estate, National University of Singapore(新加坡国立大学房地产系) ; College of Surveying and Geo-Informatics, Tongji University(同济大学测绘与地理信息学院)
AI总结 本文提出ProbGLC方法,通过概率与确定性地理定位模型结合,提升灾害响应中的位置感知和定位准确性。
交互式代理经验收集的环境缩放:综述
机构 * Hong Kong University of Science and Technology(香港科技大学) ; King’s College London(伦敦国王学院) ; DAMO Academy, Alibaba Group(阿里云达摩院)
AI总结 本文综述了交互式代理经验收集中环境缩放的方法,从环境中心视角系统回顾了任务生成、执行和反馈阶段的代表性方法,并探讨了实现框架、挑战及未来研究方向。
Comments 22 pages, 5 figures, SEA Workshop @ NeurIPS 2025
BConformeR:一种基于互采样的构象模型,用于统一预测连续和不连续抗体结合位点
机构 * Material Innovation Institute for Life Sciences and Energy The University of Hong Kong Hetao SZ-HK Cooperation Zone(生命科学与能源创新材料研究所 香港大学 河套 SZ-HK 合作区) ; Hong Kong Generative AI Research and Development Center The Hong Kong University of Science and Technology(香港生成式人工智能研究与发展中心 香港科学大学) ; Computational Immunology Centre BayVax Biotech Limited(计算免疫学中心 BayVax 生物技术有限公司) ; School of Biomedical Sciences The University of Hong Kong(生物医学科学学院 香港大学)
AI总结 BConformeR通过结合CNN和Transformer,提升对连续和不连续抗体结合位点的预测性能。
Select2Reason: 为长链推理实现高效的指令微调数据选择
机构 * DataArc Tech Ltd.(DataArc科技有限公司) ; Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; IDEA Research, International Digital Economy Academy(IDEA研究所,国际数字经济学院) ; Hithink RoyalFlush Information Network Co., Ltd(Hithink RoyalFlush信息网络有限公司)
AI总结 Select2Reason通过高效数据选择提升长链推理性能,实验证明其在多个基准测试中表现优异。
视觉感知的CoT:在统一模型中实现高保真的视觉一致性
机构 * The Hong Kong University of Science and Technology(香港科技大学) ; Kling Team, Kuaishou Technology(快手科技 Kling 团队)
AI总结 本文提出视觉感知的CoT方法,通过自适应视觉规划和迭代视觉修正提升统一模型的多模态生成能力,实现更高的视觉一致性。
Comments Project Page: https://zixuan-ye.github.io/VACoT/
动态流网络用于变形医学图像配准中组合爆炸问题
机构 * Hong Kong University of Science and Technology(香港科技大学) ; Case Western Reserve University(凯斯西储大学) ; Hong Kong Metropolitan University(香港理工大学)
AI总结 本文提出动态流网络DySNet,通过自适应流盆地和动态流注意力机制解决变形医学图像配准中的组合爆炸问题,实现更高效的特征建模与匹配。
ChemATP: 一种无需训练的化学推理框架用于大语言模型
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Nanjing University of Aeronautics and Astronautics(南京航空航天大学) ; Shanghai AI Laboratory(上海人工智能实验室)
AI总结 ChemATP通过构建原子级文本知识库,使冻结的大语言模型能够动态检索和推理化学知识,从而在无需训练的情况下实现高效的化学推理。
VizDefender:通过主动定位和意图推断揭示可视化篡改
机构 * East China Normal University(华东师范大学) ; Hong Kong University of Science and Technology(香港科学与技术大学) ; School of Computer Science and Technology(计算机科学与技术学院)
AI总结 VizDefender通过半脆弱水印和意图分析模块,主动定位和推断可视化篡改,有效检测和分析数据篡改与视觉编码篡改。
Comments IEEE Transactions on Visualization and Computer Graphics (IEEE PacificVis'26 TVCG Track)
InSight-o3: 通过通用视觉搜索增强多模态基础模型
机构 * Hong Kong University of Science and Technology(香港科技大学) ; Huawei(华为)
AI总结 InSight-o3通过通用视觉搜索任务提升多模态基础模型的推理能力,解决复杂视觉信息整合问题。
InterMT:基于人类反馈的多轮交错偏好对齐
机构 * Institute for AI, Peking University(人工智能研究院,北京大学) ; State Key Laboratory of General Artificial Intelligence, Peking University(通用人工智能国家重点实验室,北京大学) ; Hong Kong University of Science and Technology(香港科技大学)
AI总结 InterMT通过多轮多模态交互的偏好数据集探索,旨在提升多模态大模型的交互能力,结合人类反馈和专家注释,揭示多轮扩展规律。
AdaCtrl: 通过难度感知预算分配实现适应性与可控性推理
机构 * Hong Kong University of Science and Technology(香港科技大学) ; The Chinese University of Hong Kong(香港中文大学) ; Peking University(北京大学)
AI总结 AdaCtrl通过难度感知预算分配实现自适应推理控制,提升模型在不同任务中的效率与效果。
GSRender: 通过弱监督的3D高斯点划法实现去重的占用预测
机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; Houmo AI ; Dalian University of Technology(大连理工大学) ; The Hong Kong University of Science and Technology(香港科学与技术大学)
AI总结 GSRender通过弱监督3D高斯点划法和射线补偿模块,有效减少占用预测中的重复问题,提升户外自动驾驶的感知性能。
AmPLe: 通过自适应去偏集成多提示学习支持视觉-语言模型
机构 * National Key Laboratory of Space Integrated Information System, Institute of Software, Chinese Academy of Sciences(中国科学院空间信息集成系统国家重点实验室,软件研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; The Hong Kong University of Science and Technology(香港科技大学)
AI总结 AmPLe通过自适应去偏集成多提示学习方法,解决模型-提示匹配偏差和样本-提示匹配偏差,提升视觉-语言模型在下游任务中的性能。
Comments Accepted by IJCV2025
面向社交媒体的抑郁症检测综述与实验:从大型语言模型和RAG到代理
机构 * The Hong Kong Polytechnic University(香港理工大学) ; The Hong Kong University of Science and Technology(香港科学与技术大学)
AI总结 本文综述并实验了基于社交媒体的抑郁症检测方法,探讨了LLM、RAG和代理系统在提升检测可靠性与推理能力中的应用。
Comments 20 pages, 10 figures. This is an extension of ICDEW 2025
MambaMIL+: 模型长期上下文模式以处理吉像素全切片图像
机构 * Department of Computer Science and Engineering, Hong Kong University of Science and Technology(计算机科学与工程系,香港科学与技术大学) ; Department of Pathology, Nanfang Hospital, School of Basic Medical Sciences, Southern Medical University(病理学系,南方医科大学) ; th Hospital of Joint Logistic Support Force, PLA(联合后勤保障部队900医院) ; Department of Pathology, Zhujiang Hospital, Southern Medical University(病理学系,南方医科大学) ; Department of Computer Science and Engineering, the Department of Chemical and Biological Engineering, the Division of Life Science, the State Key Laboratory of Nervous System Disorders, Hong Kong University of Science and Technology(计算机科学与工程系、化学与生物工程系、生命科学 division、神经系统紊乱国家重点实验室,香港科学与技术大学) ; HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute(香港科技大学深圳-香港协同创新研究院)
AI总结 MambaMIL+通过整合空间上下文和长距离依赖建模,提升全切片图像分析的效率和鲁棒性。
Comments 18 pages, 11 figures, 10 tables
EasyHOI: 释放大模型潜力以重建真实场景中的手-物体交互
机构 * The University of Hong Kong(香港大学) ; ShanghaiTech University(上海科技大学) ; Hong Kong University of Science and Technology(香港科学大学) ; Nanyang Technological University(南洋理工大学) ; Max Planck Institute for Informatics(马克斯·普朗克研究所(信息学)) ; Texas A&M University(德克萨斯大学奥斯汀分校)
AI总结 EasyHOI利用大模型实现单视角下手-物体交互的重建,通过先验引导优化提升重建精度。
Comments Project page: https://lym29.github.io/EasyHOI-page/
世界是你的画布:通过参考图像、轨迹和文本绘画可提示事件
机构 * HKUST(香港科技大学) ; Ant Group(蚂蚁集团) ; ZJU(浙江大学) ; NEU(南京大学) ; CUHK(香港中文大学) ; NTU(南洋理工大学)
AI总结 WorldCanvas通过结合文本、轨迹和参考图像,实现可提示的多代理交互事件生成,提升世界模型的交互性和可控性。
Comments Project page and code: https://worldcanvas.github.io/
StereoPilot: 通过生成先验学习统一且高效的立体转换
机构 * HKUST(GZ)(香港科技大学(广州)) ; HKUST(香港科技大学) ; Kling Team, Kuaishou Technology(快手科技 Kling 团队) ; CUHK(香港中文大学)
AI总结 StereoPilot通过生成先验学习实现高效立体转换,提升视觉保真度和计算效率。
N3D-VLM:原生3D接地使视觉-语言模型在3D场景中实现精确的空间推理
机构 * HKUST(香港科技大学) ; Tencent AI Lab(腾讯AI实验室) ; CUHK(香港中文大学) ; ZJU(浙江大学) ; NJU(南京大学)
AI总结 N3D-VLM通过原生3D感知能力提升视觉-语言模型在3D场景中的空间推理精度与解释性。
Comments Project Page: https://n3d-vlm.github.io
基于扩散的多模态3D物体检测在恶劣天气中的修复
机构 * School of Xingzhi College, South China Normal University(星智学院,华南师范大学) ; College of Big Data and Internet, Shenzhen Technology University(大数据与互联网学院,深圳科技大学) ; School of Data Science and Engineering, Xingzhi College, South China Normal University(数据科学与工程学院,星智学院,华南师范大学) ; Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(电子与计算机工程系,香港科技大学) ; College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)
AI总结 DiffFusion通过基于扩散的修复和自适应跨模态融合,提升多模态3D物体检测在恶劣天气中的鲁棒性与清洁数据性能。
少即是多:资源高效的低秩适应
机构 * University of Macau(澳门大学) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
AI总结 EffiLoRA通过统一A矩阵和动态B矩阵更新,在多种模型中实现更高效且鲁棒的低秩适应。
Comments 18 pages, 7 figures
BigCodeArena: 通过执行揭示更多可靠的代码生成人类偏好
机构 * Monash University(墨尔本大学) ; CSIRO’s Data61(CSIRO的数据61) ; Purdue University(普渡大学) ; Independent(独立) ; HKUST (Guangzhou)(香港科技大学(广州)) ; UCSD(加州大学圣地亚哥分校) ; UVA(弗吉尼亚大学) ; CNRS, France(法国国家科学研究中心) ; IBM ; Cisco(思科) ; Comenius University in Bratislava(布拉提斯拉瓦康门纽斯大学) ; University of Notre Dame(Notre Dame大学) ; Uber ; Tano Labs(Tano实验室) ; NUS(国立大学新加坡) ; Institute of Automation, CAS(中国科学院自动化研究所) ; Tencent AI Lab(腾讯AI实验室) ; University of Washington(华盛顿大学) ; Nevsky Collective(Nevsky集体) ; ETH Zurich(苏黎世联邦理工学院) ; Detomo Inc(Detomo公司) ; University of Oxford(牛津大学) ; UIUC(伊利诺伊大学香槟分校) ; Google(谷歌) ; NVIDIA ; Singapore Management University(新加坡管理学院) ; ServiceNow Research(ServiceNow研究) ; Hugging Face
AI总结 BigCodeArena通过执行揭示更多可靠的代码生成人类偏好,提出自动评分基准评估LLM编码质量。
Comments Built with love by the BigCode community :)
OS-Oracle: 一个跨平台GUI批评模型的综合框架
机构 * Shanghai Jiaotong University(上海交通大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; CUHK MMLab(香港大学MMLab) ; The University of Hong Kong(香港大学) ; The Hong Kong University of Science and Technology(香港科学与技术大学)
AI总结 OS-Oracle提出了一种跨平台GUI批评模型框架,通过合成数据和两阶段训练方法,实现了在移动、网页和桌面平台上的高效批评评估。
EMNLP: 教育者角色道德与规范大型语言模型画像
机构 * College of Education, Zhejiang University of Technology(浙江工业大学教育学院) ; The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州)) ; Faculty of Education, East China Normal University(华东师范大学教育学院) ; GuangHua Law School, ZheJiang University(浙江大学光华法学院) ; College of Computer Science and Technology, Zhejiang University of Technology(浙江工业大学计算机科学与技术学院)
AI总结 本文提出EMNLP框架,用于评估教育者角色LLM的伦理和心理一致性,通过构建88个教师特定道德困境,揭示教师角色LLM在道德推理和情感复杂情境中的表现差异。
Comments 29pages, 15 figures, Accepted by EMNLP Main Confrence
基于可更新空间记忆的视频生成
机构 * The University of Sydney(悉尼大学) ; Microsoft Research(微软研究院) ; HKUST(香港科技大学) ; University of Waterloo(滑铁卢大学)
AI总结 Spatia通过可更新的空间记忆机制,实现视频生成中的长期空间与时间一致性,支持动态实体生成和3D交互编辑。
Comments Project page: https://zhaojingjing713.github.io/Spatia/
Qwen-Image-Layered: 通过层分解实现内在可编辑性
机构 * HKUST(GZ)(香港科技大学(广州)) ; Alibaba(阿里巴巴)
AI总结 Qwen-Image-Layered通过层分解实现图像的内在可编辑性,提出端到端扩散模型及多阶段训练策略,提升图像分解质量与编辑一致性。
Comments 12 pages, 8 figures