FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation
FREPix:频率异质流匹配用于像素空间图像生成
机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
AI总结 FREPix通过分解低频和高频组件,设计显式生成流程,提升像素空间图像生成效果,实现1.91 FID在256×256和2.38 FID在512×512的性能。
高校专区
FREPix:频率异质流匹配用于像素空间图像生成
机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
AI总结 FREPix通过分解低频和高频组件,设计显式生成流程,提升像素空间图像生成效果,实现1.91 FID在256×256和2.38 FID在512×512的性能。
H$^2$SD:混合事后自蒸馏
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Harbin Institute of Technology(哈尔滨工业大学) ; Fudan University(复旦大学) ; The Chinese University of Hong Kong(香港中文大学)
AI总结 研究针对强化学习中奖励监督问题,提出 H$^2$SD 混合事后自蒸馏框架。成功轨迹用教师概率调更新幅度,失败轨迹基于参考提示调整教师并最小化反向 KL 散度,实验表明该框架优于基线,能稳定优化且效率良好。
多目标RTA拦截中的不确定性建模与蒸馏加速
机构 * Department of Applied Mathematics(应用数学系) ; Harbin Institute of Technology, Weihai(哈尔滨工业大学威海学院) ; School of Mathematics and Statistics(数学与统计学学院) ; Shandong University(山东大学)
AI总结 本文提出UMDA框架,结合多目标学习与不确定性建模,通过知识蒸馏减少计算开销,提升RTA拦截的效率和准确性。
RM-Distiller: 利用生成式大语言模型进行奖励模型蒸馏
机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) ; School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
AI总结 RM-Distiller通过系统利用教师LLM的精炼、评分和生成能力,提升奖励模型蒸馏效果,首次系统性地探索了生成式LLM在奖励建模中的应用。
Comments Accepted to IJCAI-ECAI 2026
UAV-ON:用于空中智能体的开放世界目标导航基准测试
机构 * Harbin Institute of Technology(哈尔滨工业大学)
AI总结 介绍用于空中智能体开放世界目标导航的UAV-ON基准测试,含14个环境及1270个目标对象,通过实例级指令编码。实现多种基线方法评估,结果凸显空中导航和语义目标基础的复合挑战,推动复杂环境下无人机可扩展自主性研究。
Comments Accepted to ACM MM 2025
LMEB:长周期记忆嵌入基准
机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) ; Shenzhen Loop Area Institute (SLAI)(深圳环湖研究院) ; Peking University(北京大学)
AI总结 本文提出LMEB基准,用于评估复杂长周期记忆检索任务,涵盖22个数据集和193个零样本任务,揭示大模型并不总优于小模型,且传统检索能力不等同于长周期记忆能力。
Comments 35 pages, 9 figures, 23 tables
基于双链接熵分析的自适应视觉自回归加速
机构 * Tongji University(同济大学) ; Harbin Institute of Technology(哈尔滨工业大学) ; Shandong Academy of Sciences(山东省科学院) ; Macquarie University(麦考瑞大学)
AI总结 NOVA通过熵分析提出无需训练的VAR模型token减少加速框架,通过动态调整尺度和层的token减少比率,实现推理加速与生成质量的平衡。
Comments 12 pages, 8 figures
ViMax: 智能体视频生成
机构 * The University of Hong Kong(香港大学) ; South China University of Technology(华南理工大学) ; Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
AI总结 提出ViMax框架,通过多智能体协作实现长视频生成,利用分层叙事引擎和视觉一致性机制,保证叙事连贯性和视觉一致性。
Comments 20 pages, 13 figures
引导、思考、行动:面向视觉-语言-动作模型的交互式具身推理
机构 * Futian Laboratory(福田实验室) ; Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) ; International Digital Economy Academy (IDEA)(国际数字经济学院(IDEA)) ; School of Robotics, Hunan University(湖南大学机器人学院) ; South China University of Technology(华南理工大学) ; Visincept(Visincept公司) ; National Key Laboratory of Smart Farm Technologies and Systems(智能农业技术与系统国家重点实验室)
AI总结 本文提出GTA-VLA框架,通过允许用户用显式视觉提示引导机器人策略,实现空间可调节的具身推理。该框架整合了外部指导与内部任务规划,提升了在视觉歧义和错误恢复中的表现。
CPGRec+: 一种面向个性化视频游戏推荐的平衡框架
机构 * Harbin Institute of Technology(哈尔滨工业大学) ; School of Computer Science, Wuhan University(武汉大学计算机学院)
AI总结 本文提出CPGRec+框架,通过引入偏好引导边加权和表示生成模块,解决传统推荐系统在准确率与多样性间的平衡问题,实验表明其在Steam数据集上表现更优。
Comments Published in ACM Transactions on Information Systems (TOIS). 43 pages, 9 figures
Journal ref ACM Trans. Inf. Syst. 44, 3, Article 66 (March 2026), 44 pages
AIM-CoT:基于主动信息的多模态链式推理用于视觉-语言推理
机构 * The Chinese University of Hong Kong(香港中文大学) ; Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
AI总结 本文提出AIM-CoT框架,通过上下文增强注意力图生成、主动视觉探测和动态注意力转移触发,改进视觉语言模型的证据选择与插入触发,提升视觉-语言推理性能。
Comments Accepted by ACL 2026 Main Conference. 30 pages, 6 figures
LEGO:用于合成图像检测的基于LoRA的面向生成器框架
机构 * University of Electronic Science and Technology of China(电子科技大学) ; Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
AI总结 针对生成技术使合成图像难辨真伪,现有检测方法存在泛化和过拟合问题,提出LEGO框架。通过MLP调制多个LoRA模块,分两阶段训练,能提取特定生成器独特伪影,在少数据和少轮次训练下性能优于现有方法。
Comments 10 pages,2 figures
WINO: 一种用于变域超弹性问题的弱形式物理信息神经算子
机构 * School of Science, Harbin Institute of Technology, Shenzhen, P. R. China(哈尔滨工业大学深圳校区) ; School of Science, Harbin Institute of Technology, Shenzhen, Guangdong(哈尔滨工业大学深圳校区) ; Institute of Structural Mechanics, Bauhaus-Universität Weimar(魏玛 Bauhaus 大学结构力学研究所)
AI总结 提出一种无数据框架WINO,结合神经算子的效率与φ-有限元法的几何灵活性,通过最小化弱形式残差和惩罚项训练,实现高精度且计算时间减少50-80%。
GAP-MLLM:几何对齐预训练以激活多模态大语言模型中的3D空间感知
机构 * Harbin Institute of Technology(哈尔滨工业大学) ; School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院) ; Huawei(华为)
AI总结 本文提出GAP-MLLM,通过几何对齐预训练激活多模态大语言模型中的3D空间感知,改进了传统方法在3D空间感知上的不足。
Comments Accepted by ECCV 2026. Project page: https://gapmllm.github.io/
MMGraphRAG: 通过可解释的多模态知识图谱弥合视觉与语言的鸿沟
机构 * Harbin Institute of Technology(哈尔滨工业大学) ; Nanyang Technological University(南洋理工大学)
AI总结 MMGraphRAG通过引入可解释的多模态知识图谱,解决多模态场景下的实体链接和跨模态推理问题,实现最先进的多模态信息处理性能。
VCG-Bench:迈向统一的视觉导向基准,用于结构化生成与编辑
机构 * The Hong Kong University of Science and Technology (GuangZhou)(香港科学与技术大学(广州)) ; Huawei Technologies Co., Ltd(华为技术有限公司) ; Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) ; South China University of Technology(华南理工大学)
AI总结 本文提出VCG-Bench,一个统一的视觉导向mxGraph任务基准,通过符号逻辑和XML实现精确的图表生成与编辑,解决现有方法在结构化任务中的局限性。
Comments Accepted by ICML2026, 37 pages, 10 figures
评估基于LLM的从零开始软件生成在端到端CLI工具场景中的表现
机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) ; Independent Researcher, China(独立研究者,中国) ; Singapore Management University, Singapore(新加坡管理大学)
AI总结 本文提出CLI-Tool-Bench基准,用于评估从零生成CLI工具的能力,发现顶级模型成功率低于43%,指出生成代码倾向单体化。
Comments Data link: https://github.com/kinesiatricssxilm14/CLI-Tool-Bench
具有类别重叠和无任务标识符的流式联邦持续学习中的知识感知进化
机构 * Faculty of Computing(计算学院) ; Harbin Institute of Technology(哈尔滨工业大学)
AI总结 FedKACE通过自适应推理模型切换、梯度平衡重放和核谱边界缓冲区维护,解决流式联邦持续学习中类别重叠和无任务标识符带来的知识混淆问题。
CXRAgent:用于胸部X光解读的由主任编排的多阶段推理
机构 * School of Computer Science, Hangzhou Dianzi University(杭州电子科技大学计算机科学学院) ; Department of Automation, Tsinghua University(清华大学自动化系) ; Department of Colorectal Medical Oncology, Zhejiang Cancer Hospital(浙江省肿瘤医院结直肠医学肿瘤科) ; School of Computer and Control Engineering, University of Chinese Academy of Sciences(中国科学院大学计算机与控制工程学院) ; School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)智能科学与工程学院)
AI总结 针对胸部X光解读中模型适应性差等问题,提出CXRAgent,通过中央主任协调多阶段,包括工具调用、诊断规划和协作决策,实验证明该方法性能出色,能提供视觉证据并适应不同临床任务。
Comments 10 pages, 4 figures, 7 Tables
QDA-SQL:用于多轮文本到SQL的问题增强对话扩充
机构 * School of Cyberspace Science Harbin Institute of Technology Harbin, China(网络空间科学学院 哈尔滨工业大学 哈尔滨,中国) ; School of Computer Science(计算机科学学院) ; Technology Harbin University of Science(技术 哈尔滨理工大学) ; School of Cyberspace Science Harbin Institute of Technology, Shenzhen Shenzhen, China(网络空间科学学院 哈尔滨工业大学深圳 哈尔滨,中国)
AI总结 该研究针对多轮文本到SQL任务中微调模型面临的问题,提出QDA-SQL数据扩充方法,利用大语言模型生成多轮问答对,并引入验证和校正机制,提升了微调模型在SQL语句准确性及处理复杂问题上的能力。
Comments CAAI International Conference on Artificial Intelligence 2026 (CICAI 2026)
KnowAct-GUIClaw:深度理解、完美执行,具有自我进化记忆和技能的个人GUI助手
机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
AI总结 针对OpenClaw的不足,提出深度理解、完美执行范式,介绍KnowAct-GUIClaw框架,通过多系统实验验证其在效率、准确性和跨平台适应性方面的优势,且知识记忆和执行技能可跨基础模型迁移。
Comments 29 pages, 9 figures
RADAR: 通过语义规划和自主因果环境重置实现闭环机器人数据生成
机构 * Southern University of Science and Technology(南方科技大学) ; Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) ; Spatialtemporal AI(时空人工智能)
AI总结 RADAR通过语义规划和自主因果环境重置实现闭环机器人数据生成,具备高适应性和可扩展性,在仿真和现实部署中均表现出卓越的性能。
Comments 8 pages, 4 figures. Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026). Project page: https://radar-iros.netlify.app/
从安卓到开源鸿蒙系统的图形用户界面测试迁移实证研究
机构 * Harbin Institute of Technology(哈尔滨工业大学) ; Beijing Institute of Technology(北京理工大学) ; Beihang University(北航) ; Peking University(北京大学)
AI总结 研究从安卓到开源鸿蒙系统的图形用户界面测试迁移问题,构建数据集,选择并适配两种先进迁移方法进行评估,发现现有方法效果不佳,进而提出增强方法ITeM-HM,显著提升了测试迁移成功率。
SANTS:面向世界动作模型的状态自适应调度器
机构 * Fudan University(复旦大学) ; Harbin Institute of Technology(哈尔滨工业大学) ; Deep Computing Era Technology Co., Ltd(深计算时代科技有限公司)
AI总结 提出状态自适应噪声轨迹调度器(SANTS),通过根据视频状态动态选择去噪深度来优化视频到动作的扩散策略,在保持控制性能的同时大幅降低推理延迟。
Comments 17 pages, 5 figures, 8 tables. Project page: https://advanced-robotics-lab.github.io/SANTS/
CreatiParser: 从位图图形设计生成可编辑的图层
机构 * School of Information Science and Technology, University of Science and Technology of China(科学技术大学信息科学与技术学院) ; ByteDance Intelligent Creation(字节跳动智能创作) ; School of Computer Science and Technology, Harbin Institute of Technology (Weihai)(哈尔滨工业大学(威海)计算机科学与技术学院) ; Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合国家科学中心人工智能研究院)
AI总结 本文提出CreatiParser框架,将位图图形设计分解为可编辑的文本、背景和贴纸图层,结合视觉语言模型和多分支扩散架构,提升生成质量与编辑灵活性,实验显示在Parser-40K和Crello数据集上性能优于现有方法。
M2I2HA:基于模态内和模态间超图注意力的多模态目标检测
机构 * Harbin Institute of Technology(哈尔滨工业大学)
AI总结 针对多模态目标检测中模态内和模态间信息提取及跨模态对齐的挑战,提出基于超图理论的M2I2HA网络,通过多个模块实现多模态特征的有效处理,在多模态目标检测任务中取得了最优性能。
Comments 43 pages, 13 figures, The theoretical derivation was refined, some data was updated, and experiments were added
TopCoW挑战——用于CT和MR血管造影的拓扑感知Willis环分割
机构 * Department of Quantitative Biomedicine, University of Zurich, Zurich, Switzerland ; Institute of Computational Life Sciences, Zurich University of Applied Sciences (ZHAW), Waedenswil, Switzerland ; Department of Neuroradiology, University Hospital of Zurich, Zurich, Switzerland ; Department of Neurosurgery, Zhongnan Hospital of Wuhan University, Wuhan, China ; Department of Radiology at Weill Cornell Medicine, Cornell University, New York, USA ; Institute for Tissue Engineering ; School of Computation, Information ; Technology, Technical University of Munich, Germany ; Athinoula A. Martinos Center for Biomedical Imaging, Harvard Medical School, Boston, USA ; School of Medicine ; Health, TUM Klinikum, Technical University of Munich, Germany ; Munich Center for Machine Learning, Munich, Germany ; Department of Computing, Imperial College London, London, UK ; Image Sciences Institute, UMC Utrecht, Utrecht, The Netherlands ; Department of Neurology ; Neurosurgery, University Medical Center Utrecht, Utrecht, The Netherlands ; Department of Radiology, University Medical Center Utrecht, Utrecht, The Netherlands ; Electronic \& Information Engineering School, Harbin Institute of Technology (Shenzhen), China ; Peng Cheng Laboratory, Shenzhen, China ; Division of Medical Image Computing, German Cancer Research Center (DKFZ), Heidelberg, Germany ; Faculty of Mathematics ; Computer Science, Heidelberg University, Germany ; Helmholtz Imaging, German Cancer Research Center, Heidelberg, Germany ; Data Science School for Health, Karlsruhe/Heidelberg, Germany ; Learning Group, Department of Radiation Oncology, Heidelberg University Hospital ; Department of Radiology, University of Washington, Seattle, WA, USA ; Department of Radiology, Ren Ji Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China ; Department of Clinical Neurosciences, Division of Neurosurgery, Geneva University Hospitals, Geneva, Switzerland ; Department of Neurology, University Hospital of Zurich, Zurich, Switzerland ; Department of Physiology, University of Toronto, Canada ; Department of Neurosurgery, University Hospital of Zurich, Zurich, Switzerland ; Department of Diagnostic Imaging, National University Hospital, Singapore ; University of Chicago, USA ; Department of Diagnostic ; Interventional Neuroradiology, University Hospital Berne ; University of Berne, Berne, Switzerland ; Centre de Recherche du Centre Hospitalier de l’Université de Montréal (CRCHUM), Montréal, Québec, Canada ; DEEPNOID Inc., Seoul, South Korea ; Department of Artificial Intelligence, Korea University, Seoul, South Korea ; Charité Lab for AI in Medicine (CLAIM), Charité Universitätsmedizin Berlin, Berlin, Germany ; Lung Institute, Faculty of Medicine, Imperial College London, London, UK ; Centre for Medical Image Computing, Department of Computer Science, University College London, London, UK ; Department of Radiation Oncology, Duke University Medical Center, Durham, NC, USA ; Institute of Medical Technology, Peking University Health Science Center, Beijing, China ; Hangzhou Genlight MedTech Co., Ltd., China ; Institute of Medical Robotics, Shanghai Jiao Tong University, Shanghai, China ; Department of Automation, Shanghai Jiao Tong University, Shanghai, China ; Department of Artificial Intelligence, Sungkyunkwan University, Seoul, South Korea ; Department of Electrical ; Computer Engineering, Sungkyunkwan University, Seoul, South Korea ; Shanghai MediWorks Precision Instruments Co., Ltd., China ; Institute of Informatics, HES-SO Valais-Wallis, Switzerland ; Department of Measurement ; Electronics, AGH University of Krakow, Poland ; Laboratoire de Thermique et Energie de Nantes (LTeN), Université Nantes, Polytech’Nantes, Nantes, France ; Research Institute of Computer Vision ; Center for Precision Health, McWilliams School of Biomedical Informatics, University of Texas Health Science Center at Houston, USA ; Physense, BCN-Medtech, Department of Communication ; Information Technologies, Universitat Pompeu Fabra, Barcelona, Spain ; Department of Mathematical Modeling ; Machine Learning, University of Zurich, Zurich, Switzerland ; Laboratory of Brain Atlas ; Brain-inspired Intelligence, Institute of Automation, Chinese Academy of Sciences, Beijing, China ; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China ; School of Computer ; Information Engineering, Xiamen University of Technology, Xiamen, China ; Vascular Research Center, University at Buffalo, NY, USA ; Department of Pathology ; Anatomical Sciences, University at Buffalo, NY, USA ; Department of Neurosurgery, University at Buffalo, NY, USA ; LPIXEL Inc., Tokyo, Japan
AI总结 组织TopCoW基准挑战,发布含125对MRA和CTA扫描的注释数据集,参与者提交CoW分割和变体分类算法,经评估,最佳算法在多任务中表现出色,证明CoW分割算法对下游临床应用有可解释性效用。
Comments Summary paper for the TopCoW Challenge: 4 figures, 1 table, and supplementary material in appendix. Accepted for publication in NEJM AI. Datasets and best-performing algorithm Dockers are available at https://zenodo.org/records/15692630 and https://zenodo.org/records/15665435
Journal ref NEJM AI 2026;3(8)
CLIP中用于密集开放词汇预测的稀疏注意力
机构 * King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学) ; Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
AI总结 研究在CLIP的最终视觉自注意力层用α-entmax变换替代逐行softmax,以解决其在密集开放词汇预测时注意力分散产生噪声的问题,在开放词汇任务评估中,注意力稀疏化增益与基线注意力偏离目标类程度成正比。
UniLM-Nav:零样本最后一英里导航的统一框架
机构 * Tsinghua University(清华大学) ; State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,字节跳动公司人工智能研究院) ; Harbin Institute of Technology(哈尔滨工业大学) ; Peking University(北京大学)
AI总结 研究移动操作中最后一英里导航问题,提出UniLM-Nav统一框架,通过多模态大语言模型后端分解任务为视图选择、功能接地和姿态推理,在OVMM基准上优于现有方法,还验证了在实际机器人上的适用性。
Comments Project page: https://unilm-nav.github.io
通过表示工程解锁大语言模型和大视觉语言模型的多语言推理能力
机构 * Harbin Institute of Technology(哈尔滨工业大学) ; Peng Cheng Laboratory(鹏城实验室) ; The University of Hong Kong(香港大学) ; Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳分校) ; Huawei Technologies Co., Ltd(华为技术有限公司)
AI总结 研究针对LLMs和LVLMs在多语言推理中英语表现优于低资源语言的问题,提出无训练的推理时方法MRRE,通过在特定层注入两个预计算向量增强多语言推理能力,实验证明该方法有效提升非英语推理及输入输出语言一致性。
Comments ACL2026 main