arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Harbin Institute of Technology(哈尔滨工业大学)

共收录 1439
2607.11245 2026-07-15 cs.SE cs.AI 版本更新

An Empirical Study for Android-to-OpenHarmony GUI Test Migration

从安卓到开源鸿蒙系统的图形用户界面测试迁移实证研究

Yakun Zhang, Xinjia Chen, Yiyun Chen, Yuxia Zhang, Mingyi Zhou, Xiang Gao, Shaokun Zhang, Li Li, Yunming Ye

机构 * Harbin Institute of Technology(哈尔滨工业大学) Beijing Institute of Technology(北京理工大学) Beihang University(北航) Peking University(北京大学)

AI总结 研究从安卓到开源鸿蒙系统的图形用户界面测试迁移问题,构建数据集,选择并适配两种先进迁移方法进行评估,发现现有方法效果不佳,进而提出增强方法ITeM-HM,显著提升了测试迁移成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27947 2026-07-15 cs.RO 版本更新

SANTS: A State-Adaptive Scheduler for World Action Models

SANTS:面向世界动作模型的状态自适应调度器

Yirui Sun, Guangyu Zhuge, Keliang Liu, Jie Gu, Xinyu Bing, Zhongxue Gan, Chunxu Tian

机构 * Fudan University(复旦大学) Harbin Institute of Technology(哈尔滨工业大学) Deep Computing Era Technology Co., Ltd(深计算时代科技有限公司)

AI总结 提出状态自适应噪声轨迹调度器(SANTS),通过根据视频状态动态选择去噪深度来优化视频到动作的扩散策略,在保持控制性能的同时大幅降低推理延迟。

Comments 17 pages, 5 figures, 8 tables. Project page: https://advanced-robotics-lab.github.io/SANTS/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19632 2026-07-15 cs.CV 版本更新

CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers

CreatiParser: 从位图图形设计生成可编辑的图层

Weidong Chen, Dexiang Hong, Zhendong Mao, Yutao Cheng, Xinyan Liu, Lei Zhang, Yongdong Zhang

机构 * School of Information Science and Technology, University of Science and Technology of China(科学技术大学信息科学与技术学院) ByteDance Intelligent Creation(字节跳动智能创作) School of Computer Science and Technology, Harbin Institute of Technology (Weihai)(哈尔滨工业大学(威海)计算机科学与技术学院) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合国家科学中心人工智能研究院)

AI总结 本文提出CreatiParser框架,将位图图形设计分解为可编辑的文本、背景和贴纸图层,结合视觉语言模型和多分支扩散架构,提升生成质量与编辑灵活性,实验显示在Parser-40K和Crello数据集上性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14776 2026-07-15 cs.CV 版本更新

M2I2HA: Multi-modal Object Detection Based on Intra- and Inter-Modal Hypergraph Attention

M2I2HA:基于模态内和模态间超图注意力的多模态目标检测

Xiaofan Yang, Yubin Liu, Wei Pan, Guoqing Chu, Junming Zhang, Jie Zhao, Zhuoqi Man, Xuanming Cao

机构 * Harbin Institute of Technology(哈尔滨工业大学)

AI总结 针对多模态目标检测中模态内和模态间信息提取及跨模态对齐的挑战,提出基于超图理论的M2I2HA网络,通过多个模块实现多模态特征的有效处理,在多模态目标检测任务中取得了最优性能。

Comments 43 pages, 13 figures, The theoretical derivation was refined, some data was updated, and experiments were added

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17670 2026-07-15 cs.CV cs.LG q-bio.QM q-bio.TO 版本更新

The TopCoW Challenge -- Topology-Aware Circle of Willis Segmentation for CT and MR Angiography

TopCoW挑战——用于CT和MR血管造影的拓扑感知Willis环分割

Kaiyuan Yang, Fabio Musio, Yihui Ma, Norman Juchler, Johannes C. Paetzold, Rami Al-Maskari, Luciano Höher, Hongwei Bran Li, Ibrahim Ethem Hamamci, Anjany Sekuboyina, Suprosanna Shit, Houjing Huang, Chinmay Prabhakar, Ezequiel de la Rosa, Bastian Wittmann, Diana Waldmannstetter, Florian Kofler, Fernando Navarro, Martin J. Menten, Ivan Ezhov, Daniel Rueckert, Iris N. Vos, Ynte M. Ruigrok, Birgitta K. Velthuis, Hugo J. Kuijf, Pengcheng Shi, Wei Liu, Ting Ma, Maximilian R. Rokuss, Yannick Kirchhoff, Fabian Isensee, Klaus Maier-Hein, Chengcheng Zhu, Huilin Zhao, Philippe Bijlenga, Julien Hämmerli, Catherine Wurster, Laura Westphal, Jeroen Bisschop, Elisa Colombo, Hakim Baazaoui, Hannah-Lea Handelsmann, Andrew Makmur, James Hallinan, Amrish Soundararajan, Benedikt Wiestler, Jan S. Kirschke, Evamaria O. Riedel, Roland Wiest, Emmanuel Montagnon, Laurent Letourneau-Guillon, Kwanseok Oh, Dahye Lee, Orhun Utku Aydin, Adam Hilbert, Jana Rieger, Dimitrios Rallios, Satoru Tanioka, Alexander Koch, Dietmar Frey, Abdul Qayyum, Moona Mazher, Steven Niederer, Nico Disch, Julius C. Holzschuh, Dominic LaBella, Francesco Galati, Daniele Falcetta, Maria A. Zuluaga, Chaolong Lin, Haoran Zhao, Zehan Zhang, Minghui Zhang, Xin You, Hanxiao Zhang, Guang-Zhong Yang, Yun Gu, Sinyoung Ra, Jongyun Hwang, Hyunjin Park, Junqiang Chen, Marek Wodzinski, Henning Müller, Nesrin Mansouri, Florent Autrusseau, Cansu Yalcin, Rachika E. Hamadache, Clara Lisazo, Joaquim Salvi, Adrià Casamitjana, Xavier Lladó, Uma Maria Lal-Trehan Estrada, Valeriia Abramova, Luca Giancardo, Arnau Oliver, Paula Casademunt, Adrian Galdran, Matteo Delucchi, Oscar Camara, Jialu Liu, Haibin Huang, Yue Cui, Zehang Lin, Yusheng Liu, Shunzhi Zhu, Tatsat R. Patel, Adnan H. Siddiqui, Vincent M. Tutino, Maysam Orouskhani, Huayu Wang, Mahmud Mossa-Basha, Yuki Sato, Sven Hirsch, Susanne Wegener, Bjoern Menze

机构 * Department of Quantitative Biomedicine, University of Zurich, Zurich, Switzerland Institute of Computational Life Sciences, Zurich University of Applied Sciences (ZHAW), Waedenswil, Switzerland Department of Neuroradiology, University Hospital of Zurich, Zurich, Switzerland Department of Neurosurgery, Zhongnan Hospital of Wuhan University, Wuhan, China Department of Radiology at Weill Cornell Medicine, Cornell University, New York, USA Institute for Tissue Engineering School of Computation, Information Technology, Technical University of Munich, Germany Athinoula A. Martinos Center for Biomedical Imaging, Harvard Medical School, Boston, USA School of Medicine Health, TUM Klinikum, Technical University of Munich, Germany Munich Center for Machine Learning, Munich, Germany Department of Computing, Imperial College London, London, UK Image Sciences Institute, UMC Utrecht, Utrecht, The Netherlands Department of Neurology Neurosurgery, University Medical Center Utrecht, Utrecht, The Netherlands Department of Radiology, University Medical Center Utrecht, Utrecht, The Netherlands Electronic \& Information Engineering School, Harbin Institute of Technology (Shenzhen), China Peng Cheng Laboratory, Shenzhen, China Division of Medical Image Computing, German Cancer Research Center (DKFZ), Heidelberg, Germany Faculty of Mathematics Computer Science, Heidelberg University, Germany Helmholtz Imaging, German Cancer Research Center, Heidelberg, Germany Data Science School for Health, Karlsruhe/Heidelberg, Germany Learning Group, Department of Radiation Oncology, Heidelberg University Hospital Department of Radiology, University of Washington, Seattle, WA, USA Department of Radiology, Ren Ji Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China Department of Clinical Neurosciences, Division of Neurosurgery, Geneva University Hospitals, Geneva, Switzerland Department of Neurology, University Hospital of Zurich, Zurich, Switzerland Department of Physiology, University of Toronto, Canada Department of Neurosurgery, University Hospital of Zurich, Zurich, Switzerland Department of Diagnostic Imaging, National University Hospital, Singapore University of Chicago, USA Department of Diagnostic Interventional Neuroradiology, University Hospital Berne University of Berne, Berne, Switzerland Centre de Recherche du Centre Hospitalier de l’Université de Montréal (CRCHUM), Montréal, Québec, Canada DEEPNOID Inc., Seoul, South Korea Department of Artificial Intelligence, Korea University, Seoul, South Korea Charité Lab for AI in Medicine (CLAIM), Charité Universitätsmedizin Berlin, Berlin, Germany Lung Institute, Faculty of Medicine, Imperial College London, London, UK Centre for Medical Image Computing, Department of Computer Science, University College London, London, UK Department of Radiation Oncology, Duke University Medical Center, Durham, NC, USA Institute of Medical Technology, Peking University Health Science Center, Beijing, China Hangzhou Genlight MedTech Co., Ltd., China Institute of Medical Robotics, Shanghai Jiao Tong University, Shanghai, China Department of Automation, Shanghai Jiao Tong University, Shanghai, China Department of Artificial Intelligence, Sungkyunkwan University, Seoul, South Korea Department of Electrical Computer Engineering, Sungkyunkwan University, Seoul, South Korea Shanghai MediWorks Precision Instruments Co., Ltd., China Institute of Informatics, HES-SO Valais-Wallis, Switzerland Department of Measurement Electronics, AGH University of Krakow, Poland Laboratoire de Thermique et Energie de Nantes (LTeN), Université Nantes, Polytech’Nantes, Nantes, France Research Institute of Computer Vision Center for Precision Health, McWilliams School of Biomedical Informatics, University of Texas Health Science Center at Houston, USA Physense, BCN-Medtech, Department of Communication Information Technologies, Universitat Pompeu Fabra, Barcelona, Spain Department of Mathematical Modeling Machine Learning, University of Zurich, Zurich, Switzerland Laboratory of Brain Atlas Brain-inspired Intelligence, Institute of Automation, Chinese Academy of Sciences, Beijing, China School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China School of Computer Information Engineering, Xiamen University of Technology, Xiamen, China Vascular Research Center, University at Buffalo, NY, USA Department of Pathology Anatomical Sciences, University at Buffalo, NY, USA Department of Neurosurgery, University at Buffalo, NY, USA LPIXEL Inc., Tokyo, Japan

AI总结 组织TopCoW基准挑战,发布含125对MRA和CTA扫描的注释数据集,参与者提交CoW分割和变体分类算法,经评估,最佳算法在多任务中表现出色,证明CoW分割算法对下游临床应用有可解释性效用。

Comments Summary paper for the TopCoW Challenge: 4 figures, 1 table, and supplementary material in appendix. Accepted for publication in NEJM AI. Datasets and best-performing algorithm Dockers are available at https://zenodo.org/records/15692630 and https://zenodo.org/records/15665435

Journal ref NEJM AI 2026;3(8)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11739 2026-07-14 cs.RO 新提交

AutoPath: Learning Transferable Goal-Conditioned Stochastic Path Prior for Safe Navigation Without Human Demonstrations

AutoPath:学习可转移的目标条件随机路径先验以实现无人类示范的安全导航

Ziyang Zhang, Boyang Zhou, Zesong Yang, Haocheng Peng, Zeming Gai, Xiao Liang, Yujun Shen, Danping Zou, Ruizhen Hu, Hujun Bao, Zhaopeng Cui

机构 * Zhejiang University(浙江大学) Harbin Institute of Technology(哈尔滨工业大学) Ant Group(蚂蚁集团) Shanghai Jiao Tong University(上海交通大学) Shenzhen University(深圳大学)

AI总结 研究在复杂环境下的安全导航问题,提出学习可转移目标条件随机路径先验的方法,引入规范状态表示和结构化先验学习框架,实验证明该方法成功率高、效率有竞争力且可跨平台转移。

Comments Accepted by IEEE Robotics and Automation Letters (RA-L). 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11646 2026-07-14 cs.CV cs.RO 新提交

Event-RGB Adaptive Tracking for Nighttime Highway Perception

用于夜间高速公路感知的事件-RGB自适应跟踪

Haidong Wang, Hengxing Cai, Wanlei Li, Xiaogang Xiong, Renxin Zhong

机构 * School of Intelligent Systems Engineering, Sun Yat-sen University(中山大学智能系统工程学院) School of Intelligence Science and Engineering, Harbin Institute of Technology(哈尔滨工业大学智能科学与工程学院)

AI总结 针对高速公路夜间低光条件下RGB感知性能差的问题,提出JEAT框架,通过自适应扩展卡尔曼滤波器动态融合事件流与RGB帧,还创建SEHN数据集,为多模态融合研究提供支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10795 2026-07-14 cs.AI cs.CL 新提交

STEC: Evidence Compression for Deep Search in Open-domain Multi-Hop QA

STEC:开放域多跳问答中深度搜索的证据压缩

Xinkang Li, Rong Jiang, Xin Song, Ye Wang, Yue Han, Changjian Li

机构 * National University of Defense Technology(国防科技大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

AI总结 研究开放域多跳问答中最终答案选择难题,提出STEC证据压缩框架从候选集选答案。通过答案级证据压缩和证据引导的答案验证机制,将选择从原始轨迹比较转为候选级证据比较,实验显示其性能最佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10308 2026-07-14 cs.CV 新提交

Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis

通过虚拟模态合成将大型多模态模型推广到通用视觉模态

Shihao Yuan, Yuanze Li, Ruyi Zhang, Ming Liu, Wangmeng Zuo

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算学部)

AI总结 研究大型多模态模型推广到未见视觉模态的挑战,提出VVM-Tuning训练框架,通过模态合成和上下文让模型具备相关能力,引入VVM-Bench基准,实验证明经合成模态训练的模型在多模态上有改进且无需模态内训练。

Comments Accepted by the European Conference on Computer Vision (ECCV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10165 2026-07-14 cs.CV cs.AI 新提交

EmoStyle: Affective Conditioning of Style-Specialist Experts for Emotional Image Generation

EmoStyle:用于情感图像生成的风格专家情感调节

Dexiang Hong, Yijie Guo, Weidong Chen, Xinyan Liu, Zixuan Zou, Zhendong Mao, Yongdong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Harbin Institute of Technology(哈尔滨工业大学)

AI总结 针对情感感知艺术图像生成中训练与测试时属性不一致造成的控制差距问题,提出EmoStyle框架,利用语言模型推理器预测情感线索,编码情感字段指导生成,训练专用LoRA适配器结合风格表达情感,经视觉语言模型引导排序,在挑战赛中获佳绩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10098 2026-07-14 cs.CV cs.LG 新提交

DynaFilter: Cloud-driven Dynamic Filtering for Satellite Edge Intelligence

DynaFilter:用于卫星边缘智能的云驱动动态过滤

Ziyang Zhang, Jie Liu, Luca Mottola

机构 * Politecnico di Milano(米兰理工大学) Harbin Institute of Technology Shenzhen(哈尔滨工业大学(深圳))

AI总结 针对卫星边缘系统带宽受限等问题,设计DynaFilter动态过滤技术,通过建立云查询语义与压缩域特征的映射,让边缘设备在压缩域直接进行选择性RoI推理,实现减少数据量、节省带宽、降低能耗及加快推理延迟等效果。

Comments 15 pages, 23 figures, accepted by ACM MobiCom 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07135 2026-07-14 cs.CV 版本更新

Sparse Attention for Dense Open-Vocabulary Prediction in CLIP

CLIP中用于密集开放词汇预测的稀疏注意力

Fatimah Zohra, Chen Zhao, Shuming Liu, Bernard Ghanem

机构 * King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

AI总结 研究在CLIP的最终视觉自注意力层用α-entmax变换替代逐行softmax,以解决其在密集开放词汇预测时注意力分散产生噪声的问题,在开放词汇任务评估中,注意力稀疏化增益与基线注意力偏离目标类程度成正比。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06537 2026-07-14 cs.RO 版本更新

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation

UniLM-Nav:零样本最后一英里导航的统一框架

Zhuofan Zhang, Tianxu Wang, Guoxi Zhang, Yixiong Lin, Xilin Wang, Hongming Xu, Qing Li, Song-Chun Zhu, Lifeng Fan

机构 * Tsinghua University(清华大学) State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,字节跳动公司人工智能研究院) Harbin Institute of Technology(哈尔滨工业大学) Peking University(北京大学)

AI总结 研究移动操作中最后一英里导航问题,提出UniLM-Nav统一框架,通过多模态大语言模型后端分解任务为视图选择、功能接地和姿态推理,在OVMM基准上优于现有方法,还验证了在实际机器人上的适用性。

Comments Project page: https://unilm-nav.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17512 2026-07-14 cs.LG

One-Shot Federated Clustering of Non-Independent Completely Distributed Data

非独立完全分布式数据的一键式联邦聚类

Yiqun Zhang, Shenghong Cai, Zihua Yang, Sen Feng, Yuzhu Ji, Haijun Zhang

机构 * School of Computer Science and Technology, Guangdong University of Technology(广东工业大学计算机科学与技术学院) Department of Computer Science, Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist大学计算机科学系) Department of Computer Science, Harbin Institute of Technology(哈尔滨工业大学计算机科学系)

AI总结 本文提出GOLD框架,解决非独立完全分布数据下的联邦聚类问题,通过全局导向局部分布学习提升聚类性能。

Comments This work has been accepted for publication in IEEE Internet of Things Journal

Journal ref IEEE Internet of Things Journal, vol. 13, no. 7, pp. 14964-14978, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23231 2026-07-14 cs.CV 版本更新

Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering

通过表示工程解锁大语言模型和大视觉语言模型的多语言推理能力

Qiming Li, Xiaocheng Feng, Yixuan Ma, Zekai Ye, Ruihan Chen, Xiachong Feng, Bing Qin

机构 * Harbin Institute of Technology(哈尔滨工业大学) Peng Cheng Laboratory(鹏城实验室) The University of Hong Kong(香港大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳分校) Huawei Technologies Co., Ltd(华为技术有限公司)

AI总结 研究针对LLMs和LVLMs在多语言推理中英语表现优于低资源语言的问题,提出无训练的推理时方法MRRE,通过在特定层注入两个预计算向量增强多语言推理能力,实验证明该方法有效提升非英语推理及输入输出语言一致性。

Comments ACL2026 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13184 2026-07-14 cs.CV

Triad: Empowering LMM-based Anomaly Detection with Vision Expert-guided Visual Tokenizer and Manufacturing Process

Triad: 通过视觉专家引导的视觉标记器和制造过程增强基于LMM的异常检测

Yuanze Li, Shihao Yuan, Haolin Wang, Qizhang Li, Ming Liu, Chen Xu, Guangming Shi, Wangmeng Zuo

机构 * Harbin Institute of Technology(哈尔滨工业大学) Pengcheng Lab(鹏城实验室)

AI总结 Triad通过引入视觉专家引导的标记器和制造过程,提升基于LMM的工业异常检测性能。

Journal ref In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 21917-21926. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09481 2026-07-13 cs.CV cs.AI 新提交

Decoupling Language Guidance from Backbones for Text-Guided Medical Segmentation

用于文本引导医学分割的骨干网络与语言引导解耦

Yungeng Liu, Xuanzi Fang, Haijin Zeng, Qi Dai, Yongyong Chen

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) NingBo No.2 Hospital(宁波第二医院)

AI总结 研究针对文本引导医学分割中模型组件紧密耦合问题,提出可转移骨干层次适配器框架BTHA,通过稳定特征级接口、分层监督策略和自适应门控语义引导适配器,有效提升分割效果且计算开销适度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09084 2026-07-13 cs.LG cs.CY 新提交

A Survey on the Green Development of Large Models: From Resource-Efficient Architectures to Hardware-Software Co-Design

大模型绿色发展综述:从资源高效架构到软硬件协同设计

Linhui Xiao, Guiping Cao, Mingyue Guo, Xianchao Guan, Fan Yang, Ming Tao, Xin Li, Yuxin Peng, Yaowei Wang

机构 * Pengcheng Laboratory(鹏城实验室) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Peking University(北京大学)

AI总结 综述大模型绿色发展,涵盖资源高效架构与软硬件协同设计,回顾高效模型构建、训练部署策略、节能硬件等进展,探讨其在关键领域应用,讨论挑战与方向,为大模型可持续发展提供路线图。

Comments This paper has been accepted by CJE (2026), paper homepage: https://cje.ejournal.org.cn/article/doi/10.23919/cje.2025.00.438

Journal ref Chinese Journal of Electronics, vol. 35, no. 5, pp. 1-24, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07156 2026-07-13 cs.LG 新提交

The Anatomy of Implicit Bias: Information Allocation in Neural Network Training

神经网络优化中的信息分配动态

Zhang Gongyue, Wang Zhiyong, Liu Donghan, Ren Weihong, Sheng Yixuan, Liu Honghai

机构 * State Key Laboratory of Robotics and Systems, Harbin Institute of Technology Shenzhen(机器人系统国家重点实验室(哈尔滨工业大学深圳))

AI总结 该论文从信息分配动态角度出发,将优化器隐含偏差解释为权重与偏差类参数路径间训练信号的相对分配,通过连续预处理指数\(p\)描述调整,分析其在训练中的形成机制,揭示此相对更新分配对参数轨迹和泛化行为的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18613 2026-07-13 cs.CL cs.AI 新提交

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

LLMs 是否已准备好辅助医生?PhysAssistBench:交互式医患-电子病历辅助基准

Tianming Du, Peijie Yu, Sihan Shang, Danli Shi, My Linh Nguyen, Shengbo Gao, Guangyuan Li, Yinghong Yu, Yan Jiang, Qianlong Zhao, Behzad Bozorgtabar, Shaoxiong Ji, Jiazhen Pan, Daniel Rueckert, Jiancheng Yang

机构 * Aalto University(阿尔托大学) Tencent(腾讯) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Hong Kong Polytechnic University(香港理工大学) Aarhus University(奥胡斯大学) Technical University of Munich(慕尼黑工业大学)

AI总结 提出PhysAssistBench基准,通过构建交互式患者代理评估LLM在医患-EHR交互中的协调能力,发现当前模型不可靠,瓶颈在于多维度协调而非单一能力。

Comments 34 pages with 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09506 2026-07-13 cs.CV

Pillar-Voxel Fusion Network for 3D Object Detection in Airborne Hyperspectral Point Clouds

用于空中超光谱点云中3D目标检测的柱体-体素融合网络

Yanze Jiang, Yanfeng Gu, Xian Li

机构 * School of Electronics and Information Engineering, Harbin Institute of Technology(电子与信息工程学院,哈尔滨工业大学)

AI总结 本文提出PiV-AHPC网络,通过柱体-体素双分支编码器和多级特征融合机制,有效解决空中超光谱点云中3D目标检测的几何-光谱失真问题,实现高精度检测与泛化能力。

Journal ref Sci China Inf Sci, 2026, 69(1): 112301

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08236 2026-07-10 cs.CV 新提交

TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading

TVTA:基于轨迹感知的视素引导的基于事件的唇读时间聚合

Jingrong Zheng, Hongwei Ren, Xiangqian Wu

机构 * Harbin Institute of Technology(哈尔滨工业大学)

AI总结 针对基于事件唇读方法的局限,提出时间增强框架,引入轨迹感知差分聚合、视素引导聚合及指数移动平均师生训练策略,在DVS-Lip基准上验证有效性,消融研究验证各部分贡献,定性结果表明学到有意义视素感知结构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07287 2026-07-10 cs.RO 新提交

TouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation

TouchWorld:一种用于灵巧操作的预测性和反应性触觉基础模型

Jianyi Zhou, Feiyang Hong, Yunhao Li, Yicheng Zhao, Yongjue Cen, Zirui Liu, Jiakang Huang, Zirui Chen, Ruiyang Zhang, Weizhuo Zhu, Xuhua Song, Shuo Yang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) PHANES AI(PHANES人工智能公司)

AI总结 研究针对日常环境中灵巧操作需预测与反应的问题,提出TouchWorld模型。该模型采用分层策略,分离多项任务。通过将触摸用于预测和反馈,提升局部接触适应性。在六个相关任务中表现出色,超越最强基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13142 2026-07-10 cs.CV 版本更新

Real-World Blind Super-Resolution via Feature Matching with Implicit High-Resolution Priors

通过与隐式高分辨率先验进行特征匹配实现真实世界的盲超分辨率

Chaofeng Chen, Xinyu Shi, Yipeng Qin, Xiaoming Li, Xiaoguang Han, Tao Yang, Shihui Guo

机构 * School of Informatics, Xiamen University(厦门大学信息学院) School of Computer Science(计算机科学学院) Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) DAMO Academy, Alibaba Group(阿里巴巴集团达摩院) SSE, The Chinese University of Hong Kong(香港中文大学(深圳)SSE)

AI总结 针对现实世界图像超分辨率难题,提出FeMaSR方法,在紧凑特征空间通过匹配LR与HR图像特征及解码来恢复逼真HR图像,利用预训练HR先验、语义正则化及残差连接,实验显示其效果优于以往方法。

Comments Fix training details and some typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07314 2026-07-09 cs.LG cs.AI cs.CR 新提交

FedCVESA: Taking Away Training Data in Federated Learning via Correlation Value Encoding and Segmented Aggregation

FedCVESA:通过相关值编码和分段聚合在联邦学习中去除训练数据

Chongkai Li, Bang Zhang, Wenjian Luo

机构 * Institute of Cyberspace Security, School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)计算机科学与技术学院网络空间安全学院)

AI总结 研究联邦学习中白盒TATD攻击,提出FedCVESA方法,通过添加皮尔逊相关正则化器将私有训练数据编码到选定模型参数,并用分段聚合减少覆盖,实验证明该方法能窃取私有训练图像且保持主任务效用,揭示联邦学习存在的安全隐患。

Comments 16 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06442 2026-07-08 cs.RO 新提交

SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models

SIEVE:用于基于VLA模型的模仿学习的结构感知数据选择

Changti Wu, Bin Yu, Zhaolong Shen, Shijie Lian, Xiaopeng Lin, Cong Huang, Zhirui Zhang, Lei Zhang, Kai Chen

机构 * East China Normal University(东华师范大学) Zhongguancun Academy(中关村学院) Harbin Institute of Technology(哈尔滨工业大学) Beihang University(北航大学) Huazhong University of Science and Technology(华中科技大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院) DeepCybo Training Demonstrations Policy Execution(深科博培训演示政策执行)

AI总结 研究针对VLA模型模仿学习中数据冗余等问题,提出结构感知数据选择方法SIEVE。该方法将演示视为原语和接口组合,通过发现原语、分配预算和选择轨迹来选数据,实验证明其优于基线,能在少数据少步骤下超越全数据训练。

Comments The code is available at \href{https://github.com/ChangtiWu/SIEVE}{SIEVE}

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05787 2026-07-08 cs.CV 新提交

DeSeG: Decoupling Semantic Intent and Geometric Constraints for Physically Plausible Human-Scene Interaction

DeSeG:解耦语义意图和几何约束以实现物理上合理的人机场景交互

Jiakun Li, Zhe Li, Wenqiang Wu, Zheng Chang, Mingqi Gao, Jinyu Yang, Feng Zheng

机构 * Southern University of Science and Technology(南方科技大学) Spatialtemporal AI(时空人工智能) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) University of Sheffield(谢菲尔德大学)

AI总结 针对合成物理合理人机场景交互时语义 - 几何纠缠问题,提出DeSeG分层框架,通过残差语义规划器和物理正则化扩散执行器解耦语义意图与几何约束,实验表明该方法性能达最优,有效降低场景穿透率并提高语义对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22807 2026-07-08 cs.CL 新提交

KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking

KaLM-Reranker-V1:快速但非晚期交互的压缩文档重排序

Xinping Zhao, Jiaxin Xu, Ziqi Dai, Xin Zhang, Shouzheng Huang, Danyu Tang, Xinshuo Hu, Meishan Zhang, Baotian Hu, Min Zhang

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Shenzhen Loop Area Institute (SLAI)(深圳循环区域研究所)

AI总结 提出KaLM-Reranker-V1,一种基于编码器-解码器架构的快速但非晚期交互重排序器,通过解耦查询和段落计算并利用交叉注意力保持相关性建模,在BEIR等基准上达到最先进性能。

Comments Technical Report; Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04634 2026-07-07 cs.RO cs.CR cs.MA 新提交

Governed Caste Reassignment in Heterogeneous Swarms: An Asymmetric-Trust Protocol with Audited Operator Countersignature

异构群体中的受管种姓重新分配:一种带审核操作员副签名的不对称信任协议

Xue Qin, Simin Luan, Cong Yang, Zhijun Li

机构 * School of Software, Harbin Institute of Technology(哈尔滨工业大学软件学院) School of Future Science and Engineering, Soochow University(苏州大学未来科学与工程学院)

AI总结 研究异构机器人群体中种姓重新分配,提出不对称信任协议,自动收紧重分配可自动进行,受限放宽需操作员副签名,经评估能抵御多种攻击,还构建分布式审计层并证明相关特性。

Comments 28 pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04613 2026-07-07 cs.AI cs.CR 新提交

Governed Individuation: Cryptographically Decoupling an Agent's Learning from Its Authority

受治理的个体化:通过密码学方式将智能体的学习与其授权方解耦

Xue Qin, Simin Luan, Cong Yang, Zhijun Li

机构 * School of Software, Harbin Institute of Technology(哈尔滨工业大学软件学院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) School of Future Science and Engineering, Soochow University(苏州大学未来科学与工程学院)

AI总结 研究智能体在部署中学习时,运行系统是否仍受操作员授权限制的问题。提出受治理的个体化方法,通过绑定身份摘要和基于语义效果的动作门控保证限制,证明学习等不会扩大权限,实证验证了该方法效果。

Comments Technical companion report. 25 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏