arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Huazhong University of Science and Technology(华中科技大学)

共收录 702
2603.09943 2026-05-26 cs.AI

PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMs

PathMem: 面向病理学多模态大模型的认知对齐记忆转换

Jinyue Li, Yuci Liang, Qiankun Li, Xinheng Lyu, Jiayu Qian, Huabao Chen, Kun Wang, Zhigang Zeng, Anil Anthony Bharath, Yang Liu

机构 * University of Science and Technology of China(中国科学技术大学) Shenzhen University(深圳大学) Nanyang Technological University(南洋理工大学) Imperial College London(伦敦帝国学院) Huazhong University of Science and Technology(华中科技大学)

AI总结 提出PathMem框架,通过长期记忆与工作记忆的动态转换机制,实现结构化病理知识整合与可解释记忆控制,在WSI报告生成和开放诊断任务上达到SOTA。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03827 2026-05-26 cs.CV cs.RO

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

LIBERO-PRO:超越记忆的视觉-语言-动作模型鲁棒与公平评估

Xueyang Zhou, Yangming Xu, Guiyao Tie, Yongchao Chen, Guowen Zhang, Duanfeng Chu, Pan Zhou, Lichao Sun

机构 * Huazhong University of Science and Technology(华中科技大学) College of AI, Tsinghua University(清华大学人工智能学院) Wuhan University of Technology(武汉理工大学) Lehigh University(莱斯大学)

AI总结 针对LIBERO基准评估中的记忆偏差问题,提出LIBERO-PRO扩展基准,通过在操作对象、初始状态、任务指令和环境四个维度施加合理扰动,揭示现有VLA模型性能从90%以上骤降至0.0%的严重缺陷,并呼吁采用鲁棒评估方法。

Comments 10 pages,7 figures, 0 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.07863 2026-05-26 cs.CV cs.AI cs.MM

You Can Ground Earlier than See: An Effective and Efficient Pipeline for Temporal Sentence Grounding in Compressed Videos

你可以比看见更早定位:一种用于压缩视频中时序句子定位的高效流程

Xiang Fang, Daizong Liu, Pan Zhou, Guoshun Nan

机构 * The Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology(大数据安全湖北工程研究中心,网络安全科学与工程学院,华中科技大学) Peking University(北京大学) Beijing University of Posts and Telecommunications(北京邮电大学)

AI总结 提出一种三分支压缩域时空融合框架(TCSF),直接从压缩视频中提取I帧、运动向量和残差特征,实现高效准确的时序句子定位。

Comments Accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.11572 2026-05-26 cs.CV cs.AI cs.IR cs.MM

Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval

多模态跨域对齐网络用于视频时刻检索

Xiang Fang, Daizong Liu, Pan Zhou, Yuchong Hu

机构 * Hubei Key Laboratory of Distributed System Security(湖北分布式系统安全重点实验室) Hubei Engineering Research Center on Big Data Security(湖北大数据安全工程研究中心) School of Cyber Science and Engineering(网络安全学院) Huazhong University of Science and Technology(华中科技大学) Wangxuan Institute of Computer Technology(王轩计算机技术研究所) Peking University(北京大学) School of Computer Science and Technology(计算机科学与技术学院) Key Laboratory of Information Storage System Ministry of Education of China(信息存储系统教育部重点实验室)

AI总结 提出多模态跨域对齐网络,通过域对齐、跨模态对齐和特定对齐三个模块,解决跨域视频时刻检索中域差异和语义鸿沟问题。

Comments Accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.14882 2026-05-26 cs.MM cs.CL cs.CV cs.IR

Hierarchical Local-Global Transformer for Temporal Sentence Grounding

层次化局部-全局Transformer用于时间语句定位

Xiang Fang, Daizong Liu, Pan Zhou, Zichuan Xu, Ruixuan Li

机构 * Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology(大数据安全湖北工程研究中心,华中科技大学网络安全科学与工程学院) Wangxuan Institute of Computer fTechnology, Peking University(王宣计算机技术研究院,北京大学) School of software, Dalian University of Technology(软件学院,大连理工大学) School of Computer Science, and Technology, Huazhong University of Science, and Technology(计算机科学与技术学院,华中科技大学)

AI总结 提出层次化局部-全局Transformer(HLGT),通过建模视频和查询的不同粒度层次及跨模态交互,实现更细粒度的多模态表示,并在三个数据集上取得最先进性能。

Comments Publish in IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.11194 2026-05-26 cs.LG cs.CV cs.NE

V3H: View Variation and View Heredity for Incomplete Multi-view Clustering

V3H: 面向不完整多视图聚类的视图变异与视图遗传

Xiang Fang, Yuchong Hu, Pan Zhou, Dapeng Oliver Wu

机构 * School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学大数据安全工程研究中心) Department of Electrical and Computer Engineering, University of Florida(佛罗里达大学电子与计算机工程系)

AI总结 提出一种受遗传学启发的视图变异与视图遗传方法(V3H),通过分解子空间为变异矩阵和遗传矩阵分别学习各视图的独特信息和所有视图的一致信息,并利用可调低秩表示恢复底层数据结构,在不完整多视图聚类中同时捕获一致与独特信息,在15个基准数据集上超越现有方法。

Comments Publisheded in IEEE Transactions on Artificial Intelligence

Journal ref IEEE Transactions on Artificial Intelligence 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.10396 2026-05-26 cs.LG cs.AI

Double Self-weighted Multi-view Clustering via Adaptive View Fusion

双自加权多视图聚类:通过自适应视图融合

Xiang Fang, Yuchong Hu

机构 * School of Computer Science and Technology, Key Laboratory of Information Storage System Ministry of Education of China, Huazhong University of Science and Technology(计算机科学与技术学院,信息存储系统教育部重点实验室,华中科技大学)

AI总结 提出双自加权多视图聚类框架(DSMC),通过自适应权重矩阵和权重因子分别对特征和图进行加权,去除冗余和噪声,并融合多图进行聚类。

Comments Corresponding author: Xiang Fang

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.10331 2026-05-26 cs.CV cs.LG

ANIMC: A Soft Framework for Auto-weighted Noisy and Incomplete Multi-view Clustering

ANIMC: 一种自动加权噪声与不完整多视图聚类的软框架

Xiang Fang, Yuchong Hu, Pan Zhou, Dapeng Oliver Wu

机构 * Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology(大数据安全湖北工程研究中心,信息科学与工程学院,华中科技大学) School of Computer Science and Technology, Huazhong University of Science and Technology(计算机科学与技术学院,华中科技大学) Key Laboratory of Information Storage System Ministry of Education of China, Huazhong University of Science and Technology(信息存储系统教育部重点实验室,华中科技大学) Department of Electrical and Computer Engineering, University of Florida(电气与计算机工程系,佛罗里达大学)

AI总结 提出ANIMC框架,通过软自动加权策略和双软正则回归模型,处理多视图聚类中的缺失实例和噪声问题。

Comments Publisheded in IEEE Transactions on Artificial Intelligence

Journal ref IEEE Transactions on Artificial Intelligence 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.10254 2026-05-26 cs.LG cs.AI stat.ML

Unbalanced Incomplete Multi-view Clustering via the Scheme of View Evolution: Weak Views are Meat; Strong Views do Eat

通过视图演化方案的不平衡不完整多视图聚类:弱视图为食,强视图为食

Xiang Fang, Yuchong Hu, Pan Zhou, Dapeng Oliver Wu

机构 * School of Computer Science and Technology, Key Laboratory of Information Storage System Ministry of Education of China, Huazhong University of Science and Technology(计算机科学与技术学院,信息存储系统教育部重点实验室,华中科技大学) Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology(大数据安全工程研究中心,网络安全学院,华中科技大学) Department of Electrical and Computer Engineering, University of Florida(电气与计算机工程系,佛罗里达大学)

AI总结 针对不同视图不完整程度不平衡的问题,受生物进化理论启发,提出基于视图演化的不平衡不完整多视图聚类方法UIMC,通过加权多视图子空间聚类和低秩鲁棒表示恢复数据,显著提升聚类性能。

Comments Accepted by IEEE Transactions on Emerging Topics in Computational Intelligence

Journal ref IEEE Transactions on Emerging Topics in Computational Intelligence 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15284 2026-05-26 cs.LG math.ST stat.TH

Small Ensemble-based Data Assimilation: A Machine Learning-Enhanced Data Assimilation Method with Limited Ensemble Size

基于小集合的数据同化:一种机器学习增强的有限集合数据同化方法

Zhilin Li, Zhou Yao, Xianglong Li, Zeng Liu, Zhaokuan Lu, Shanlin Xu, Seungnam Kim, Guangyao Wang

机构 * Centre for Regional Oceans, Department of Ocean Science and Technology, and State Key Laboratory of Internet of Things for Smart City, University of Macau(澳门地区海洋研究中心、海洋科学与技术学院及智能城市物联网国家重点实验室,澳门大学) School of Naval Architecture and Ocean Engineering, Huazhong University of Science and Technology(华中科技大学船舶与海洋工程学院) Ningbo Institute of Dalian University of Technology(大连理工大学宁波研究院) College of Civil Engineering, Zhejiang University of Technology(浙江工业大学土木工程学院) Department of Naval Architecture and Ocean Engineering, Hongik University(成均馆大学船舶与海洋工程学院) State Key Laboratory of Internet of Things for Smart City, University of Macau(智能城市物联网国家重点实验室,澳门大学) Zhuhai UM Science and Technology Research Institute(珠海UM科技研究院)

AI总结 提出一种结合集合卡尔曼滤波与全连接神经网络的机器学习数据同化方法,通过小集合生成初步分析状态并用神经网络预测修正项,在几乎不增加计算成本下提升精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05004 2026-05-26 cs.CL

Can Large Language Models Resolve Semantic Discrepancy in Self-Destructive Subcultures? Evidence from Jirai Kei

大语言模型能否解决自我毁灭亚文化中的语义差异?来自Jirai Kei的证据

Peng Wang, Xilin Tao, Siyi Yao, Jiageng Wu, Yuntao Zou, Zhuotao Tian, Libo Qin, Dagang Li

机构 * School of Computer Science and Engineering, Macau University of Science and Technology(澳门科技大学计算机科学与工程学院) SKLPlanets, Macau University of Science and Technology(澳门科技大学SKLPlanets) College of Software, Northeastern University(东北大学软件学院) School of Energy and Power Engineering, Huazhong University of Science and Technology(华中科技大学能源与动力工程学院) School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)计算机科学与技术学院)

AI总结 针对亚文化中自我毁灭行为检测面临的知识滞后和语义错位问题,提出多智能体框架SAS,通过自动检索和亚文化对齐显著提升LLM检测性能,并优于现有先进方法。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10054 2026-05-26 cs.LG cs.AI cs.CL cs.CV

Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs

Uni-DPO:大语言模型动态偏好优化的统一范式

Shangpin Peng, Weinong Wang, Zhuotao Tian, Senqiao Yang, Xing Wu, Haotian Xu, Chengquan Zhang, Takashi Isobe, Baotian Hu, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Xi’an Jiaotong University(西安交通大学) The Chinese University of Hong Kong(香港中文大学) University of Chinese Academy of Sciences(中国科学院大学) Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 针对现有DPO方法忽略数据质量和学习难度差异的问题,提出Uni-DPO统一框架,通过自适应重加权偏好对实现更有效的数据利用和更优性能。

Comments Accepted by ICLR 2026. Code & models: https://github.com/pspdada/Uni-DPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23204 2026-05-25 cs.AI

AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery

AutoResearch AI:迈向人工智能驱动的科研自动化以实现科学发现

Guiyao Tie, Jiawen Shi, Dingjie Song, Yixiao Huang, Ziji Sheng, Xueyang Zhou, Daizong Liu, Pan Zhou, Yongchao Chen, Ran Xu, Lifang He, Qingsong Wen, Manling Li, Cong Lu, Shuai Li, Pengtao Xie, Yixuan Yuan, Rui Meng, Lei Xing, Lichao Sun, Caiming Xiong, Philip S. Yu, Jianfeng Gao

机构 * Huazhong University of Science and Technology(华中科技大学) Lehigh University(莱斯大学) Tsinghua University(清华大学) Wuhan University(武汉大学) Salesforce Research(Salesforce研究) Squirrel AI Learning(Squirrel AI学习) Northwestern University(西北大学) Independent(独立) Shanghai Jiao Tong University(上海交通大学) University of California San Diego(加州大学圣地亚哥分校) Chinese University of Hong Kong(香港中文大学) University of Illinois Chicago(伊利诺伊大学香槟分校) Stanford University(斯坦福大学) Google Cloud AI Research(谷歌云AI研究) Recursive Superintelligence(递归超级智能) Microsoft Research(微软研究院)

AI总结 本文综述了AI驱动的科研工作流自动化(AutoResearch)的发展,分析了从任务级AI到工作流级研究自动化的转变,并提出了五个评估维度(新颖性、有效性、影响力、可靠性和溯源),指出自主性受领域条件限制。

Comments 49 pages, 12 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22506 2026-05-22 cs.CR cs.LG

EnCAgg: Enhanced Clustering Aggregation for Robust Federated Learning against Dynamic Model Poisoning

EnCAgg: 增强型聚类聚合用于对抗动态模型中毒的联邦学习

Tianyun Zhang, Zhen Yang, Haozhao Wang, Ru Zhang, Yongfeng Huang

机构 * School of Cyberspace Security, Beijing University of Posts and Telecommunications(信息安全部门,北京邮电大学) School of Computer Science and Technology, Huazhong University of Science and Technology(计算机科学与技术学院,华中科技大学) Department of Electronic Engineering, Tsinghua University(电子工程系,清华大学)

AI总结 本文提出了一种新的鲁棒聚合方法,通过利用少量已知的良性客户端作为参考,准确识别和过滤恶意梯度,同时保留尽可能多的良性梯度,即使恶意客户端的数量未知且变化。方法包括密度基低维梯度聚类、增强聚类低维梯度生成模型和低维梯度重新聚类。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20820 2026-05-21 cs.CV

AIR: Amortized Image Reconstruction Framework for Self-Supervised Feed-Forward 2D Gaussian Splatting

AIR: 一种用于自监督前馈2D高斯点散射的 amortized 图像重建框架

Zhaojie Zeng, Yuesong Wang, Yawei Luo, Tao Guan

机构 * School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) School of Software Technology, Zhejiang University(浙江大学软件学院)

AI总结 本文提出了一种自监督前馈框架AIR,通过将迭代高斯拟合 amortized 到单次网络传递中,消除了每张图像测试时的优化需求。该框架采用分阶段残差架构,逐步从重建残差中预测额外的高斯原始体,并结合显式的阶段控制机制,仅在欠重建区域激活新的原始体。通过预测-优化-蒸馏训练策略,稳定了多阶段预测,最终实现了更高效的图像重建。

Comments preprint version

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20273 2026-05-21 cs.LG cs.AI

Modality-Decoupled Online Recursive Editing

模态解耦的在线递归编辑

Siyuan Li, Youyuan Zhang, Fangming Liu, Jing Li

机构 * Harbin Institute of Technology, Shenzhen, China.(哈尔滨工业大学(深圳)) Peng Cheng Laboratory, China.(鹏城实验室) Huazhong University of Science and Technology, China(华中科技大学)

AI总结 本文提出M-ORE,一种用于持续多模态大语言模型适应的模态解耦在线递归编辑器,通过统一的近端投影公式和Sherman-Morrison递归实现常数级的每编辑开销,从而在保持模块局部统计信息和固定正交低秩编辑子空间的同时,减少长周期干扰,提升可靠性、通用性和局部性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19594 2026-05-20 cs.RO

MCNav: Memory-Aware Dynamic Cognitive Map for Zero-shot Goal-oriented Navigation

MCNav: 用于零样本目标导向导航的记忆感知动态认知图

Jingyu Li, Zhe Liu, Wenxiao Wu, Li Zhang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) University of Hong Kong(香港大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 本文提出MCNav,一种记忆感知的动态认知图导航框架,通过高效查询已探索区域的相关物体信息,解决零样本目标导向导航中目标丢失或误识别的问题,通过目标再验证和遗漏目标再探索策略,结合黑名单和双检机制,实现最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19390 2026-05-20 cs.CV

LMM-Track4D: Eliciting 4D Dynamic Reasoning in LMMs via Trajectory-Grounded Dialogue

LMM-Track4D: 通过轨迹引导的对话激发LMM中的4D动态推理

Chaoyue Li, Yongxue Xu, Jie Feng, Jiayu Ding

机构 * Huazhong University of Science and Technology(华中科技大学) Sun Yat-sen University(中山大学) Beihang University(北航) Peking University(北京大学)

AI总结 本文提出LMM-Track4D任务,通过轨迹引导的多轮时空对话,结合RTGE、TRK和OSK-RA解码器,提升LMM在4D动态推理中的性能,实验表明显式动态状态建模是有效设计原则。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25646 2026-05-20 cs.CV cs.RO

SAMe: A Semantic Anatomy Mapping Engine for Robotic Ultrasound

SAMe:一种用于机器人超声的语义解剖映射引擎

Jing Zhang, Duojie Chen, Wentao Jiang, Zihan Lou, Jianxin Liu, Xinwu Cui, Qinghong Zhao, Bo Du, Christoph F. Dietrich, Dacheng Tao

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Hubei Center for Applied Mathematics, Wuhan University(湖北应用数学中心,武汉大学) Department of Ultrasound, The Central Hospital of Wuhan(武汉市中心医院超声科) Department of Medical Ultrasound, Tongji Hospital, Tongji Medical College, Huazhong University of Science and Technology(同济医院,同济医学院,华中科技大学医学影像科) Department of Ultrasound in Medicine, Renmin Hospital of Wuhan University(武汉大学仁医医院医学超声科) University Hospital, Johann-Wolfgang-Goethe University Frankfurt am Main(法兰克福歌德大学医学院大学医院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

AI总结 该研究提出SAMe,一种语义解剖映射引擎,通过提供显式的解剖先验层,解决机器人超声扫描初始化问题,实现了基于临床症状的解剖目标识别和控制指令生成,提高了自动扫描的准确性和效率。

Comments Supplementary information included. Code will be released at https://github.com/MiliLab/Echo-SAMe

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12296 2026-05-20 cs.LG cs.AI eess.SP

Synthetic Data Generation for Brain-Computer Interfaces: Overview, Benchmarking, and Future Directions

脑机接口中的合成数据生成:概述、基准测试与未来方向

Ziwei Wang, Zhentao He, Xingyi He, Hongbin Wang, Tianwang Jia, Jingwei Luo, Siyang Li, Xiaoqing Chen, Dongrui Wu

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan, China(华中科技大学人工智能与自动化学院,武汉,中国) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)

AI总结 本文综述了用于脑机接口的合成脑数据生成方法,讨论了不同生成方法的分类、基准实验、评估指标和应用,以及未来研究方向,旨在提升数据效率和隐私保护的脑机接口系统。

Comments 33 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05709 2026-05-20 cs.AI

Nonlinearity as Rank: Generative Low-Rank Adapter with Radial Basis Functions

非线性作为秩:基于径向基函数的生成低秩适配器

Yihao Ouyang, Shiwei Li, Haozhao Wang, Xiandi Luo, Zhuoqi Hu, Yuetong Song, Qiyu Qin, Yichen Li, Ruixuan Li

机构 * Huazhong University of Science and Technology(华中科技大学) Hebei University of Technology(河北工业大学)

AI总结 本文提出GenLoRA,通过使用轻量级非线性函数生成径向基函数来替代传统低秩适配器中显式的基向量存储,从而提高参数效率和细调性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00404 2026-05-20 cs.CV

Hard-Label Black-Box Attacks on 3D Point Clouds

针对3D点云的硬标签黑盒攻击

Daizong Liu, Yunbo Tao, Junhao Dong, Keke Tang, Pan Zhou, Wei Hu, Yew-Soon Ong

机构 * Institute for Math & AI(数学与人工智能研究院) Wuhan University(武汉大学) Huazhong University of Science and Technology(华中科技大学) Shenzhen Huazhong University of Science and Technology Research Institute(深圳华中科技大学研究机构) College of Computing and Data Science(计算与数据科学学院) Nanyang Technological University(南洋理工大学) Cyberspace Institute of Advanced Technology(先进技术网络空间研究院) Guangzhou University(广州大学) Wangxuan Institute of Computer Technology(王轩计算机技术研究所) Peking University(北京大学)

AI总结 本文提出了一种基于硬标签黑盒攻击的3D点云攻击方法,通过引入新的频谱感知决策边界算法生成高质量对抗样本,以提升攻击性能和对抗质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18733 2026-05-19 cs.CV

Advancing Narrative Long Video Generation via Training-Free Identity-Aware Memory

通过无训练的身份感知记忆推进叙事长视频生成

Jinzhuo Liu, Jiangning Zhang, Wencan Jiang, Yabiao Wang, Dingkang Liang, Zhucun Xue, Ran Yi, Yong Liu

机构 * Zhejiang University(浙江大学) Tencent Youtu Lab(腾讯云智实验室) Huazhong University of Science and Technology(华中科技大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 本文提出了一种无训练的身份感知记忆框架IAMFlow,通过显式建模和跟踪持久实体身份,实现一致的生成,同时引入NarraStream-Bench基准测试,在叙事流视频生成中取得最佳性能。

Comments Project page: https://eddie0521.github.io/projects/iamflow/ Code: https://github.com/Eddie0521/IAMFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17927 2026-05-19 cs.RO

Learning-Based Adaptive Control for Surgical Robotic Exposure Task on Deformable Tissues

基于学习的自适应控制用于变形组织手术机器人暴露任务

Jiayi Liu, Kaiqi Wei, Yiwei Wang, Huan Zhao, Han Ding

机构 * Huazhong University of Science and Technology(华中科技大学)

AI总结 本文提出了一种基于学习的自适应控制框架,用于解决手术中因覆盖组织的不规则几何形状、非线性生物力学特性及有限视野导致的自动组织牵开挑战,通过在线优化控制输入和深度变形估计模型实现零样本适应。

Comments Accepted to Robotics: Science and Systems (RSS) 2026. 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17806 2026-05-19 cs.LG

AMO: Adaptive Muon Orthogonalization

AMO:自适应缪子正交化

Xinlin Zhuang, Panyi Ouyang, Yichen Li, Jiangming Shi, Yizhang Chen, Shuman Liu, Ying Qian, Weiyang Liu, Haibo Zhang, Imran Razzak

机构 * The Chinese University of Hong Kong(香港中文大学) Shopee MBZUAI East China Normal University(华东师范大学) Huazhong University of Science and Technology(华中科技大学) Xiamen University(厦门大学)

AI总结 本文研究了缪子优化中正交化过程的异质性,提出自适应缪子正交化方法,通过测量权重几何特性动态分配NS预算,提升预训练性能。

Comments preprint, under-review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13977 2026-05-19 cs.CV

ROVR-Open-Dataset: A Large-Scale Depth Dataset for Autonomous Driving

ROVR-Open-Dataset: 一个大规模深度数据集用于自动驾驶

Xianda Guo, Ruijun Zhang, Yiqun Duan, Ruilin Wang, Matteo Poggi, Keyuan Zhou, Wenzhao Zheng, Wenke Huang, Gangwei Xu, Yanlun Peng, Yuan Si, Qin Zou

机构 * Wuhan University(武汉大学) Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所) University of Technology Sydney(悉尼大学) University of Bologna(博洛尼亚大学) Zhejiang University(浙江大学) University of California, Berkeley(加州大学伯克利分校) Huazhong University of Science and Technology(华中科技大学) Great Wall Motor(长城汽车) ROVR Labs, Inc.(ROVR实验室)

AI总结 本文提出ROVR-Open-Dataset,一个大规模、多样化且成本效益高的深度数据集,用于提升自动驾驶中空间感知的能力,通过提供丰富的场景、光照和天气条件数据,以及经过验证的地面真实数据,支持鲁棒的模型训练,并识别出当前架构共享的三种失败模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16909 2026-05-19 cs.AI

TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents

TOBench:面向真实世界工具使用代理的任务导向多模态基准

Zhiqiang Liu, Wenhui Dong, Yilang Tan, Yuwen Qu, Haochen Yin, Chenyang Si

机构 * Nanjing University(南京大学) Huazhong University of Science and Technology(华中科技大学) Southwest Jiaotong University(西南交通大学) The Chinese University of Hong Kong(香港中文大学)

AI总结 本文提出TOBench,一个面向真实世界工具使用代理的多模态基准,通过闭环多模态验证设计,评估和推动下一代多模态工具使用代理的发展。

Comments Github: https://github.com/Pi3AI/TOBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00952 2026-05-19 cs.CV

Decoupling Motion and Geometry in 4D Gaussian Splatting

分离运动与几何的4D高斯点散射

Yi Zhang, Yulei Kang, Jiangxin Sun, Beihao Xia, Jisheng Dang, Jian-Fang Hu

机构 * Sun Yat-sen University(中山大学) University of Trento(特伦特大学) Huazhong University of Science and Technology(华中科技大学) Lanzhou University(兰州大学)

AI总结 本文提出VeGaS框架,通过引入伽利略剪切矩阵和几何变形网络,分离高斯运动与几何属性,提升复杂非线性运动建模能力,实验表明其在公开数据集上达到最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16438 2026-05-19 cs.IR cs.AI

OPERA: A Reinforcement Learning--Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop Retrieval

OPERA: 一种增强强化学习的协调规划-执行架构用于面向推理的多跳检索

Yu Liu, Yanbing Liu, Fangfang Yuan, Cong Cao, Youbang Sun, Kun Peng, Weizhuo Chen, Jianjun Li, Zhiyuan Ma

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) Department of Electronic Engineering, Tsinghua University(清华大学电子工程系)

AI总结 OPERA通过协调规划-执行架构解决多跳检索中推理规划、检索和过滤的不足,采用MAPGRPO方法提升复杂任务性能。

Comments Accepted by AAAI 2026. Extended version

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15640 2026-05-18 cs.CV

Learning Disentangled Representations for Generalized Multi-view Clustering

学习解耦表示以实现通用多视图聚类

Xin Zou, Ruimeng Liu, Chang Tang, Zhenglai Li, Xinwang Liu, Kunlun He, Wanqing Li

机构 * AI Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能方向) School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) School of Computer, National University of Defense Technology(国防科技大学计算机学院) Medical Big Data Research Center, Medical Engineering Laboratory of Chinese PLA General Hospital(中国人民解放军总医院医学大数据研究中心,医学工程实验室) School of Computing and Information Technology, University of Wollongong(沃林根大学计算与信息学院)

AI总结 本文提出GMAE框架,通过解耦表示学习保留多视图互补性,提升聚类效果。实验表明其在完整和不完整多视图聚类任务中均优于现有方法。

Comments accepted by IEEE TPAMI 2026 (IEEE Transactions on Pattern Analysis and Machine Intelligence)

详情

展开后加载摘要…

URL PDF HTML 收藏