arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Huazhong University of Science and Technology(华中科技大学)

共收录 701
2605.27367 2026-06-01 cs.CV

SpatialBench: Is Your Spatial Foundation Model an All-Round Player?

SpatialBench: 你的空间基础模型是全能选手吗?

Haosong Peng, Hao Li, Jiaqi Chen, Yuhao Pan, Runmao Yao, Yalun Dai, Fushuo Huo, Fangzhou Hong, Zhaoxi Chen, Haozhao Wang, Dingwen Zhang, Ziwei Liu, Wenchao Xu

机构 * Hong Kong University of Science and Technology(香港科技大学) Nanyang Technological University(南洋理工大学) Northwestern Polytechnical University(西北工业大学) Southeast University(东南大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 提出SpatialBench基准,通过跨范式、多域、确定性采样的评估,揭示当前空间基础模型在多样化下游任务中的泛化能力不足,并引入DA-Next-5M数据集和DA-Next模型推动空间表示学习。

Comments Project Page: https://ropedia.github.io/SpatialBench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15197 2026-06-01 cs.AI cs.CL cs.CV cs.RO

LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries

LangForce: 通过潜在动作查询对视觉语言动作模型进行贝叶斯分解

Shijie Lian, Bin Yu, Xiaopeng Lin, Laurence T. Yang, Zhaolong Shen, Changti Wu, Yuzhuo Miao, Cong Huang, Kai Chen

机构 * Huazhong University of Science and Technology(华中科技大学) Beijing Zhongguancun Academy(北京中关村学院) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院) Harbin Institute of Technology(哈尔滨工业大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Zhengzhou University(郑州大学) Beihang University(北航) East China Normal University(东华大学) DeepCybot Co., Ltd.(DeepCybot有限公司)

AI总结 针对VLA模型在训练中因数据偏差导致语言信息被忽略的问题,提出LangForce框架,通过贝叶斯分解和潜在动作查询构建双分支架构,最大化动作与指令的点互信息,无需新数据即可显著提升泛化能力。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25906 2026-06-01 cs.LG

Federated Learning with Enhanced Privacy via Model Splitting and Random Client Participation

通过模型拆分和随机客户端参与增强隐私的联邦学习

Yiwei Li, Shuai Wang, Zhuojun Tian, Xiuhua Wang, Shijian Su

机构 * School of Optoelectronic & Communication Engineering, Xiamen University of Technology(厦门理工学院光电信息与通信工程学院) National Key Laboratory of Wireless Communications, University of Electronic Science and Technology of China(电子科技大学信息与通信国家重点实验室) Division of Information Science and Engineering, KTH Royal Institute of Technology(皇家理工学院信息科学与工程系) School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学网络安全科学与工程学院) School of Engineering, Huaqiao University(华侨大学工程学院)

AI总结 提出MS-PAFL框架,通过将模型拆分为私有和公共子模型并仅向公共子模型注入噪声,结合随机客户端参与和本地数据子采样的隐私放大分析,在强隐私保证下实现更优的隐私-效用权衡。

Comments Accepted for publication in IEEE Transactions on Cognitive Communications and Networking

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30211 2026-05-29 cs.CV

Cycle Consistency in Video Object-Centric Learning

视频目标中心学习中的循环一致性

Rongzhen Zhao, Zhiyuan Li, Ruonan Wei, Juho Kannala, Joni Pajarinen

机构 * Department of Electrical Engineering and Automation, Aalto University, Espoo, Finland(艾洛大学电气工程与自动化系,芬兰 Espoo) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan, China(华中科技大学人工智能与自动化学院,中国 Wuhan) Department of Computer Science, Aalto University, Espoo, Finland(艾洛大学计算机科学系,芬兰 Espoo) Center for Machine Vision and Signal Analysis, University of Oulu, Oulu, Finland(奥卢大学机器视觉与信号分析中心,芬兰 Oulu)

AI总结 针对视频目标中心学习中潜在槽空间难以直接应用循环一致性的问题,提出隐式循环一致性(ICC),将约束从槽空间转移到连续重建流形,避免特征坍塌并提升性能。

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29812 2026-05-29 cs.CV

Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval Using Language

并非所有输入都有效:面向开放集视频时刻检索的语言方法

Xiang Fang, Wanlong Fang, Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou, Renfu Li, Zichuan Xu, Lixing Chen, Panpan Zheng, Yu Cheng

机构 * Huazhong University of Science and Technology(华中科技大学) Peking University(北京大学) Zhejiang Gongshang University(浙江工商大学) Dalian University of Technology(大连理工大学) Shanghai Jiao Tong University(上海交通大学) Xinjiang University(新疆大学) The Chinese University of Hong Kong(香港中文大学)

AI总结 针对开放集场景下视频时刻检索任务中无关查询导致错误检索的问题,提出基于归一化流的开放集视频时刻检索模型OpenVMR,实现分布内查询的精确检索与分布外查询的拒绝。

Comments Published in ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29793 2026-05-29 cs.CV

Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using Language

更少步骤,更优性能:基于语言的高效跨模态视频片段修剪用于视频时刻检索

Xiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou, Zichuan Xu, Wenzheng Xu, Junyang Chen, Renfu Li

机构 * Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science of Technology(湖北大数据安全工程研究中心,网络安全学院,华中科技大学) Peking University(北京大学) Henan University(河南大学) Dalian University of Technology(大连理工大学) Sichuan University(四川大学) Shenzhen University(深圳大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 提出SpotVMR方法,通过可学习的片段搜索模型和低成本语义索引特征,高效修剪查询相关视频片段,作为即插即用模块提升现有VMR方法的效率与性能。

Comments Published in AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29776 2026-05-29 cs.CV

Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning

通过打破尾部对齐改进CLIP适应:用于源无关跨域小样本学习

Shuai Yi, Yixiong Zou, Yuhua Li, Ruixuan Li

机构 * School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China(华中科技大学计算机科学与技术学院) Institute of Artificial Intelligence, Huazhong University of Science and Technology, Wuhan, China(华中科技大学人工智能研究院)

AI总结 针对CLIP在跨域小样本学习中的性能下降问题,提出自适应尾头对齐策略(ATHA),通过有选择地削弱低相似度图像令牌的对齐来减少过拟合,在四个基准上取得最优结果。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29708 2026-05-29 cs.CL

Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs

理解混合专家大语言模型中的安全敏感专家行为

Zhibo Zhang, Yuxi Li, Zhen Ouyang, Ling Shi, Kailong Wang

机构 * Huazhong University of Science and Technology, Wuhan, China(华中科技大学,武汉,中国) Nanyang Technological University, Singapore(南洋理工大学,新加坡)

AI总结 通过提出RASET框架,研究混合专家大语言模型中安全对齐与路由专家专业化之间的关系,发现路由模式主要由主题驱动,而安全行为可通过调整少数专家改变而不影响路由路径。

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29602 2026-05-29 cs.CV

CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning

CogniVerse: 用认知反思与几何推理革新多模态检索增强生成

Xiang Fang, Wanlong Fang, Changshuo Wang

机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) Nanyang Technological University, Singapore(新加坡南洋理工大学) University College London(伦敦大学学院)

AI总结 提出CogniVerse框架,通过认知反思模块、基于黎曼流形对齐的多模态检索模块和最优传输层次生成模块,解决多模态检索增强生成中的噪声检索、跨模态语义错位和生成不连贯问题。

Comments Accepted in CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19282 2026-05-29 cs.CL cs.AI

Less Is More: Elevating RAG via Performance-Driven Context Compression

少即是多:通过性能驱动的上下文压缩提升RAG

Ziqiang Cui, Yunpeng Weng, Xing Tang, Peiyang Liu, Shiwei Li, Bowei He, Jiamin Chen, Yansen Zhang, Xiuqiang He, Chen Ma

机构 * City University of Hong Kong, Hong Kong SAR, China(香港城市大学) Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE(阿布扎赫尔 Mohamed bin Zayed 人工智能大学) Huazhong University of Science and Technology(华中科技大学) Peking University, Beijing, China(北京大学) Shenzhen Technology University, Shenzhen, China(深圳技术大学)

AI总结 提出CORE-RAG框架,利用任务性能作为反馈信号迭代优化压缩策略,在3%压缩率下平均精确匹配得分提升3.3点。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12747 2026-05-29 cs.CV

Privacy Protection Against Personalized Text-to-Image Synthesis via Cross-image Consistency Constraints

针对个性化文本到图像合成的跨图像一致性约束隐私保护

Guanyu Wang, Kailong Wang, Yihao Huang, Mingyi Zhou, Geguang Pu, Li Li

机构 * Beihang University(北京航空航天大学) Huazhong University of Science and Technology(华中科技大学) East China Normal University(东华大学)

AI总结 提出跨图像反个性化框架,通过强制扰动图像间的风格一致性并采用动态比率调整策略,增强对扩散模型个性化攻击的抵抗能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28791 2026-05-28 cs.CL cs.AI

Skill-Conditioned Gated Self-Distillation for LLM Reasoning

技能条件门控自蒸馏用于大语言模型推理

Jiazhen Huang, Xiao Chen, Xiao Luo, Yong Dai, Senkang Hu, Yuzhi Zhao

机构 * Tsinghua University(清华大学) Fudan University(复旦大学) City University of Hong Kong(香港城市大学) Huazhong University of Science and Technology(华中科技大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

AI总结 提出技能条件门控自蒸馏(SGSD),通过从经验技能库中检索技能-错误对构建多教师池,并利用验证器验证教师极性,以鲁棒门控目标蒸馏信息性师生差异,在弱先验信息假设下提升数学推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27920 2026-05-28 cs.CV

Rethinking Video-Language Model from the Language Input Perspective

从语言输入角度重新思考视频-语言模型

Xiang Fang, Wanlong Fang, Changshuo Wang, Xiaoye Qu, Daizong Liu

机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) Nanyang Technological University, Singapore(新加坡南洋理工大学) University College London(伦敦大学学院) Huazhong University of Science and Technology(华中科技大学) Wuhan University(武汉大学)

AI总结 本文从语言输入角度出发,提出一种即插即用的框架,通过生成正负文本、属性文本推理和自加权损失,提升视频-语言模型的性能。

Comments Published in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27896 2026-05-28 cs.CL cs.CE

FinBoardBench: Benchmarking Dynamic Wealth Management and Strategic Financial Reasoning of LLMs via Board Game Simulations

FinBoardBench: 通过棋盘游戏模拟基准测试大语言模型的动态财富管理和战略金融推理

Xuesi Hu, Peng Wang, Jinpeng Miao, Xilin Tao, Caiwei Li, Yue Ma, Jie He, Qiancheng Zhang, Yuntao Zou, Dagang Li

机构 * School of Computer Science and Engineering, Macau University of Science and Technology, Macau, China(1 计算机科学与工程学院,澳门科技大学,澳门,中国) School of Economics, Anhui University, Anhui, China(2 经济学院,安徽大学,安徽,中国) SKLPlanets, Macau University of Science and Technology, Macau, China(3 SKLPlanets,澳门科技大学,澳门,中国) Department of Computer and Information Science, University of Macau, Macau, China(4 计算机与信息科学系,澳门大学,澳门,中国) School of Energy and Power Engineering, Huazhong University of Science and Technology, Hubei, China(5 能源与动力工程学院,华中科技大学,湖北,中国)

AI总结 提出基于三款经典金融棋盘游戏的评估套件FinBoardBench,测试大语言模型在动态财富管理、企业投资收购和竞争谈判等综合金融技能,发现模型虽具备基本规划能力但无法将静态推理转化为成功动态决策。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27894 2026-05-28 cs.CV

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

面向不完整多模态输入的统一视觉-语言模型

Xiang Fang, Wanlong Fang, Changshuo Wang, Keke Tang, Daizong Liu, Siyi Wang, Wei Ji

机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) Nanyang Technological University, Singapore(新加坡南洋理工大学) University College London(伦敦大学学院) Guangzhou University(广州大学) Wuhan University(武汉大学) Nanjing University(南京大学)

AI总结 针对视频-语言模型在传感器失效导致模态不完整数据下的训练-测试不一致问题,提出首个统一的不完整视频-语言模型作为即插即用模块,提升多模态任务性能。

Comments Published in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27374 2026-05-28 cs.CL

ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment

ICG: 通过基于MLLM的提示和个性化偏好对齐改进封面图像生成

Zhipeng Bian, Jieming Zhu, Qijiong Liu, Wang Lin, Guohao Cai, Zhaocheng Du, Jiacheng Sun, Zhou Zhao, Zhenhua Dong

机构 * Huazhong University of Science and Technology(华中科技大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Hong Kong Polytechnic University(香港理工大学) Zhejiang University(浙江大学)

AI总结 提出ICG框架,利用多模态大语言模型和扩散模型,通过元标记提取语义特征、用户嵌入个性化对齐及多奖励学习策略,实现高质量、个性化封面图像生成。

Comments Published in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 12268-12278, EMNLP 2025. Official version: https://doi.org/10.18653/v1/2025.emnlp-main.617

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (Main Track) EMNLP 2025 12268-12278

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16358 2026-05-28 cs.LG cs.CL

SaFeR-Steer: Evolving Multi-Turn MLLMs via Synthetic Bootstrapping and Feedback Dynamics

SaFeR-Steer:通过合成引导和反馈动力学进化多轮多模态大语言模型

Haolong Hu, Hanyu Li, Tiancheng He, Huahui Yi, An Zhang, Qiankun Li, Kun Wang, Yang Liu, Zhigang Zeng

机构 * Huazhong University of Science and Technology(华中科技大学) Beijing University of Posts and Telecommunications(北京邮电大学) West China Biomedical Big Data Center, Sichuan University(四川大学西部生物医学大数据中心) School of Public Policy and Administration, Chongqing University(重庆大学公共政策与管理学院) Nanyang Technological University(南洋理工大学)

AI总结 提出SaFeR-Steer框架,通过分阶段合成引导和导师参与的GRPO训练单学生模型,并引入轨迹一致总结奖励(TCSR)以解决多轮安全对齐中的长上下文安全衰减问题,显著提升多轮安全性和有用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.03808 2026-05-28 cs.LG eess.SP stat.ML

ECGadv: Generating Adversarial Electrocardiogram to Misguide Arrhythmia Classification System

ECGadv: 生成对抗性心电图以误导心律失常分类系统

Huangxun Chen, Chenyu Huang, Qianyi Huang, Qian Zhang, Wei Wang

机构 * The Hong Kong University of Science and Technology(香港科技大学) Southern University of Science and Technology, Peng Cheng Laboratory(南方科技大学鹏城实验室) Huazhong University of Science and Technology(华中科技大学)

AI总结 本文针对基于深度神经网络的心电图诊断系统,分析心电图特性并设计两种攻击模型下的对抗攻击方案,揭示系统盲点,呼吁采取对策。

Comments Accepted by AAAI 2020

Journal ref Proceedings of the AAAI conference on artificial intelligence 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27186 2026-05-27 cs.CL

MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation

MAIGO: 通过历史清理的在线策略自蒸馏缓解对话丢失

Haoyu Zheng, Yun Zhu, Shu Yuan, Shangming Chen, Qing Wang, Wenqiao Zhang, Jun Xiao, Yueting Zhuang

机构 * Zhejiang University(浙江大学) Shanghai AI Laboratory(上海人工智能实验室) Huazhong University of Science and Technology(华中科技大学) Fuzhou University(福州大学) Tencent(腾讯)

AI总结 针对大语言模型在多轮对话中性能下降(对话丢失)的问题,提出MAIGO方法,通过在线策略自蒸馏和清理历史助手回复来减少自污染,无需验证器或推理时辅助,显著提升多轮对话准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27083 2026-05-27 cs.CL cs.CR

On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning

反事实知识训练在LLM遗忘中的隐藏代价

Xiaotian Ye, Xiaohan Wang, Mengqi Zhang, Shu Wu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Huazhong University of Science and Technology(华中科技大学) Shandong University(山东大学) NLPR, MAIS, Institute of Automation, Chinese Academy of Sciences(神经网络计划、人工智能研究所、中国科学院自动化研究所)

AI总结 本文发现反事实微调(CFT)在LLM遗忘中存在知识冲突和幻觉溢出两大问题,并引入扩展基准RWKU+及诊断工具进行系统分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27003 2026-05-27 cs.CV cs.AI

Timestep-Aware SVDQuant-GPTQ for W4A4 Quantization of Wan2.2-I2V

时间步感知的 SVDQuant-GPTQ 用于 Wan2.2-I2V 的 W4A4 量化

Junhao Wu, Dezhong Yao, Hai Jin

机构 * National Engineering Research Center for Big Data Technology and System(大数据技术与系统国家工程研究中心) Services Computing Technology and System Lab(服务计算技术与系统实验室) Cluster and Grid Computing Lab(集群与网格计算实验室) School of Computer Science and Technology(计算机科学与技术学院) Huazhong University of Science and Technology(华中科技大学)

AI总结 针对 Wan2.2-I2V 视频扩散 Transformer 的 W4A4 量化,提出结合 SVDQuant 低秩异常补偿、GPTQ 重建感知残差权重量化和时间步分箱逐层激活裁剪比搜索的后训练量化框架,在 OpenS2V-Eval 上降低 59.3% 峰值显存且仅损失 0.9% VBench 平均分。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26637 2026-05-27 cs.RO

Enabling Extensible Embodied Capabilities with Tools

利用工具实现可扩展的具身能力

Xueyang Zhou, Zijia Wang, Qianjiang Li, Yibo Hu, Guiyao Tie, Li Wan, Yidan Liu, Pan Zhou, Lichao Sun, Yongchao Chen

机构 * Huazhong University of Science and Technology(华中科技大学) Hebei University of Technology(河北工业大学) Tianjin University(天津大学) Lehigh University(莱特大学) College of AI, Tsinghua University(清华大学人工智能学院)

AI总结 提出一种通过外部化能力为工具、并借助标准化协议ETP动态调用工具的方法,在仿真和真实平台上平均提升具身性能31%-36%,但揭示了工具使用在认知和感知方面增益显著而在执行方面有限。

Comments 51 pages, 20 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26501 2026-05-27 cs.CV cs.AI

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization

揭示视觉-语言模型的脆弱性:通过纹理约束扰动和跨模态优化的多模态对抗协同

Xiang Fang, Wanlong Fang, Changshuo Wang

机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) Nanyang Technological University, Singapore(新加坡南洋理工大学) University College London(伦敦大学学院)

AI总结 提出多模态对抗协同框架,通过纹理约束的通用对抗扰动和可学习的文本提示扰动,在黑盒设置下联合优化,揭示视觉-语言模型在多模态攻击下的脆弱性。

Comments Publish in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26441 2026-05-27 cs.CV cs.AI

Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective

从博弈视角重新思考弱监督视频时间定位

Xiang Fang, Zeyu Xiong, Wanlong Fang, Xiaoye Qu, Chen Chen, Jianfeng Dong, Keke Tang, Pan Zhou, Yu Cheng, Daizong Liu

机构 * Hubei Key Laboratory of Distributed System Security(湖北分布式系统安全重点实验室) Hubei Engineering Research Center on Big Data Security(大数据安全工程研究中心) School of Cyber Science and Engineering(网络安全科学与工程学院) Huazhong University of Science and Technology(华中科技大学) University of Central Florida(佛罗里达中央大学) Zhejiang Gongshang University(浙江工商大学) Guangzhou University(广州大学) The Chinese University of Hong Kong(香港中文大学) Peking University(北京大学)

AI总结 本文从博弈论视角出发,通过多元合作博弈建模帧与词的不确定对应关系,实现多级跨模态交互,从而在弱监督下提升视频时间定位的准确性。

Comments Published in ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09038 2026-05-27 cs.DB cs.AI

Scaling GraphLLM with Bilevel-Optimized Sparse Querying

基于双层优化稀疏查询的GraphLLM扩展

Yangzhe Peng, Haiquan Qiu, Quanming Yao, Kun He

机构 * Huazhong University of Science and Technology(华中科技大学) Tsinghua University(清华大学)

AI总结 提出BOSQ框架,通过自适应稀疏查询策略选择性调用LLM,在降低计算成本的同时保持或提升图节点任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25941 2026-05-26 cs.CV

Where Concept Erasure Should Occur: Concept-Layer Alignment in Text-to-Video Diffusion Models

概念擦除应发生在何处:文本到视频扩散模型中的概念-层对齐

Yiwei Xie, Ping Liu, Zheng Zhang

机构 * The School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China(人工智能与自动化学院,华中科技大学,武汉430074,中国) Department of Computer Science and Engineering, University of Nevada, Reno, NV, USA(计算机科学与工程系,内华达大学里诺分校,内华达州,美国)

AI总结 本文通过识别概念-层拓扑对齐瓶颈,提出基于可分离性优化的CLEAR框架,在文本到视频扩散模型中实现精确的概念擦除并保持生成质量。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25799 2026-05-26 cs.CV

Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning

应对源自由跨域小样本学习中加剧的注意力汇聚问题

Shuai Yi, Yixiong Zou, Yuhua Li, Ruixuan Li

机构 * Huazhong University of Science and Technology(华中科技大学)

AI总结 针对跨域小样本学习中标准微调加剧注意力汇聚导致判别性下降的问题,提出基于令牌动态重加权的方法抑制简单令牌依赖并增强困难令牌学习,实现新最优性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25621 2026-05-26 cs.CV

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering

StreamOV: 通过证据引导记忆与响应触发的流式全视频理解

Ming Xie, Zizheng Huang, Xudong Tan, Chao Wang, Xiangyu Zeng, Wenxiao Wu, Tao Chen, Limin Wang, Yanwei Fu

机构 * Shanghai Innovation Institute(上海创新研究院) Fudan University(复旦大学) Nanjing University(南京大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Huazhong University of Science and Technology(华中科技大学)

AI总结 提出StreamOV框架,利用多模态证据引导的长短期记忆和隐状态驱动的触发机制,实现流式全视频理解中的在线推理与主动响应,并在新基准SOVBench上取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09943 2026-05-26 cs.AI

PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMs

PathMem: 面向病理学多模态大模型的认知对齐记忆转换

Jinyue Li, Yuci Liang, Qiankun Li, Xinheng Lyu, Jiayu Qian, Huabao Chen, Kun Wang, Zhigang Zeng, Anil Anthony Bharath, Yang Liu

机构 * University of Science and Technology of China(中国科学技术大学) Shenzhen University(深圳大学) Nanyang Technological University(南洋理工大学) Imperial College London(伦敦帝国学院) Huazhong University of Science and Technology(华中科技大学)

AI总结 提出PathMem框架,通过长期记忆与工作记忆的动态转换机制,实现结构化病理知识整合与可解释记忆控制,在WSI报告生成和开放诊断任务上达到SOTA。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03827 2026-05-26 cs.CV cs.RO

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

LIBERO-PRO:超越记忆的视觉-语言-动作模型鲁棒与公平评估

Xueyang Zhou, Yangming Xu, Guiyao Tie, Yongchao Chen, Guowen Zhang, Duanfeng Chu, Pan Zhou, Lichao Sun

机构 * Huazhong University of Science and Technology(华中科技大学) College of AI, Tsinghua University(清华大学人工智能学院) Wuhan University of Technology(武汉理工大学) Lehigh University(莱斯大学)

AI总结 针对LIBERO基准评估中的记忆偏差问题,提出LIBERO-PRO扩展基准,通过在操作对象、初始状态、任务指令和环境四个维度施加合理扰动,揭示现有VLA模型性能从90%以上骤降至0.0%的严重缺陷,并呼吁采用鲁棒评估方法。

Comments 10 pages,7 figures, 0 tables

详情

展开后加载摘要…

URL PDF HTML 收藏