arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Zhejiang University(浙江大学)

2026-08-03 至 2026-08-03 共收录 3
2607.29531 2026-08-03 cs.CV q-bio.NC 新提交

Multi-Source Multi-View Graph Domain Adaptation with Hyperbolic Residual Encoding for Cross-Site MDD Identification from rs-fMRI

基于双曲残差编码的多源多视图图域适应用于跨站点静息态fMRI的重度抑郁症识别

Zhanpeng Zheng, Xiran Chen, Haiteng Jiang, Renjie Tian, Qinyu Cai, Jiexi Liu, Xiaofeng Chen, Weikai Li, Yansu Wang

机构 * School of Computer and Artificial Intelligence, Shandong Jianzhu University(山东建筑大学计算机与人工智能学院) Institute of Fundamental and Frontier Sciences, University of Electronic Science and Technology of China(电子科技大学基础与前沿研究院) School of Mathematics and Statistics, Chongqing Jiaotong University(重庆交通大学数学与统计学院) State Key Laboratory of Brain-machine Intelligence, Zhejiang University(浙江大学脑机智能国家重点实验室) Institute of Computer Vision and Traffic Image Understanding, School of Information Science and Engineering, Chongqing Jiaotong University(重庆交通大学信息科学与工程学院计算机视觉与交通图像理解研究所) School of Life Sciences, Westlake University(西湖大学生命科学学院) School of Computer and Artificial Intelligence, Nanjing University of Finance and Economics(南京财经大学计算机与人工智能学院)

AI总结 该研究针对跨站点rs-fMRI的MDD识别难题,提出结合双曲残差编码、双流自适应融合及类别级对齐的多源多视图图域适应框架,在七个目标域取得73.60%平均准确率,实现有效泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29079 2026-08-03 cs.CL 新提交

Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models

更快但不同:加速多模态扩散语言模型中的内容漂移诊断与控制

Yaoxuan Dou, Yang Shu

机构 * School of Mathematics and Statistics, Beijing Institute of Technology(北京理工大学数学与统计学院) Zhejiang University(浙江大学)

AI总结 该研究针对加速多模态扩散语言模型的内容漂移问题,通过实验诊断出漂移源于过期状态,提出缩短KV缓存刷新区间的控制方法,在1.3倍加速时实现近乎完全一致,且未发现事实错误差异。

Comments 9 pages, 4 figures, 6 tables. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28678 2026-08-03 cs.AI cs.CV 新提交

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

ViSAGE:为长视频理解构建自修正记忆

Xinkui Zhao, Enbo Chen, Yifan Zhang, Chang Liu, Guanjie Cheng, Naibo Wang, Yueshen Xu

机构 * Zhejiang University(浙江大学) Xidian University(西安电子科技大学)

AI总结 ViSAGE是一种多模态智能体记忆框架,通过跨模态绑定、双向记忆精调及多智能体交叉验证解决长视频理解中的实体混淆等问题,较最强基线准确率提升5.9%。

Comments Accept by ACMMM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏