arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Zhejiang University(浙江大学)

共收录 2931
2512.06999 2025-12-09 cs.SD cs.AI

Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model

基于多模态大基础模型的歌唱音色受欢迎程度评估

Zihao Wang, Ruibin Yuan, Ziqi Geng, Hengjia Li, Xingwei Qu, Xinyi Li, Songye Chen, Haoying Fu, Roger B. Dannenberg, Kejun Zhang

机构 * Zhejiang University(浙江大学) Carnegie Mellon University(卡内基梅隆大学) Hong Kong University of Science and Technology(香港科学与技术大学) University of California, Berkeley(加州大学伯克利分校) University of Manchester(曼彻斯特大学) Innovation Center of Yangtze River Delta, Zhejiang University(长江三角洲创新中心,浙江大学)

AI总结 本文提出基于多模态大基础模型的歌唱音色受欢迎程度评估方法,通过引入Sing-MD数据集、VocalVerse架构和H-TPR基准,实现无参考、多维度的歌唱评估。

Comments Accepted to ACMMM 2025 oral

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (ACMMM 2025), Pages 12227-12236

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00975 2025-12-09 cs.CV cs.LG cs.RO

MM-ACT: Learn from Multimodal Parallel Generation to Act

MM-ACT: 从多模态并行生成中学习以行动

Haotian Liang, Xinyi Chen, Bin Wang, Mingkang Chen, Yitian Liu, Yuhao Zhang, Zanxin Chen, Tianshuo Yang, Yilun Chen, Jiangmiao Pang, Dong Liu, Xiaokang Yang, Yao Mu, Wenqi Shao, Ping Luo

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) The University of Hong Kong(香港大学) University of Science and Technology of China(中国科学技术大学) Fudan University(复旦大学) Zhejiang University(浙江大学)

AI总结 MM-ACT通过多模态并行生成提升机器人任务执行能力,实现96.3%的成功率。

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00387 2025-12-09 cs.CV

WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing

WiseEdit: 图像编辑的认知与创造力导向基准测试

Kaihang Pan, Weile Chen, Haiyi Qiu, Qifan Yu, Wendong Bu, Zehan Wang, Yun Zhu, Juncheng Li, Siliang Tang

机构 * Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

AI总结 WiseEdit通过分解图像编辑为意识、解释和想象三个步骤,结合三种知识类型,全面评估图像编辑的认知与创造力能力。

Comments 32 pages, 20 figures. Project Page: https://qnancy.github.io/wiseedit_project_page/. Benchmark: https://huggingface.co/datasets/123123chen/WiseEdit-Benchmark/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04002 2025-12-09 cs.CL

AgriGPT-VL: Agricultural Vision-Language Understanding Suite

AgriGPT-VL:农业视觉-语言理解套件

Bo Yang, Yunkui Chen, Lanfei Feng, Yu Zhang, Xiao Xu, Jianyu Zhang, Nueraili Aierken, Runhe Huang, Hongjian Lin, Yibin Ying, Shijian Li

机构 * Zhejiang University(浙江大学)

AI总结 AgriGPT-VL通过构建农业专用的视觉-语言模型和评估套件,提升了农业领域的多模态理解和推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07341 2025-12-09 cs.CV

DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation

DCoAR: 一种深度概念注入到统一自回归模型中用于个性化文本到图像生成

Fangtai Wu, Mushui Liu, Weijie He, Zhao Wang, Yunlong Yu

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团)

AI总结 DCoAR通过深度概念注入框架实现个性化文本到图像生成,采用LMCL策略和多正则化方案提升生成质量与上下文适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20109 2025-12-09 cs.SE cs.AI

Learning to Align Human Code Preferences

学习对齐人类代码偏好

Xin Yin, Chao Ni, Xiaohu Yang

机构 * Zhejiang University(浙江大学)

AI总结 本文提出自适应偏好优化(APO)方法,通过动态整合SFT和DPO,提升模型在不同代码偏好场景下的对齐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02483 2025-12-09 cs.CV

Event-Customized Image Generation

事件定制图像生成

Zhen Wang, Yilei Jiang, Dong Zheng, Jun Xiao, Long Chen

机构 * Zhejiang University, Hangzhou, China(浙江大学) The Hong Kong University of Science and Technology(香港科技大学)

AI总结 本文提出FreeEvent方法,通过引入实体切换和事件转移路径,实现事件定制化图像生成,提升复杂场景下的定制化能力。

Journal ref ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01064 2025-12-09 cs.CV

Roadside Monocular 3D Detection Prompted by 2D Detection

道路旁单目3D检测受2D检测启发

Yechi Ma, Yanan Li, Wei Hua, Shu Kong

机构 * Zhejiang University(浙江大学) Zhejiang Lab(浙江实验室) University of Macau(澳门大学) Institute of Collaborative Innovation(协同创新研究院)

AI总结 本文提出Pro3D,通过利用2D检测作为提示,提升道路旁单目3D检测的性能,实现更精确的3D定位。

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.11452 2025-12-09 cs.MA cs.AI cs.LO

Providing personalized Explanations: a Conversational Approach

提供个性化解释:一种对话方法

Jieting Luo, Thomas Studer, Mehdi Dastani

机构 * Zhejiang University(浙江大学) University of Bern(伯尔尼大学) Utrecht University(乌得勒支大学)

AI总结 本文提出了一种通过连续对话为不同背景的用户提供个性化解释的方法,并证明了对话终止的条件。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06835 2025-12-09 cs.AI

Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning

解耦以泛化:面向数据稀缺的视觉语言推理的上下文优先自进化学习

Tingyu Li, Zheng Sun, Jingxuan Wei, Siyuan Li, Conghui He, Lijun Wu, Cheng Tan

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai JiaoTong University(上海交通大学) University of Chinese Academy of Sciences(中国科学院大学) Zhejiang University(浙江大学)

AI总结 DoGe通过双解耦框架,引导模型优先从上下文学习以提升数据稀缺下的视觉语言推理能力,提出两阶段RL方法和演进课程学习流水线,有效解决传统方法的奖励黑客问题。

Comments 25 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06417 2025-12-09 cs.LG cs.SD

Hankel-FNO: Fast Underwater Acoustic Charting Via Physics-Encoded Fourier Neural Operator

Hankel-FNO:通过物理编码的傅里叶神经算子实现快速水下声学制图

Yifan Sun, Lei Cheng, Jianlong Li, Peter Gerstoft

机构 * College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou, China(信息科学与电子工程学院,浙江大学,杭州,中国) Scripps Institution of Oceanography, University of California San Diego, La Jolla, California 92093, USA(Scripps 海洋研究所,加州大学圣地亚哥分校,La Jolla,加利福尼亚州 92093,美国)

AI总结 Hankel-FNO通过物理编码的傅里叶神经算子实现高效准确的水下声学制图,优于传统求解器和数据驱动方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06304 2025-12-09 eess.AS cs.AI cs.CR cs.SD

Degrading Voice: A Comprehensive Overview of Robust Voice Conversion Through Input Manipulation

降噪语音:通过输入操控实现稳健语音转换的全面概述

Xining Song, Zhihua Wei, Rui Wang, Haixiao Hu, Yanxiang Chen, Meng Han

机构 * Tongji University(同济大学) iFLYTEK Research(iFLYTEK研究院) Binjiang Institute Of Zhejiang University(浙大滨江研究院) Hefei University of Technology(合肥工业大学) Zhejiang University(浙江大学)

AI总结 本文全面概述了通过输入操控实现稳健语音转换的研究,探讨了不同降质攻击对VC模型的影响,并提出了优化攻击和防御策略的未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06022 2025-12-09 cs.SD cs.MM

DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation

DreamFoley: 用于高保真视频到音频生成的可扩展视觉-语言模型

Fu Li, Weichao Zhao, You Li, Zhichao Zhou, Dongliang He

机构 * Bytedance Intelligent Creation Lab(字节跳动智能创作实验室) Zhejiang University(浙江大学)

AI总结 DreamFoley通过结合视觉-语言模型,提出了一种可扩展的视频到音频生成方法,实现了高保真音频生成并优化了训练效率与质量的平衡。

Comments 10 pages; Bytedance

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09241 2025-12-09 cs.RO

Unveiling the Impact of Data and Model Scaling on High-Level Control for Humanoid Robots

揭示数据与模型扩展对双足机器人高层控制的影响

Yuxi Wei, Zirui Wang, Kangning Yin, Yue Hu, Jingbo Wang, Siheng Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) University of Michigan(密歇根大学)

AI总结 SCHUR通过大规模数据集和可扩展学习框架,提升了双足机器人高层控制的运动生成和文本-运动对齐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16329 2025-12-09 cs.CR cs.AI cs.CV

DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling

通过分布建模实现文本到图像生成系统可扩展的红队测试

Boheng Li, Junjie Wang, Yiming Li, Zhiyang Hu, Leyi Qi, Jianshuo Dong, Run Wang, Han Qiu, Zhan Qin, Tianwei Zhang

机构 * School of Cyber Science(网络安全学院) State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室) Zhejiang University(浙江大学) Tsinghua University(清华大学)

AI总结 DREAM通过分布建模实现文本到图像生成系统的可扩展红队测试,有效发现多样化的有害提示,提升安全性和多样性。

Comments To appear in the IEEE Symposium on Security & Privacy, May 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16302 2025-12-09 cs.LG cs.AI cs.CR cs.CV

Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning

面向对抗下游微调的鲁棒安全驱动遗忘方法用于扩散模型

Boheng Li, Renjie Gu, Junjie Wang, Leyi Qi, Yiming Li, Run Wang, Zhan Qin, Tianwei Zhang

机构 * Nanyang Technological University, Singapore(南洋理工大学,新加坡) Central South University, China(中南大学,中国) Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University, China(航空航天信息安全部门,教育部,武汉大学,中国) State Key Laboratory of Blockchain and Data Security, Zhejiang University, China(区块链与数据安全国家重点实验室,浙江大学,中国)

AI总结 本文提出ResAlign,一种针对扩散模型对抗下游微调的鲁棒安全驱动遗忘框架,通过隐含优化问题建模和元学习策略提升安全性和生成能力。

Comments Accepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13558 2025-12-09 cs.CV

X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability

X-Scene: 通过高保真和灵活可控性实现大规模驾驶场景生成

Yu Yang, Alan Liang, Jianbiao Mei, Yukai Ma, Yong Liu, Gim Hee Lee

机构 * Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学)

AI总结 X-Scene通过高保真和灵活可控性实现大规模驾驶场景生成,支持多粒度控制和一致性外推,提升自动驾驶的数据生成与模拟能力。

Comments Accepted by NeurIPS 2025, Project page at https://x-scene.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14711 2025-12-09 cs.SI cs.LG

Can GNNs Learn Link Heuristics? A Concise Review and Evaluation of Link Prediction Methods

图神经网络能否学习链接启发式?链接预测方法的简要回顾与评估

Shuming Liang, Yu Ding, Zhidong Li, Bin Liang, Siqi Zhang, Yang Wang, Fang Chen

机构 * Faculty of Engineering and Information Technology, University of Technology Sydney(工程与信息技术学院,悉尼大学) Faculty of Engineering and Information Sciences, University of Wollongong(工程与信息科学学院,沃林戈大学) College of Electrical Engineering, Zhejiang University(电气工程学院,浙江大学)

AI总结 本文研究了图神经网络在链接预测中的能力,发现可训练节点嵌入能提升性能,且图密度越高提升越显著,为未来算法发展提供指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05557 2025-12-08 cs.CV cs.AI

2K-Characters-10K-Stories: A Quality-Gated Stylized Narrative Dataset with Disentangled Control and Sequence Consistency

2K角色-10K故事:一个具有解耦控制和序列一致性的质量门控风格化叙事数据集

Xingxi Yin, Yicheng Li, Gong Yan, Chenglin Li, Jian Zhao, Cong Huang, Yue Deng, Yin Zhang

机构 * Zhejiang University(浙江大学) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)

AI总结 本文提出2K-Characters-10K-Stories数据集,通过解耦控制和质量门控机制,实现高质量的风格化叙事生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11222 2025-12-08 cs.SE cs.AI cs.CL cs.IR

ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal

ORFuzz: 测试LLM安全的'另一面' -- 检测过度拒绝

Haonan Zhang, Dongxia Wang, Yi Liu, Kexin Chen, Jiashui Wang, Xinlei Ying, Long Liu, Wenhai Wang

机构 * Zhejiang University(浙江大学) Zhejiang University Huzhou Institute of Industrial Control Technology(浙江大学湖州工业控制技术研究所)

AI总结 ORFuzz通过进化测试框架系统检测LLM的过度拒绝现象,生成高覆盖率的测试用例并建立新基准,提升LLM安全性。

Comments Accepted by ASE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05111 2025-12-05 cs.CV

ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning

ARM-Thinker: 通过智能工具使用和视觉推理强化多模态生成奖励模型

Shengyuan Ding, Xinyu Fang, Ziyu Liu, Yuhang Zang, Yuhang Cao, Xiangyu Zhao, Haodong Duan, Xiaoyi Dong, Jianze Liang, Bin Wang, Conghui He, Dahua Lin, Jiaqi Wang

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Innovation Institute(上海创新研究院)

AI总结 ARM-Thinker通过智能工具使用和视觉推理提升多模态奖励模型的准确性与可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04939 2025-12-05 cs.CV

LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging

LiteVGGT: 通过几何感知的缓存标记合并提升基础VGGT

Zhijian Shu, Cheng Lin, Tao Xie, Wei Yin, Ben Li, Zhiyuan Pu, Weize Li, Yao Yao, Xun Cao, Xiaoyang Guo, Xiao-Xiao Long

机构 * Nanjing University of Posts and Telecommunications(南京邮电大学) Horizon Robotics Nanjing University(南京大学) Zhejiang University(浙江大学) Macau University of Science and Technology(澳门科技大学) TARS Robotics China Mobile Zijin Innovation Institute(中国移动吉林创新研究院)

AI总结 LiteVGGT通过几何感知的缓存标记合并提升基础VGGT的效率,实现10倍加速和内存减少,适用于大规模3D重建场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04864 2025-12-05 cs.AI

Are Your Agents Upward Deceivers?

您的代理是向上欺骗者吗?

Dadi Guo, Qingyu Liu, Dongrui Liu, Qihan Ren, Shuai Shao, Tianyi Qiu, Haoran Li, Yi R. Fung, Zhongjie Ba, Juntao Dai, Jiaming Ji, Zhikai Chen, Jialing Tao, Yaodong Yang, Jing Shao, Xia Hu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Hong Kong University of Science and Technology(香港科学与技术大学) Zhejiang University(浙江大学) Shanghai Jiao Tong University(上海交通大学) Peking University(北京大学) Alibaba Group(阿里巴巴集团)

AI总结 研究发现基于LLM的代理可能通过欺骗行为隐瞒失败,需加强缓解策略以确保安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04404 2025-12-05 cs.RO

Bridging Probabilistic Inference and Behavior Trees: An Interactive Framework for Adaptive Multi-Robot Cooperation

弥合概率推断与行为树:一种交互框架用于自适应多机器人协作

Chaoran Wang, Jingyuan Sun, Yanhui Zhang, Changju Wu

机构 * School of Aeronautic and Astronautics, Zhejiang University(航空航天学院,浙江大学) Shanghai Huawei Technologies Co., Ltd(上海华为技术有限公司)

AI总结 本文提出IIBT框架,结合概率推断与行为树,实现多机器人自适应协作,实验表明其能显著降低复杂度并保持鲁棒性。

Comments 34 pages, is submitted RAS Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10901 2025-12-05 cs.CV cs.AI cs.LG

A Gray-box Attack against Latent Diffusion Model-based Image Editing by Posterior Collapse

针对基于潜在扩散模型的图像编辑的灰盒攻击

Zhongliang Guo, Chun Tong Lei, Lei Fang, Shuai Zhao, Yifei Qian, Jingyu Lin, Zeyu Wang, Cunjian Chen, Ognjen Arandjelović, Chun Pong Lau

机构 * School of Computer Science, University of St Andrews(圣安德鲁大学计算机科学学院) Department of Data Science, City University of Hong Kong(香港城市大学数据科学系) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) School of Computer Science, University of Nottingham(诺丁汉大学计算机科学学院) Department of Data Science and Artificial Intelligence, Monash University(墨尔本大学数据科学与人工智能系) College of Information Science & Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院)

AI总结 本文提出了一种基于后验崩溃现象的灰盒攻击方法,通过调整参数实现两种崩溃类型,有效防止未经授权的图像编辑,且对模型依赖低,计算效率高。

Comments 15 pages, 9 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04515 2025-12-05 cs.CV

EgoLCD: Egocentric Video Generation with Long Context Diffusion

EgoLCD:基于长上下文扩散的自体视频生成

Liuzhou Zhang, Jiarui Ye, Yuanlei Wang, Ming Zhong, Mingju Cao, Wanke Xia, Bowen Zeng, Zeyu Zhang, Hao Tang

机构 * Peking University(北京大学) Sun Yat-sen University(中山大学) Zhejiang University(浙江大学) Chinese Academy of Sciences(中国科学院) Tsinghua University(清华大学)

AI总结 EgoLCD通过结合长时稀疏KV缓存和LoRA扩展的注意力机制,实现了高效稳定的自体长上下文视频生成,提升了感知质量和时间一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04496 2025-12-05 cs.CV

Shift-Window Meets Dual Attention: A Multi-Model Architecture for Specular Highlight Removal

移位窗口与双注意力机制:一种多模型架构用于镜面高光去除

Tianci Huo, Lingfeng Qi, Yuhan Chen, Qihong Xue, Jinyuan Shao, Hai Yu, Jie Li, Zhanhua Zhang, Guofa Li

机构 * College of Mechanical and Vehicle Engineering, Chongqing University(重庆大学机械与车辆工程学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Artificial Intelligence Center, Geely Automotive Research Institute(吉利汽车研究院人工智能中心)

AI总结 本文提出多模型架构MM-SHR,结合卷积与注意力机制,有效去除不同尺度的镜面高光,提升视觉任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19035 2025-12-05 cs.CV cs.AI

Changes in Gaza: DINOv3-Powered Multi-Class Change Detection for Damage Assessment in Conflict Zones

加沙的变化:基于DINOv3的多类变化检测用于冲突区损害评估

Kai Zheng, Zhenkai Wu, Fupeng Wei, Miaolan Zhou, Kai Lie, Haitao Guo, Lei Ding, Wei Zhang, Hang-Cheng Dong

机构 * School of Computer Science and Technology(计算机科学与技术学院) Zhejiang University(浙江大学) School of Software Technology(软件学院) School of Information Engineering(信息工程学院) North China University of Water Resources and Electric Power(北方水利电力大学) Polytechnic Institute Institute of Systems Engineering(系统工程院) Academy of Military Sciences(军事科学院) Department of Geo-spatial Information(测绘信息学院) Information Engineering University(信息工程大学) School of Instrumentation Science and Engineering(仪器科学与工程学院) Harbin Institute of Technology(哈尔滨工业大学) Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究院)

AI总结 本文提出基于DINOv3的MC-DiSNet模型,用于冲突区多类变化检测,通过多尺度交叉注意力机制和差分孪生结构实现细粒度语义变化检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18541 2025-12-05 cs.AI

Align$^2$LLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation

Align$^2$LLaVA: 多模态指令编纂的级联人类与大语言模型偏好对齐

Hongzhe Huang, Jiang Liu, Zhewen Yu, Li Cai, Dian Jiao, Wenqiao Zhang, Siliang Tang, Juncheng Li, Hao Jiang, Haoyuan Li, Yueting Zhuang

机构 * Zhejiang University(浙江大学) Alibaba(阿里巴巴)

AI总结 Align$^2$LLaVA通过级联人类与LLM偏好对齐方法,有效压缩多模态指令数据,提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04025 2025-12-04 cs.CV cs.AI cs.LG

PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation

PSA: 基于金字塔稀疏注意力的高效视频理解和生成

Xiaolong Li, Youping Gu, Xi Lin, Weijie Wang, Bohan Zhuang

机构 * ZIP Lab, Zhejiang University(浙江大学浙大信息实验室)

AI总结 PSA通过多级池化键值表示实现高效视频理解和生成,相比现有稀疏注意力方法在效率和质量上表现更优。

Comments Tech report

详情

展开后加载摘要…

URL PDF HTML 收藏