Helios: Real Real-Time Long Video Generation Model
Helios:实时长视频生成模型
机构 * Peking University(北京大学)
AI总结 Helios是首个无需并行框架、能实时生成长视频且质量媲美基线模型的14B视频生成模型。
Comments Page: pku-yuangroup.github.io/Helios-Page
高校专区
Helios:实时长视频生成模型
机构 * Peking University(北京大学)
AI总结 Helios是首个无需并行框架、能实时生成长视频且质量媲美基线模型的14B视频生成模型。
Comments Page: pku-yuangroup.github.io/Helios-Page
IPD: 通过离线规划蒸馏提升序列策略
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Peking University(北京大学)
AI总结 IPD通过引入离线规划蒸馏技术,提升离线强化学习中序列策略的决策稳定性和性能。
谱手术:通过梯度引导的奇异值重新加权实现LoRA的训练免费优化
机构 * School of Computing(计算学院) ; Information Systems, Singapore Management University(信息系统,新加坡管理大学) ; School of Computing, National University of Singapore(计算学院,新加坡国立大学) ; State Key Laboratory for Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室,北京大学)
AI总结 Spectral Surgery通过梯度引导的奇异值重新加权优化LoRA适配器,提升模型性能。
视差对齐所有内容:一种用于分布式多视图图像压缩的 OmniParallax 注意机制
机构 * The National Engineering Laboratory for Video Technology, School of Computer Science, Peking University(国家视频技术工程实验室,计算机科学学院,北京大学) ; Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究院(AIR),清华大学) ; Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)
AI总结 ParaHydra通过OmniParallax注意机制和多信息融合模块,实现了分布式多视图图像压缩的高效编码与解码,显著提升压缩效率并降低计算开销
Comments Accepted by CVPR 2026
EvalMVX: 一种用于神经3D重建的统一基准测试,适用于多样化的多视角设置
机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ; School of Mechanical Engineering, Shanghai Jiao Tong University(上海交通大学机械工程学院) ; Xiong’an Aerospace Information Research Institute(雄安航空航天信息研究所) ; State Key Laboratory of Multimedia Information Processing, School of Computer Science Peking University(北京大学计算机科学学院多媒体信息处理国家重点实验室)
AI总结 EvalMVX是一个包含25个物体的真实世界数据集,用于评估神经3D重建方法在不同多视角设置下的性能。
CLASH:通过增强的仿真到现实混合化进行碰撞学习以弥合现实差距
机构 * School of Mathematical Sciences, Peking University(北京大学数学科学学院) ; School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) ; School of Mathematical Sciences, Beijing Normal University(北京师范大学数学科学学院)
AI总结 CLASH通过增强的仿真到现实混合化方法,有效提升碰撞预测精度并减少仿真计算时间,增强机器人策略在现实中的鲁棒性。
通过感知与推理增强改进医疗视觉强化微调
机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ; Emory University(埃默里大学) ; Sichuan University(四川大学) ; Peking University(北京大学)
AI总结 本研究提出VRFT-Aug框架,通过增强感知与推理能力,改进医疗影像领域的强化微调效果,提升模型在医学应用中的可靠性与推理能力。
Comments CPAL 2026
Journal ref 2026 Conference on Parsimony and Learning (CPAL)
TIGeR: 视觉-语言模型中集成工具的几何推理
机构 * Beihang University(北洋大学) ; Beijing Academy of Artificial Intelligence(北京人工智能研究院) ; State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机科学学院多媒体信息处理国家重点实验室)
AI总结 TIGeR通过集成工具实现视觉-语言模型的几何推理,提升机器人操作的精确度。
Comments 8 pages, 6 figures
ToolVQA: 一个用于多步推理VQA的外部工具数据集
机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机科学技术研究院)
AI总结 ToolVQA是一个用于多步推理VQA的外部工具数据集,通过真实场景和复杂推理任务提升模型在现实工具使用中的泛化能力。
Comments Project page: https://fugtemypt123.github.io/ToolVQA-website/
稀疏自编码器的局限性:一种理论框架和加权修复方法
机构 * State Key Lab of General AI, School of Intelligence Science and Technology, Peking University(1 人工智能通用基础研究实验室,智能科学与技术学院,北京大学) ; Amazon AGI SF Lab(2 亚马逊人工智能实验室) ; Institute for Artificial Intelligence, Peking University(3 人工智能研究院,北京大学)
AI总结 本文提出了一种理论框架和加权修复方法,以解决稀疏自编码器在恢复真实单义特征时的局限性,通过改进特征恢复能力提升可解释性。
Comments Accepted to ICLR2026
困难示例损害无监督对比学习:一种理论视角
机构 * State Key Lab of General AI, School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院人工智能国家重点实验室) ; School of Engineering, Hong Kong University of Science and Technology(香港科技大学工程学院) ; Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)
AI总结 本文从理论角度探讨了困难示例对无监督对比学习的影响,发现移除困难示例可提升下游分类性能,并通过理论分析和实验证明了其有效性。
Comments Accepted to ICLR 2026 as an Oral Presentation
FSMLP: 基于简单理论的多层感知机在频域中建模通道依赖性
机构 * Communication University of China(通信大学) ; Peking University(北京大学) ; Zhejiang University(浙江大学) ; Tsinghua University(清华大学)
AI总结 FSMLP通过引入简单体MLP层,有效缓解时间序列预测中通道依赖的过拟合问题,提升预测精度与效率。