arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Peking University(北京大学)

2026-05-26 至 2026-05-26 共收录 37
2605.26086 2026-05-26 cs.AI

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

Claw-Anything: 对更广泛访问用户数字世界的始终在线个人助手的基准测试

Yusong Lin, Xinyuan Liang, Haiyang Wang, Qipeng Gu, Siqi Cheng, Jiangui Chen, Shuzhe Wu, Feiyang Pan, Lue Fan, Sanyuan Zhao, Dandan Tu

机构 * Beijing Institute of Technology(北京理工大学) Huawei Technologies Co., Ltd(华为技术有限公司) Peking University(北京大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

AI总结 提出Claw-Anything基准测试,通过扩展长期活动历史、相互依赖的后端服务以及跨多设备的GUI和CLI交互三个维度,评估大型语言模型代理在始终在线环境下的性能,发现GPT-5.5仅达34.5% pass@1,并发布自动化数据生成管道提升基线模型23.7%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26003 2026-05-26 cs.CV

Towards 3D heart mesh generation using contactless radar imaging and physics-informed neural network

基于非接触式雷达成像和物理信息神经网络的3D心脏网格生成

Jinye Li, Chenxi Fu, Minghang Zheng, Yang Liu, Xiahai Zhuang, Qingchao Chen

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Fudan University(复旦大学) Peking University(北京大学)

AI总结 提出SAR2Mesh框架,通过粗到细的网格变形过程,结合几何感知特征投影和物理信息雷达损失,从合成孔径雷达图像重建高保真3D心脏几何结构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25922 2026-05-26 cs.CV

Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models

闭环双向提示用于视觉语言模型的对抗鲁棒性

Xiao Liu, Jiaxiang Liu, Boci Peng, Boren Hu, Yusong Wang, Xiwen Chen, Prayag Tiwari, Liming Zhang, Mingkun Xu

机构 * University of Macau(澳门大学) Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) Peking University(北京大学) Independent Researcher(独立研究员) Institute of Science Tokyo(东京科学研究院) Morgan Stanley(摩根大通) Halmstad University(哈马碧大学)

AI总结 针对视觉语言模型在对抗扰动下跨模态语义对齐脆弱的问题,提出闭环双向提示方法,通过动态反馈循环恢复跨模态一致性,并引入语义锚点约束循环更新,实现实例自适应保护,在11个数据集上达到最先进的鲁棒性和泛化性能。

Comments 24 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25851 2026-05-26 cs.RO

RePlan-Bot: Multi-Level Replanning for Embodied Instruction Following

RePlan-Bot:面向具身指令跟随的多级重规划

Xicheng Gong, Guozheng Sun, Peiran Xu, Yadong Mu

机构 * Peking University(北京大学) Tsinghua University(清华大学)

AI总结 提出RePlan-Bot,通过多级连续重规划(高层LLM审计器、常识引导搜索、轻量级ViT校正器)解决具身指令跟随中的长时规划和不可逆状态变化问题,在ALFRED基准上取得最佳性能。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25813 2026-05-26 cs.RO

Extending Embodied Question Answering from Perception to Decision

将具身问答从感知扩展到决策

Xicheng Gong, Qiwei Li, Peiran Xu, Yadong Mu

机构 * Peking University(北京大学) XYZ Embodied AI(XYZ具身AI)

AI总结 提出大规模具身问答数据集EQA-Decision和基线模型RoboDecision,系统覆盖静态场景构建、空间理解、任务动态推理和即时决策四个维度,以统一框架评估具身环境中的感知、推理和行动级决策。

Comments 11 pages,4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25802 2026-05-26 cs.CV

Rethinking VLM Representation for VLA Initialization

重新思考用于VLA初始化的VLM表示

Weifeng Lin, Siyuan Huang, Hao Li, Tingwei Chen, Ruichuan An, Xinyu Wei, Jianbo Liu, Hongsheng Li

机构 * CUHK(香港中文大学) PolyU Peking University(北京大学) ACE Robotics(ACE机器人)

AI总结 本文通过控制表示设计问题,沿能力级具身VQA监督、参数更新策略和机器人数据预训练三个轴,研究VLA初始化,发现保留预训练VLM表示对动作性能至关重要,而LoRA比全微调提供更可靠的初始化,分阶段基于LoRA的训练获得最强变体。

Comments 9 main-text pages, 5 appendix pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25786 2026-05-26 cs.LG cs.AI

NPSolver: Neural Poisson Solver with Iterative Physics Supervision

NPSolver: 具有迭代物理监督的神经泊松求解器

Bocheng Zeng, Rui Zhang, Runze Mao, Mengtao Yan, Xuan Bai, Yang Liu, Zhi X. Chen, Hao Sun

机构 * Gaoling School of Artificial Intelligence(高岭人工智能学院) Renmin University of China(中国人民大学) School of Mechanics and Engineering Science(力学与工程科学学院) Peking University(北京大学) AI for Science Institute(AI for Science研究院) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 提出NPSolver,通过迭代物理监督(利用少量PCG步骤)训练无标签的神经泊松求解器,并引入边界感知Transolver架构,在2D/3D不规则几何上优于物理信息和数据驱动基线。

Comments kdd 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25725 2026-05-26 cs.CV

TriDP-PTM: a three-stage distortion-perception tradeoff guides the pre-training model for radar cardiac sensing

TriDP-PTM:三阶段失真-感知权衡引导的预训练模型用于雷达心脏感知

Jinye Li, Aidong Men, Yang Liu, Qingchao Chen

机构 * National Institute of Health Data Science, Peking University(北京大学国家健康数据科学研究院) Institute of Medical Technology, Peking University(北京大学医学技术研究院) Beijing University of Posts and Telecommunications(北京邮电大学) School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)

AI总结 提出三阶段失真-感知预训练模型(TriDP-PTM),通过雷达-心电图-任务间接路径和复合损失函数,在合作竞争阶段实现最佳下游临床精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25698 2026-05-26 cs.LG cs.AI

How Should LLMs Consume High-Quality Data? Optimal Data Scheduling via Quality-Aware Functional Scaling Laws

LLM应如何消费高质量数据?通过质量感知的功能缩放定律实现最优数据调度

Zhitao Zhu, Xili Wang, Shizhe Wu, Jiawei Fu, Xiaoqing Liu

机构 * Peking University(北京大学) Meituan(美团)

AI总结 本文通过引入数据质量维度扩展功能缩放定律,解析求解了联合数据质量和批次大小调度问题,揭示了高质量数据的双重角色,并提出了Drop-Stable-Rampup调度策略,在15B MoE模型上相比WSD和余弦衰减分别提升平均准确率+1.70和+2.98。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25563 2026-05-26 cs.CV

CodecSplat: Ultra-Compact Latent Coding for Feed-Forward 3D Gaussian Splatting

CodecSplat: 用于前馈式3D高斯泼溅的超紧凑潜在编码

Pengpeng Yu, Runqing Jiang, Qi Zhang, Dingquan Li, Jing Wang, Yulan Guo

机构 * Sun Yat-sen University(中山大学) Peking University(北京大学) Pengcheng Laboratory(鹏城实验室)

AI总结 提出CodecSplat框架,通过将压缩集成到前馈式高斯生成流水线中,利用结构化中间特征表示实现超紧凑场景编码,显著降低存储和传输开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25354 2026-05-26 cs.AI

Context-CoT: Enhancing Context Learning via High-Quality Reasoning Synthesis

Context-CoT:通过高质量推理合成增强上下文学习

Hongbo Jin, Mingnan Zhu, Jingqi Tian, Xu Jiang, Zhongjing Du, Haoran Tang, Siyi Xie, Qiaoman Zhang, Jiayu Ding

机构 * Peking University(北京大学) Xiamen University(厦门大学) Tsinghua University(清华大学)

AI总结 针对大语言模型在动态提取和应用新知识方面的上下文学习能力不足,提出Context-CoT方法,通过合成高质量推理链来增强上下文学习,在CL-Bench上显著提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23163 2026-05-26 cs.CL

Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving

Fast-dDrive:面向自动驾驶的高效块扩散视觉语言模型

Kewei Zhang, Jin Wang, Sensen Gao, Chengyue Wu, Yulong Cao, Songyang Han, Boris Ivanovic, Langechuan Liu, Marco Pavone, Song Han, Daquan Zhou, Enze Xie

机构 * Peking University(北京大学) NVIDIA The University of Hong Kong(香港大学) MIT(麻省理工学院)

AI总结 提出Fast-dDrive,一种块扩散视觉语言动作模型,通过语义单元内双向细化与跨单元因果约束,结合结构化令牌冻结、分段感知训练和推测解码,实现高保真轨迹规划与高效推理,在WOD-E2E和nuScenes上达到最优性能,推理速度提升12倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08213 2026-05-26 cs.CV cs.AI

EditCaption: Human-Refined SFT and HAE-DPO for Image Editing Instruction Synthesis

EditCaption: 用于图像编辑指令合成的人工精炼SFT与HAE-DPO

Xiangyuan Wang, Honghao Cai, Yunhao Bai, Chao Hui, Tianze Zhou, Haohua Chen, Hao Shi, Yuling Wu, Yao Hu, Xu Tang, Yibo Chen, Wei Zhu

机构 * Peking University(北京大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Tsinghua University(清华大学) Beihang University(北京航空航天大学) Xiaohongshu Inc.(小红书公司)

AI总结 提出EditCaption两阶段后训练流程,通过人工精炼SFT和基于难度自适应错误感知DPO(HAE-DPO)提升图像编辑指令合成质量,显著降低关键错误率并超越现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05550 2026-05-26 cs.CL cs.CE

AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery

AutoSOTA:面向最先进AI模型发现的端到端自动化研究系统

Yu Li, Chenyang Shao, Xinyang Liu, Ruotong Zhao, Peijie Liu, Hongyuan Su, Zhibin Chen, Qinglong Yang, Anjie Xu, Yi Fang, Qingbin Zeng, Tianxing Li, Jingbo Xu, Fengli Xu, Yong Li, Tie-Yan Liu

机构 * Department of Electronic Engineering, BNRist, Tsinghua University(电子工程系,北京理工大学,清华大学) Zhongguancun Academy(中关村学院) Peking University(北京大学) University of Science and Technology of China(中国科学技术大学)

AI总结 提出AutoSOTA系统,采用多智能体架构实现从论文复现到模型优化的全自动化,成功发现105个超越原始方法的新SOTA模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04707 2026-05-26 cs.CV

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models

OpenWorldLib: 高级世界模型的统一代码库与定义

DataFlow Team, Bohan Zeng, Daili Hua, Kaixin Zhu, Yifan Dai, Bozhou Li, Yuran Wang, Chengzhuo Tong, Yifan Yang, Mingkun Chang, Jianbin Zhao, Zhou Liu, Hao Liang, Xiaochen Ma, Ruichuan An, Junbo Niu, Zimo Meng, Tianyi Bai, Meiyi Qiang, Huanyao Zhang, Zhiyou Xiao, Tianyu Guo, Qinhan Yu, Runhao Zhao, Zhengpin Li, Xinyi Huang, Yisheng Pan, Yiwen Tang, Juanxi Tian, Yang Shi, Yue Ding, Xinlong Chen, Hongcheng Gao, Minglei Shi, Jialong Wu, Zekun Wang, Yuanxing Zhang, Xintao Wang, Pengfei Wan, Yiren Song, Mike Zheng Shou, Wentao Zhang

机构 * Peking University(北京大学) Zhongguancun Academy(中关村学院) Tsinghua University(清华大学) National University of Singapore(新加坡国立大学) Shanghai Jiao Tong University(上海交通大学) Sun Yat-sen University(中山大学) Beijing Key Laboratory of Data Intelligence and Security(北京数据智能与安全重点实验室) Nanyang Technological University(南洋理工大学)

AI总结 本文提出OpenWorldLib框架,基于对世界模型演化的分析给出清晰定义,并系统分类其核心能力,实现多任务模型的统一集成与高效推理。

Comments 28 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04279 2026-05-26 cs.CL

ECG-R1: Protocol-Guided and Modality-Agnostic MLLM for Reliable ECG Interpretation

ECG-R1: 协议引导且模态无关的可靠心电图解读多模态大语言模型

Jiarui Jin, Haoyu Wang, Xingliang Wu, Xiaocheng Fang, Xiang Lan, Zihan Wang, Deyun Zhang, Bo Liu, Yingying Zhang, Xian Wu, Hongyan Li, Shenda Hong

机构 * School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) National Institute of Health Data Science, Peking University(北京大学健康数据科学国家研究院) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) Tianjin Institute of Cardiology, the Second Hospital of Tianjin Medical University(天津医科大学第二医院心内科) National University of Singapore(新加坡国立大学) Jarvis Lab, Tencent(腾讯 Jarvis实验室) HeartVoice Medical Technology(HeartVoice医疗科技)

AI总结 提出ECG-R1,通过协议引导数据生成、模态解耦架构和强化学习,实现可靠的心电图解读。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15264 2026-05-26 cs.CV

DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion

DriveGen3D: 通过高效视频扩散提升前馈驾驶场景生成

Weijie Wang, Jiagang Zhu, Zeyu Zhang, Xiaofeng Wang, Zheng Zhu, Guosheng Zhao, Chaojun Ni, Haoxiao Wang, Guan Huang, Xinze Chen, Yukun Zhou, Wenkang Qin, Duochao Shi, Haoyun Li, Yicheng Xiao, Donny Y. Chen, Jiwen Lu

机构 * Zhejiang University(浙江大学) GigaAI Tsinghua University(清华大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Peking University(北京大学) Monash University(墨尔本大学)

AI总结 提出DriveGen3D框架,结合快速视频扩散Transformer(FastDrive-DiT)和前馈3D重建模块(FastRecon3D),实现高质量、可控的动态3D驾驶场景生成,在长视频和3D一致性上达到最优。

Comments ICME 2026 Oral, Project Page: https://lhmd.top/drivegen3d

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.07863 2026-05-26 cs.CV cs.AI cs.MM

You Can Ground Earlier than See: An Effective and Efficient Pipeline for Temporal Sentence Grounding in Compressed Videos

你可以比看见更早定位:一种用于压缩视频中时序句子定位的高效流程

Xiang Fang, Daizong Liu, Pan Zhou, Guoshun Nan

机构 * The Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology(大数据安全湖北工程研究中心,网络安全科学与工程学院,华中科技大学) Peking University(北京大学) Beijing University of Posts and Telecommunications(北京邮电大学)

AI总结 提出一种三分支压缩域时空融合框架(TCSF),直接从压缩视频中提取I帧、运动向量和残差特征,实现高效准确的时序句子定位。

Comments Accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.11572 2026-05-26 cs.CV cs.AI cs.IR cs.MM

Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval

多模态跨域对齐网络用于视频时刻检索

Xiang Fang, Daizong Liu, Pan Zhou, Yuchong Hu

机构 * Hubei Key Laboratory of Distributed System Security(湖北分布式系统安全重点实验室) Hubei Engineering Research Center on Big Data Security(湖北大数据安全工程研究中心) School of Cyber Science and Engineering(网络安全学院) Huazhong University of Science and Technology(华中科技大学) Wangxuan Institute of Computer Technology(王轩计算机技术研究所) Peking University(北京大学) School of Computer Science and Technology(计算机科学与技术学院) Key Laboratory of Information Storage System Ministry of Education of China(信息存储系统教育部重点实验室)

AI总结 提出多模态跨域对齐网络,通过域对齐、跨模态对齐和特定对齐三个模块,解决跨域视频时刻检索中域差异和语义鸿沟问题。

Comments Accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.14882 2026-05-26 cs.MM cs.CL cs.CV cs.IR

Hierarchical Local-Global Transformer for Temporal Sentence Grounding

层次化局部-全局Transformer用于时间语句定位

Xiang Fang, Daizong Liu, Pan Zhou, Zichuan Xu, Ruixuan Li

机构 * Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology(大数据安全湖北工程研究中心,华中科技大学网络安全科学与工程学院) Wangxuan Institute of Computer fTechnology, Peking University(王宣计算机技术研究院,北京大学) School of software, Dalian University of Technology(软件学院,大连理工大学) School of Computer Science, and Technology, Huazhong University of Science, and Technology(计算机科学与技术学院,华中科技大学)

AI总结 提出层次化局部-全局Transformer(HLGT),通过建模视频和查询的不同粒度层次及跨模态交互,实现更细粒度的多模态表示,并在三个数据集上取得最先进性能。

Comments Publish in IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25077 2026-05-26 cs.CV

WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models

WorldCraft: 从相机导航到交互式视频世界模型中的物体操控

Bohai Gu, Taiyi Wu, Yueyang Yuan, Jian Liu, Xiaocheng Lu, Dazhao Du, Jie Zhang, Jinxiang Lai, Shuai Yang, Xiaotong Zhao, Alan Zhao, Song Guo

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) AI Technology Center, Tencent Video, Tencent(腾讯视频AI技术中心,腾讯) Wuhan University(武汉大学) Peking University(北京大学)

AI总结 提出WorldCraft框架,通过轨迹控制管道(NWT、SP-LoRA、TASP)将交互式视频世界模型从相机导航扩展到物体级轨迹操控,实现用户指定路径下的物体运动与相机导航共存。

Comments Project page: https://nevsdev.github.io/WorldCraft/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24959 2026-05-26 cs.CV

Three-Step Conditional Diffusion 3D Reconstruction for Light-Field Microscopy

三步条件扩散光场显微三维重建

Qihong Zhao, Shaokang Yan, Zhimin Qiao, Jinjia Wang, Bo Xiong

机构 * Yanshan University(雁山大学) Peking University(北京大学)

AI总结 针对光场显微成像中传统算法分辨率低、伪影重、计算成本高,以及现有学习方法重建精度和泛化能力不足的问题,提出一种基于三步条件扩散的高保真三维重建方法,通过确定性三步采样和轻量条件U-Net实现快速准确重建,并引入类间检测模块增强稳定性。

Comments 10 pages, 6 figures. Accepted to CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24958 2026-05-26 cs.CL cs.AI

SEP-Attack: A Simple and Effective Paradigm for Transfer-Based Textual Adversarial Attack

SEP-Attack:一种简单有效的基于迁移的文本对抗攻击范式

Han Liu, Zhi Xu, Xiaotong Zhang, Feng Zhang, Xiaoming Xu, Wei Wang, Fenglong Ma, Hong Yu

机构 * Dalian University of Technology(大连理工大学) Peking University(北京大学) Macao Polytechnic University(澳门理工学院) The Pennsylvania State University(宾夕法尼亚州立大学)

AI总结 提出SEP-Attack,利用行列式点过程生成多样化的代理集成权重,通过新指标评估预测置信度以计算词重要性并生成对抗样本,在多个数据集和API上显著优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24914 2026-05-26 cs.IR cs.DB cs.LG

MVR-cache: Optimizing Semantic Caching via Multi-Vector Retrieval and Learned Prompt Segmentation

MVR-cache:通过多向量检索和学习型提示分割优化语义缓存

Ali Noshad, Zishan Zheng, Yinjun Wu

机构 * School of Computer Science, Peking University, Beijing, China(北京大学计算机科学学院,北京,中国) School of Information, Renmin University of China, Beijing, China(中国人民大学信息学院,北京,中国)

AI总结 提出MVR-cache方法,利用多向量检索和学习型提示分割模型,通过强化学习优化缓存命中率,在保证正确性的前提下将缓存命中率提升高达37%。

Comments Published in ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24592 2026-05-26 cs.RO

MuGen: Multi-Skill Generative Locomotion Controller for Humanoid Robots

MuGen: 人形机器人的多技能生成式运动控制器

Yusen Feng, Xiang Wang, Heyuan Yao, Zixi Kang, Xinyu Huo, Boyang Yu, Pengyun Qiu, Ruijie Zhao, Baoquan Chen, Libin Liu

机构 * Peking University(北京大学)

AI总结 提出MuGen框架,利用VQ-VAE和教师-学生策略蒸馏,从异构人类运动数据中学习生成式运动表示,使人形机器人能够执行多技能运动并模仿未见过的动作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24530 2026-05-26 cs.CL cs.CV

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval

Unveil: 统一视觉-文本集成与蒸馏的多模态文档检索

Hao Sun, Yingyan Hou, Jiayan Guo, Bo Wang, Chunyu Yang, Jinsong Ni, Yan Zhang

机构 * State Key Laboratory of General Artificial Intelligence, Peking University(北京理工大学通用人工智能国家重点实验室) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空航天信息研究所) Key Laboratory of Target Cognition and Application Technology(目标认知与应用技术重点实验室) Beijing Institute of Technology(北京理工大学) Ucap Cloud(Ucap云)

AI总结 提出Unveil框架,通过视觉-文本嵌入和知识蒸馏实现鲁棒的文档检索,兼顾布局与语义信息。

Comments ACL 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11182 2026-05-26 cs.AI

The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes

在线策略蒸馏的多种面貌:陷阱、机制与修复

Siqi Zhu, Xuyan Ye, Hongyu Lu, Weiye Shi, Ge Liu

机构 * UIUC(伊利诺伊大学香槟分校) Renmin University of China(中国人民大学) Peking University(北京大学)

AI总结 本文通过实证研究分析了在线策略蒸馏(OPD)和在线策略自蒸馏(OPSD)在大语言模型后训练中的有效性、失败机制及修复方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01284 2026-05-26 cs.CV cs.AI cs.CL cs.IR

Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation

证据链:面向迭代检索增强生成的像素级视觉归因

Peiyang Liu, Ziqiang Cui, Xi Wang, Di Liang, Wei Ye

机构 * National Engineering Research Center for Software Engineering, Peking University(软件工程国家级工程研究中心,北京大学) City University of Hong Kong(香港城市大学) Peking University(北京大学) Tencent Technology(腾讯科技)

AI总结 提出Chain of Evidence (CoE)框架,利用视觉语言模型直接对检索到的文档截图进行推理,输出精确边界框以可视化完整推理链,解决迭代检索增强生成中的粗粒度归因和视觉语义丢失问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06626 2026-05-26 cs.LG cs.AI

Grouter: Decoupling Routing from Representation for Accelerated MoE Training

Grouter: 将路由与表示解耦以加速MoE训练

Yuqi Xu, Rizhen Hu, Zihan Liu, Mou Sun, Kun Yuan

机构 * School of Mathematical Sciences, Peking University, Beijing, China(北京大学数学科学学院) Center for Machine Learning Research, Peking University, Beijing, China(北京大学机器学习研究中心) Yuanpei College, Peking University, Beijing, China(北京大学元培学院) Zhejiang Lab, Hangzhou, China(浙江实验室)

AI总结 提出Grouter方法,通过从预训练MoE模型中蒸馏高质量结构作为固定路由器,解耦结构优化与权重更新,显著加速模型收敛并提升训练吞吐量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23916 2026-05-26 cs.CV cs.AI

Topology-Driven Transferability Estimation of Medical Foundation Models for Segmentation

基于拓扑驱动的医学基础模型分割迁移性估计

Jiaqi Tang, Shaoyang Zhang, Xiaoqi Wang, Jiaying Zhou, Yang Liu, Qingchao Chen

机构 * Peking University(北京大学) Hohai University(河海大学) Beijing Normal University-Hong Kong Baptist University United International College(北京师范大学-香港 Baptist大学联合国际学院) National Institute of Health Data Science, Peking University(健康数据科学国家研究院,北京大学) Institute of Medical Technology, Peking University(北京大学医学技术研究院) State Key Laboratory of General Artificial Intelligence, Peking University(通用人工智能国家重点实验室,北京大学)

AI总结 提出拓扑驱动迁移性估计框架,通过全局表示拓扑散度、局部边界感知拓扑一致性和任务自适应融合,无需微调即可高效选择医学基础模型,在OpenMind基准上加权Kendall指标相对提升约31%。

详情

展开后加载摘要…

URL PDF HTML 收藏