Text2SQL-Flow: A Robust SQL-Aware Data Augmentation Framework for Text-to-SQL
Text2SQL-Flow:一种鲁棒的SQL感知数据增强框架用于文本到SQL
机构 * Peking University, Beijing, China(北京大学,北京,中国)
AI总结 Text2SQL-Flow提出了一种SQL感知的数据增强框架,通过生成高质量的文本到SQL数据集,提升大型语言模型在问题解决和检索任务中的性能。
高校专区
Text2SQL-Flow:一种鲁棒的SQL感知数据增强框架用于文本到SQL
机构 * Peking University, Beijing, China(北京大学,北京,中国)
AI总结 Text2SQL-Flow提出了一种SQL感知的数据增强框架,通过生成高质量的文本到SQL数据集,提升大型语言模型在问题解决和检索任务中的性能。
MILR:通过测试时潜在推理改进多模态图像生成
机构 * University of Science and Technology of China(中国科学技术大学) ; State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室) ; Peking University(北京大学) ; Tsinghua University(清华大学) ; University of California, Los Angeles(加州大学洛杉矶分校)
AI总结 MILR通过测试时潜在推理提升多模态图像生成性能,实现跨模态推理和统一潜在空间优化。
Comments 21 pages,14 figures,9 tables
关于并行文本生成的综述:从并行解码到扩散语言模型
机构 * Peking University(北京大学) ; University of Illinois Chicago(伊利诺伊大学香槟分校) ; Tsinghua University(清华大学) ; XPENG ; Alibaba Group(阿里巴巴集团) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
AI总结 本文综述了并行文本生成技术,分析了基于AR和非AR的方法,评估了其在速度、质量和效率上的权衡,并指出了未来研究方向。
选择性KV缓存共享以缓解LLM推理中的定时侧信道
机构 * University of Connecticut Storrs(康涅狄格大学斯托尔斯分校) ; Independent(独立研究者) ; Indiana University Bloomington(印第安纳大学布卢明顿分校) ; UC Santa Cruz(圣克拉拉大学) ; Peking University(北京大学)
AI总结 SafeKV通过选择性KV缓存共享缓解LLM推理中的定时侧信道问题,实现隐私保护与性能优化的平衡。
Comments 14 pages,15 figures
Cochain: 在LLM代理工作流中平衡不足与过度的合作
机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) ; Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学计算机科学与工程系) ; School of Urban Planning and Design, Peking University(北京大学城市规划与设计学院)
AI总结 Cochain通过整合知识图谱和提示树,有效解决业务工作流中LLM代理合作问题,优于基线模型。
Comments 35 pages, 23 figures
神经力场:基于少量样本的学习通用物理推理
机构 * Institute for AI, Peking University(人工智能研究院,北京大学) ; School of Psychological and Cognitive Sciences, Peking University(心理学与认知科学学院,北京大学) ; School of EECS, Peking University(电子工程学院,北京大学) ; School of Integrated Circuits, Peking University(集成电路学院,北京大学) ; State Key Lab of General AI, Peking University(通用人工智能国家重点实验室,北京大学) ; Beijing Key Laboratory of Behavior and Mental Health, Peking University(北京行为与心理健康重点实验室,北京大学)
AI总结 神经力场通过力场表示实现基于少量样本的学习通用物理推理,有效提升物理动态的泛化能力。
Comments 27 pages, ICLR 2026
高效且稳定的扩散语言模型强化学习
机构 * State Key Laboratory of Cognitive Intelligence, University of Science(认知智能国家重点实验室,科学大学) ; City University of Hong Kong, Hong Kong, China(香港城市大学,香港,中国) ; Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China(人工智能学院,中国人民大学,北京,中国) ; School of Software and Microelectronics, Peking University, Beijing, China(软件与微电子学院,北京大学,北京,中国)
AI总结 本文提出STP框架,通过空间和时间剪枝提升扩散语言模型强化学习的效率和稳定性。
Comments 13 pages, 3 figures
TiFRe: 基于文本的视频帧减少用于高效视频多模态大语言模型
机构 * Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学)
AI总结 TiFRe通过文本引导的帧采样和帧匹配机制,有效减少视频输入帧数,同时保留关键信息,提升视频多模态大语言模型的效率和性能。
FlattenGPT: Transformer中基于层扁平化的深度压缩
机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学) ; Tsinghua University(清华大学)
AI总结 FlattenGPT通过层扁平化技术实现Transformer模型的深度压缩,有效提升效率并保持性能,适用于多种模型类型和参数规模。
Comments Submitted to ICML 2026
面向时间知识图谱的更优进化建模
机构 * Xidian University China Xi'an ; Peking University China Peking ; University of Electronic Science ; XI'AN University of Posts\&Telecommunications China Xi'an ; Xidian University ; Peking University ; XI'AN University of Posts\&Telecommunications
AI总结 本文提出TKG进化基准测试,旨在解决现有评估中固有偏差和简化任务的问题,通过四个偏见校正数据集和两个新型任务提升对时间知识图谱进化建模的理解。
Comments 13 pages, 11 figures
OPE:通过大纲引导路径探索克服并行思维中的信息饱和
机构 * National Engineering Research Center for Software Engineering, Peking University, Beijing, China(软件工程国家工程研究中心,北京大学,北京,中国)
AI总结 OPE通过生成多样化的推理大纲,减少信息冗余,提升并行思维在复杂问题中的推理性能。
Tighnari v2: 通过专家混合和弱监督学习缓解多模态植物分布预测中的标签噪声和分布偏移
机构 * The University of Sydney, Sydney, New South Wales, Australia(悉尼大学) ; The University of New South Wales, Sydney, New South Wales, Australia(新南威尔士大学) ; School of Software and Microelectronics, Peking University, Beijing, China(北京大学软件与微电子学院) ; Shandong University of Technology, Zibo, Shandong, China(山东科技大学)
AI总结 Tighnari v2通过专家混合和弱监督学习缓解多模态植物分布预测中的标签噪声和分布偏移,提升预测性能。
新技能还是更尖锐的原语?从概率角度探讨RLVR中推理能力的出现
机构 * University of Science(科学技术大学) ; The Chinese University of Hong Kong(香港中文大学) ; Peking University(北京大学) ; Nanjing University(南京大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Shanghai Jiao Tong University(上海交通大学)
AI总结 RLVR通过提升原子步骤概率,使模型在多步骤推理中克服指数衰减,展现新能力而非单纯激发潜在痕迹。
Comments 15 pages
OLion:通过谱和ℓ∞隐偏见的交集接近Hadamard理想
机构 * Peking University(北京大学) ; Microsoft Research Asia(微软亚洲研究院)
AI总结 OLion通过结合谱控制与ℓ∞隐偏见,高效近似在Hadamard-like集上取最大步长,提升大规模语言和视觉任务的优化性能。
Comments 23 pages
同时触觉-视觉感知用于学习多模态机器人操作
机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) ; School of Psychological and Cognitive Sciences, Peking University(北京大学心理与认知科学学院) ; Beijing Key Lab of Behavior and Mental Health, Peking University(北京大学行为与心理健康北京市重点实验室) ; Beijing Institute for General Artificial Intelligence(北京一般人工智能研究院) ; State Key Lab for General Artificial Intelligence(一般人工智能国家重点实验室) ; Embodied Intelligence Lab, PKU-Wuhan Institute for Artificial Intelligence(具身智能实验室,北京大学武汉人工智能研究院) ; Department of Computer Science and Technology, University of Cambridge(剑桥大学计算机科学与技术系)
AI总结 TacThru-UMI通过结合同时触觉-视觉感知与现代学习框架,实现了高精度多模态机器人操作。
UniLiP: 适配CLIP以实现统一的多模态理解、生成与编辑
机构 * Center for Data Science, Peking University(北京大学数据科学中心) ; Alibaba Group(阿里巴巴集团) ; CASIA ; Center for Machine Learning Research, Peking University(北京大学机器学习研究中心) ; State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
AI总结 UniLIP通过两阶段训练和双条件架构,提升CLIP在多模态理解、生成和编辑任务中的性能,实现高效参数下的高表现。
通过DropTriple损失实现动与文本的跨模态检索
机构 * School of Artificial Intelligence, Chongqing University of Technology, China(重庆理工大学人工智能学院) ; College of Computer Science, Sichuan University, China(四川大学计算机学院) ; Key Laboratory of Machine Perception, Shenzhen Graduate School, Peking University, China(北京大学深圳研究生院机器感知重点实验室)
AI总结 本文提出DropTriple损失,用于提升人类动作与文本之间的跨模态检索性能,实验表明在HumanML3D数据集上实现了较高的检索准确率。
Comments This paper has been accepted by ACM MM Asia 2023 (Best Paper Candidate)
TextVQA挑战2021获胜团队Mia:基于预训练序列到序列模型的视觉-语言表示学习
机构 * SFE Deeplearning Platform Ping An Health Technology Beijing China(平安健康科技北京分公司) ; Peking University Beijing China(北京大学) ; Visual Computing Group Ping An Property & Casualty Insurance Company Shenzhen China(平安财产保险股份有限公司视觉计算组) ; Jinan University Guangzhou China(暨南大学) ; Central University of Finance and Economics Beijing China(中央财经大学) ; Ping An Health Cloud Company Limited Shenzhen China(平安健康云有限公司) ; Ping An International Smart City Technology Co Ltd Shenzhen China(平安国际智慧城市科技有限公司)
AI总结 TextVQA挑战2021中,Mia团队利用预训练序列到序列模型T5,通过融合多模态信息和优化预训练任务,提升视觉-语言推理能力。
Comments Winner of TextVQA 2021
弱到强:基于VLM的伪标签作为多模态视频隐藏情绪理解任务中的弱监督训练策略
机构 * The University of New South Wales, Sydney, New South Wales, Australia(新南威尔士大学) ; The University of Sydney, Sydney, New South Wales, Australia(悉尼大学) ; School of Software and Microelectronics, Peking University, Zibo, Shandong, China(北京大学软件与微电子学院) ; Shandong University of Technology, Beijing, China(山东理工大学)
AI总结 本文提出基于VLM的伪标签弱监督方法,用于多模态视频隐藏情绪识别,提升准确率至0.69并建立新基准。
V-ABFT:基于方差的自适应阈值用于混合精度深度学习中的容错矩阵乘法
机构 * Peking University(北京大学) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
AI总结 V-ABFT通过基于方差的自适应阈值算法,提升了混合精度深度学习中矩阵乘法的容错性能,实现更精确的误差检测和更低的阈值与误差比。
从超快运动模糊图像中恢复3D形状
机构 * Shandong University(山东大学) ; Nanjing University(南京大学) ; Peking University(北京大学)
AI总结 本文提出了一种高效的逆渲染方法,用于从超快运动模糊图像中恢复3D形状,通过优化计算瓶颈提升模拟速度和精度。
Comments Accepted by 3DV 2026. Project page: https://maxmilite.github.io/rec-from-ultrafast-blur/
AD-MIR:通过结构化推理弥合从感知到说服的广告视频理解鸿沟
机构 * Peking University, Beijing, China(北京大学) ; Tsinghua University, Beijing, China(清华大学) ; Xi'an Jiaotong University, Xi'an, China(西安交通大学) ; South China University of Technology, Guangzhou, China(华南理工大学)
AI总结 AD-MIR通过结构化推理框架,结合语义检索与精确关键词匹配,有效解码广告意图,提升广告视频理解的准确性与说服策略分析能力。
基于相似性的专家重新路由:用于MoE模型高效批量解码
机构 * Taobao & Tmall Group of Alibaba(淘宝与天猫集团) ; Shenzhen Graduate School, Peking University(北京大学深圳研究生院)
AI总结 SERE通过基于相似性的专家重新路由方法,提升MoE模型在批量解码中的效率,实现2倍速度提升并减少质量损失。
Comments Published as a conference paper at ICLR 2026
远程人类识别:挑战、方法与HID 2025竞赛结果
机构 * Shenzhen Polytechnic University(深圳职业技术大学) ; Fujitsu Research and Development Center Company Ltd.(富士通研发有限公司) ; Beijing Jiaotong University(北京交通大学) ; South China University of Technology(华南理工大学) ; Shandong University(山东大学) ; EVERSPRY ; Wuhan University of Technology(武汉理工大学) ; Sun Yat-sen University(中山大学) ; Shanghai Jiao Tong University(上海交通大学) ; Fujitsu Ltd.(富士通株式会社) ; Huazhong University of Science and Technology(华中科技大学) ; Hefei University of Technology(合肥工业大学) ; Peking University, Shenzhen Graduate School(北京大学深圳研究生院) ; Beijing Normal University(北京师范大学) ; Watrix Technology Limited Co. Ltd.(华智科技有限公司) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; University of Cordoba(科尔多瓦大学) ; University of East London(东伦敦大学) ; Southern University of Science and Technology(南方科技大学)
AI总结 HID 2025竞赛评估了步态识别在远程人类识别中的挑战与进展,通过SUSTech-Competition数据集验证了算法性能的提升,达到94.2%的准确率。
Comments Accepted by IJCB 2025(https://ijcb2025.ieee-biometrics.org/competitions/)
基于帕累托最优的流水线:在移动MOBA游戏中蒸馏轻量级AI代理
机构 * School of Computer Science Peking University Beijing China(计算机科学系 首都大学 北京 中国) ; TiMi L1 Studio Tencent Chengdu China(TiMi L1工作室 腾讯 成都 中国) ; Peking University(首都大学) ; Tencent(腾讯)
AI总结 本文提出基于帕累托最优的流水线,设计高效学生架构搜索空间,在移动MOBA游戏中实现轻量级AI代理的蒸馏,提升推断速度与能效。
非参数贝叶斯优化用于一般奖励
机构 * Guanghua School of Management, Peking University(北京大学光华管理学院)
AI总结 本文提出一种非参数贝叶斯优化算法,通过无限高斯过程与汤普森采样结合,实现一般奖励下的无遗憾保证,适用于非平稳和病态奖励场景。
前瞻性验证:用于扩散大语言模型在无上下文自由文法下的可靠约束解码
机构 * College of AI, Tsinghua University(清华大学人工智能学院) ; School of Computer Science, Peking University(北京大学计算机学院) ; School of Computer Science and Engineering, Beihang University(北航计算机科学与工程学院) ; School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院) ; School of Computing, National University of Singapore(新加坡国立大学计算机学院)
AI总结 LAVE通过并行预测标记分布实现可靠的约束解码,提升扩散大语言模型在上下文无关文法下的生成准确性。
RealPDEBench: 一个整合真实世界数据的复杂物理系统基准
机构 * School of Engineering, Westlake University(西湖大学工程学院) ; Global College, Shanghai Jiao Tong University(上海交通大学全球学院) ; Academy of Mathematics and Systems Science, Chinese Academy of Sciences(中国科学院数学与系统科学研究院) ; Department of Geotechnical Engineering, Tongji University(同济大学地质工程系) ; School of Physics, Peking University(北京大学物理学院) ; Key Laboratory for Power Machinery and Engineering of M. O. E., Shanghai Jiao Tong University(上海交通大学机械工程重点实验室)
AI总结 RealPDEBench通过整合真实世界数据与模拟数据,为科学机器学习提供了一个新的基准,旨在弥合仿真与现实之间的差距。
Comments iclr26 oral; 46 pages, 21 figures
通过复制粘贴缓解大语言模型的幻觉
机构 * Department of Computer Science, Tianjin University of Technology(天津理工大学计算机学院) ; National Institute of Health Data Science, Peking University(北京大学国家健康数据科学研究院) ; Tencent Jarvis Lab(腾讯Jarvis实验室)
AI总结 通过复制粘贴方法提升大语言模型上下文忠实性,减少幻觉并提高测试性能。
Comments Accepted to ICLR 2026
重新思考可解释疾病预测:通过反思认知架构协同提升准确性和可靠性
机构 * Institute for Artificial Intelligence, Peking University ; School of Software \& Microelectronics, Peking University ; School of Computer Science, Peking University ; School of Life Sciences, Peking University ; Department of Comprehensive Oncology, National Cancer Center/National Clinical Research Center for Cancer/Cancer Hospital, Chinese Academy of Medical Sciences ; Peking Union Medical College Beijing, China
AI总结 本文提出反思认知架构(RCA),通过经验与反思直接从表格数据学习,实现预测精度与解释可靠性的协同提升。
Comments under review