SageBwd: A Trainable Low-bit Attention
SageBwd: 一种可训练的低比特注意力
机构 * Tsinghua University(清华大学) ; UC Berkeley(伯克利大学)
AI总结 SageBwd通过降低比特数实现高效的注意力机制,通过理论分析和实验发现QK-范数和K-平滑对训练稳定性至关重要,从而在预训练中实现与全精度注意力的性能匹配。
高校专区
SageBwd: 一种可训练的低比特注意力
机构 * Tsinghua University(清华大学) ; UC Berkeley(伯克利大学)
AI总结 SageBwd通过降低比特数实现高效的注意力机制,通过理论分析和实验发现QK-范数和K-平滑对训练稳定性至关重要,从而在预训练中实现与全精度注意力的性能匹配。
MAP-Diff: 多锚点引导的扩散模型用于渐进式三维全身低剂量PET去噪
机构 * School of Engineering, Zurich University of Applied Sciences, CH ; Bioengineering Department ; Imperial-X, Imperial College London, UK ; DAMTP, University of Cambridge, UK ; Lucerne University Teaching ; Research Hospital, CH ; Lung Institute, Imperial College London, UK ; Cardiovascular Research Centre, Royal Brompton Hospital, UK ; School of Biomedical Engineering \& Imaging Sciences, King's College London, UK ; Yau Mathematical Sciences Center, Tsinghua University, CN
AI总结 MAP-Diff通过多锚点引导的扩散模型实现低剂量PET图像的渐进式去噪,提升PSNR和SSIM,降低NMAE,优于多种基线方法。
Comments 8 pages, 3 figures
学以致用:用于检测LLM生成文本的距离学习
机构 * Tsinghua University(清华大学) ; University of Birmingham(伯明翰大学) ; London School of Economics and Political Science(伦敦政治经济学院)
AI总结 本文提出了一种自适应学习的距离方法,用于更有效地检测LLM生成的文本,通过几何方法揭示了重写检测算法的原理,并在多种LLM上实现了显著的性能提升。
Comments Accepted by ICLR2026
StockBench: LLM代理在真实市场中能否盈利地进行股票交易?
机构 * Tsinghua University(清华大学) ; Beijing University of Posts and Telecommunications(北京邮电大学)
AI总结 STOCKBENCH评估LLM在真实股票市场中的交易能力,发现多数模型难以超越简单买入持有策略,部分模型展现出更高的收益和风险管理潜力。
论文未告诉你什么:为自动化论文复现恢复隐性知识
机构 * School of Software, Shandong University China ; Beijing Institute of Technology China ; Zhejiang University China ; Dept.\ of Math \& Statistics, Boston University USA ; College of AI, Tsinghua University China ; Alibaba Group China ; The Hong Kong Polytechnic University Hong Kong SAR, China ; School of Software, Shandong University ; Beijing Institute of Technology ; Zhejiang University ; Dept.\ of Math \& Statistics, Boston University ; College of AI, Tsinghua University ; Alibaba Group ; The Hong Kong Polytechnic University
AI总结 本文提出一种基于图的智能体框架,通过恢复关系型、身体型和集体型隐性知识,提升自动化论文复现性能。
Comments 32 pages (+ appendix), 8 figures. Lehui Li and Ruining Wang contributed equally. Yongshun Gong is the corresponding author
DriveCombo:自主驾驶中组合交通规则推理的基准测试
机构 * Autolab, Westlake University(西lake大学自动化实验室) ; Li Auto Inc(Li Auto公司) ; Tsinghua University(清华大学) ; The University of Hong Kong(香港大学)
AI总结 DriveCombo提出了一种基于文本和视觉的基准,用于评估自动驾驶中复杂交通规则的推理能力,通过五级认知阶梯和Rule2Scene代理提升模型在多规则冲突场景下的性能。
TQCodec:面向高质量音乐流媒体的神经音频编解码器
机构 * Tencent Music Entertainment(腾讯音乐娱乐) ; The Chinese University of Hong Kong(香港中文大学) ; Southeast University(东南大学) ; Tsinghua University(清华大学)
AI总结 TQCodec是一种针对高质量音乐流媒体设计的神经音频编解码器,通过改进的网络架构和感知驱动的比特分配策略,在高比特率下实现卓越的音频质量。
探究基于扩散变压器的文本到音频生成中的群体相对策略优化
机构 * Tsinghua University(清华大学) ; Microsoft(微软)
AI总结 本文基于扩散变压器架构,利用群体相对策略优化提升文本到音频生成的保真度和提示遵循性。
通过FSM驱动的流式推理管道提升AI可靠性:一个工业案例
机构 * School of Software, BNRist, Tsinghua University, China(软件学院、BNRist、清华大学、中国) ; Tianyi Technology Co., Ltd, China(天翼科技有限公司、中国)
AI总结 本文提出一种基于FSM驱动的流式推理管道,通过整合先验知识提升AI在工业场景中的鲁棒性和预测准确性。
Comments Preprint. The work was done in 2024
FATE: 闭环可行性感知任务生成与主动修复的物理 grounded 机器人课程
机构 * School of Aerospace Engineering, Tsinghua University, Beijing, China(航空航天工程系,清华大学,北京,中国) ; Department of Automation, Tsinghua University, Beijing, China(自动化系,清华大学,北京,中国) ; School of Integrated Circuits, Tsinghua University, Beijing, China(集成电路学院,清华大学,北京,中国) ; State Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence (BIGAI), Beijing, China(通用人工智能国家重点实验室,北京通用人工智能研究院(BIGAI),北京,中国)
AI总结 FATE通过闭环验证与主动修复机制,生成物理 grounded 的机器人任务课程,有效减少执行失败率。
Comments 16 Pages, 4 Figures
VidDoS:针对基于视频的大语言模型的通用拒绝服务攻击
机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; The Chinese University of Hong Kong, Hong Kong SAR(香港中文大学) ; Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳研究院) ; Shenzhen Loop Area Institute, China(深圳环园院) ; McGill University(麦吉尔大学) ; The Hong Kong University of Science and Technology, Guangzhou(香港科学与技术大学)
AI总结 VidDoS是一种针对视频大语言模型的通用拒绝服务攻击方法,通过通用优化生成实例无关的触发器,导致模型推理延迟和token扩展显著增加,引发安全问题。
第一人称共驾:面向辅助第一人称AI的网页原生智能眼镜代理
机构 * Shenzhen International Graduate School Tsinghua University Shenzhen China(深圳国际研究生院清华大学深圳中国) ; Independent Researcher London United Kingdom(独立研究者伦敦英国) ; Queen Mary University of London London United Kingdom(女王玛丽大学伦敦英国) ; Imperial College London London United Kingdom(帝国理工学院伦敦英国) ; University Of Surrey Guildford United Kingdom(Surrey大学Guildford英国)
AI总结 Egocentric Co-Pilot通过网页原生智能眼镜代理实现第一人称AI的持续辅助,结合神经符号框架和多模态意图层,展示了在日常生活中提升可及性和情境感知的实用路径。
Comments 14 pages, 6 figures, WWW 2026
通过从失败中显式学习解锁VLA在自动驾驶中的潜力
机构 * Tsinghua University(清华大学) ; University of Macau(澳门大学) ; Beijing Jiaotong University(北京交通大学)
AI总结 本文提出ELF-VLA框架,通过显式学习失败来提升自动驾驶中VLA模型的性能,实现关键场景的解决并取得SOTA效果。
张量输出的组合带状贝叶斯优化
机构 * Department of Industrial Engineering, Tsinghua University, Beijing 100084, China(清华大学工业工程系) ; College of Economics and Management, Nanjing University of Aeronautics and Astronautics, Nanjing 211106, China(南京航空航天大学经济管理学院)
AI总结 本文提出了一种针对张量输出的组合带状贝叶斯优化方法,通过引入张量输出高斯过程和UCB获取函数,有效处理部分观测的张量输出问题,并建立了理论遗憾界。
LSPRAG: 语言无关的实时单元测试生成中的LSP引导RAG
机构 * Tsinghua University(清华大学) ; East China Normal University(华东师范大学) ; Tencent(腾讯)
AI总结 LSPRAG通过利用语言服务器协议实现实时语言无关单元测试生成,显著提高了测试覆盖率。
Comments 13pages, 6 figures
跨视角视觉:评估视觉-语言模型在机器人场景中的空间推理能力
机构 * Tsinghua University(清华大学) ; Peking University(北京大学) ; Fudan University(复旦大学) ; Microsoft Research Asia(微软亚洲研究院) ; Hong Kong University of Science and Technology(香港科技大学) ; Zhejiang University(浙江大学)
AI总结 本文提出MV-RoboBench基准,评估视觉-语言模型在机器人场景中的多视角空间推理能力,揭示其在多视角机器人感知中的挑战。
Comments Accepted to ICLR 2026. Camera-ready version. Project page: https://aaronfengzy.github.io/MV-RoboBench-Webpage/
Journal ref International Conference on Learning Representations (ICLR), 2026
Fly-CL: 一种受飞虫嗅觉电路启发的框架,用于提升预训练模型连续表示学习中的高效去相关性和减少训练时间
机构 * Department of Automation, Tsinghua University(清华大学自动化系) ; Academy of Medical Engineering and Translational Medicine, Tianjin University(天津大学医学工程与转化医学学院)
AI总结 Fly-CL通过生物启发设计提升预训练模型连续学习效率,减少训练时间并保持高性能。
Comments ICLR 2026 accepted paper
无变分自编码器的潜在扩散模型
机构 * Department of Automation, Tsinghua University(自动化系,清华大学) ; Kling Team, Kuaishou Technology(快手科技 Kling 团队)
AI总结 SVG提出了一种无需变分自编码器的潜在扩散模型,通过自监督表示提升视觉生成的效率和质量。
Comments Accepted by ICLR 2026
DA-MMP:学习协调且准确的投掷动作:动态感知的运动流形原语
机构 * Shanghai Qi Zhi Institute(上海启智研究院) ; Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院)
AI总结 DA-MMP通过动态感知的运动流形原语,实现了高协调性投掷动作的生成,优于人类专家并能推广至新目标。
Comments Accepted to ICRA 2026. Project page: https://cc299792458.github.io/da-mmp/
VLA-Reasoner: 通过在线蒙特卡洛树搜索增强视觉-语言-动作模型的推理能力
机构 * School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电子与电气工程学院) ; Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; School of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications(北京邮电大学智能工程与自动化学院)
AI总结 VLA-Reasoner通过在线蒙特卡洛树搜索增强视觉-语言-动作模型,提升长时间轨迹任务的推理能力与执行效率。
Comments 8 pages, 6 figures, Accepted by ICRA 2026
PMark: 向鲁棒且无失真语义级水印技术迈进:在通道约束下
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; The Hong Kong University of Science and Technology(香港科技大学) ; Peking University(北京大学) ; National University of Singapore(新加坡国立大学) ; Tsinghua University(清华大学)
AI总结 PMark通过代理函数框架实现鲁棒且无失真的语义级水印技术,提升对改写攻击的鲁棒性并优化采样效率。
Comments ICLR 2026 Poster
OJBench: 一个面向大语言模型的竞赛级代码基准
机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ; Tsinghua University(清华大学) ; University of Chinese Academy of Sciences(中国科学院大学) ; Peking University(北京大学) ; Moonshot AI
AI总结 OJBench是一个用于评估大语言模型竞赛级代码推理能力的基准,通过232个编程竞赛问题揭示了现有模型在复杂推理任务中的局限性。
Comments 9 pages, 5 figures
CityLens:评估大型视觉-语言模型用于城市社会经济感知
机构 * Information Hub, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)信息中心) ; Department of Electronic Engineering, BNRist, Tsinghua University(清华大学电子工程系) ; School of Electronic and Information Engineering, Beijing Jiaotong University(北京交通大学电子与信息工程学院)
AI总结 CityLens通过评估大型视觉-语言模型在城市社会经济指标预测中的表现,揭示了其在可持续城市发展中的潜力与局限性。
Comments Accepted by ICLR 2026
AReaL: 一种用于语言推理的大规模异步强化学习系统
机构 * IIIS, Tsinghua University(清华信息学院) ; Ant Group(蚂蚁集团) ; HKUST(香港科技大学)
AI总结 AReaL通过异步机制提升大语言模型推理训练效率,实现GPU利用率显著提升
OneTwoVLA:一种具有自适应推理能力的统一视觉-语言-动作模型
机构 * Tsinghua University(清华大学) ; Shanghai Qi Zhi Institute(上海启智研究院)
AI总结 OneTwoVLA通过自适应推理机制实现视觉-语言-动作统一,提升机器人在复杂任务中的规划、交互与执行能力。
MoMa:一种用于材料属性预测的模块化深度学习框架
机构 * Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究院(AIR),清华大学) ; Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学) ; School of Artificial Intelligence, Nanjing University(人工智能学院,南京大学) ; Department of Chemical Engineering, Tsinghua University(化学工程系,清华大学)
AI总结 MoMa提出了一种模块化深度学习框架,通过训练专门模块并自适应组合以提升材料属性预测的性能,实验表明其在多个数据集上表现优异。
Comments Accepted to ICLR 2026
大语言模型中上下文长度扩展的内在熵
机构 * Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院) ; Carnegie Mellon University(卡内基梅隆大学) ; CPHOS ; University of Washington(华盛顿大学)
AI总结 本研究提出'内在熵'理论,通过实验验证长上下文对语言模型的影响,揭示训练数据集大小与最佳上下文长度的关系。
Comments 36 pages, 18 figures, 2 tables
Thoth:中期训练使LLM具备时间序列理解能力
机构 * School of Software, BNRist, Tsinghua University(软件学院、BNRist、清华大学)
AI总结 Thoth通过中期训练和Book-of-Thoth语料库,使LLM具备时间序列理解能力,并在多个基准测试中表现优异。
UD-SfPNet: 一种用于水下散射消除的偏振形状从偏振网络用于3D法线重建
机构 * School of Mechanical Engineering and Automation, Fuzhou University(福州大学机械与自动化学院) ; The Research Institute of Highway, Ministry of Transport(交通部高速公路研究 institute) ; The State Key Laboratory of Precision Measurement Technology and Instruments, Department of Precision Instruments, Tsinghua University(清华大学精密测量技术与仪器国家重点实验室)
AI总结 UD-SfPNet通过结合偏振成像的散射消除与形状从偏振技术,实现了水下3D法线重建的高精度与高效性。
神经潜在任意拉格朗日-欧拉网格用于流固耦合
机构 * School of Computer Science, Peking University(北京大学计算机科学学院) ; Global Innovation Exchange, Tsinghua University(清华大学全球创新交流中心) ; School of Electrical and Computer Science, University of Southampton(南安普顿大学电子与计算机科学学院)
AI总结 Fisale通过多尺度潜在ALE网格和分区耦合模块,实现复杂双向流固耦合问题的高效建模与学习。
Comments Proceedings of the 14th International Conference on Learning Representations