BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
BagelVLA: 通过交错视觉-语言-动作生成增强长时程操作
机构 * Tsinghua University(清华大学)
AI总结 BagelVLA通过整合语言规划、视觉预测和动作生成,提升复杂长时程操作任务的性能。
高校专区
BagelVLA: 通过交错视觉-语言-动作生成增强长时程操作
机构 * Tsinghua University(清华大学)
AI总结 BagelVLA通过整合语言规划、视觉预测和动作生成,提升复杂长时程操作任务的性能。
WorldArena: 一个用于评估具身世界模型感知与功能效用的统一基准
机构 * Tsinghua University, Beijing, China(清华大学) ; Shanghai Jiao Tong University, Shanghai, China(上海交通大学) ; The University of Hong Kong, Hong Kong SAR, China(香港大学) ; Princeton University, Princeton, NJ, USA(普林斯顿大学) ; Chinese Academy of Sciences, Beijing, China(中国科学院) ; University of Science(科学技术大学) ; Peking University, Beijing, China(北京大学) ; National University of Singapore, Singapore(新加坡国立大学)
AI总结 WorldArena是一个统一的基准,用于评估具身世界模型的感知和功能效用,揭示高视觉质量与强具身任务能力之间的差距。
针对多模态大语言模型的语音音频组合攻击及其通过SALMONN-Guard的缓解
机构 * Tsinghua University(清华大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; University of Cambridge(剑桥大学)
AI总结 针对多模态LLMs的安全性,提出SACRED-Bench评估语音音频组合攻击,并通过SALMONN-Guard将攻击成功率降低至20%。
DiffuTester: 通过挖掘结构模式加速扩散大语言模型的单元测试生成
机构 * College of AI, Tsinghua University(清华大学人工智能学院) ; School of Software, Beihang University(北航软件学院) ; School of Computer Science and Engineering, Beihang University(北航计算机科学与工程学院)
AI总结 DiffuTester通过挖掘结构模式提升扩散大语言模型在单元测试生成中的效率,实现高质量测试用例生成。
Comments Update format and add some experimental results
PhysUniBench: 一个面向本科生层面的多模态物理推理基准测试
机构 * The University of Sydney(悉尼大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) ; Fudan University(复旦大学) ; The Chinese University of Hong Kong(香港中文大学) ; Michigan State University(密歇根州立大学) ; Tsinghua University(清华大学) ; Beihang University(北航大学)
AI总结 PhysUniBench是一个面向本科生层面的多模态物理推理基准测试,旨在评估和提升多模态大语言模型在物理问题上的推理能力。
从无害到有害:通过对抗隐喻 jailbreak 语言模型
机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; People’s Public Security University of China(中国人民公安大学) ; Tsinghua University(清华大学)
AI总结 通过对抗隐喻诱导语言模型生成有害内容,提出AVATAR框架以提升jailbreak攻击效果。
Comments arXiv admin note: substantial text overlap with arXiv:2412.12145
DexImit:从单目人类视频中学习双臂灵巧操作
机构 * Shanghai AI Laboratory(上海人工智能实验室) ; Tsinghua University(清华大学) ; The Chinese University of Hong Kong(香港中文大学) ; NVIDIA(英伟达)
AI总结 DexImit通过四阶段生成流程,将单目人类视频转化为物理合理的机器人数据,以解决双臂灵巧操作的泛化问题。
UniVTAC:一种用于视觉-触觉操作数据生成、学习和评估的统一仿真平台
机构 * ScaleLab, Shanghai Jiao Tong University(上海交通大学ScaleLab) ; D-Robotics ; ViTai Robotics ; The University of Hong Kong(香港大学) ; Nanjing University(南京大学) ; Shenzhen University(深圳大学) ; Wuhan University(武汉大学) ; Fudan University(复旦大学) ; Tsinghua University(清华大学)
AI总结 UniVTAC提出了一种统一的视觉-触觉数据生成平台,通过训练触觉中心的编码器和基准任务提升机器人操作的成功率。
Comments Website: https://univtac.github.io/
RoboInter: 一种面向机器人操控的综合中间表示套件
机构 * University of Science and Technology of China(中国科学技术大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Beihang University(北航) ; Nanyang Technological University(南洋理工大学) ; Zhejiang University(浙江大学) ; Tsinghua University(清华大学) ; The Chinese University of Hong Kong(香港中文大学)
AI总结 RoboInter通过统一的中间表示套件,提升机器人操控中视觉-语言-动作系统的泛化能力和推理能力。
Comments Published to ICLR 2026, 69 pages, 40 figures
利用局部物理信息瓶颈训练深度物理神经网络
机构 * Department of Precision Instrument, Tsinghua University(精密仪器系,清华大学) ; Laboratoire Kastler Brossel, École Normale Supérieure - Paris Sciences et Lettres (PSL) Research University, Sorbonne Université, Centre National de la Recherche Scientifique (CNRS), UMR 8552, Collège de France(Kastler Brossel实验室,巴黎科学与文学研究大学(PSL)研究大学,索邦大学,法国国家科学研究中心(CNRS),UMR 8552,法国学院) ; School of Integrated Circuits, Beijing Advanced Innovation Center for Integrated Circuits, BNRist, Tsinghua University(集成电路学院,北京集成电路先进创新中心,BNRist,清华大学) ; Department of Electrical and Electronic Engineering, The University of Hong Kong(电子与电气工程系,香港大学) ; State Key Laboratory of Precision Space-time Information Sensing Technology, Department of Precision Instrument, Tsinghua University(精密时空信息感知技术国家重点实验室,精密仪器系,清华大学) ; Key Laboratory of Photonic Control Technology, Ministry of Education, Tsinghua University(光电控制技术重点实验室,教育部,清华大学) ; Department of Electrical Engineering, City University of Hong Kong(电气工程系,城市大学)
AI总结 本文提出物理信息瓶颈框架,通过整合信息论与局部学习,实现深度物理神经网络在任意物理动态下的高效训练,适用于多种物理基质。
Comments 9 pages, 4 figures
通过相关性感知支付增强仿射最大化拍卖
机构 * CFCS, School of Computer Science, Peking University(计算机科学系,北京大学) ; Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学)
AI总结 本文提出CA-AMA框架,通过相关性感知支付增强AMAs,解决估值相关分布下的收益优化问题。
Comments 22 pages. Work in progress
敏捷非对称多足运动:通过几何力学和自旋模型对偶性进行接触规划
机构 * Department of Mechanical Engineering, Pennsylvania State University(宾夕法尼亚州立大学机械工程系) ; MIT Improbable AI Lab, Maschusetts Institute of Technology(麻省理工学院Improbable AI实验室) ; Department of Physics, Georgia Institute of Technology(佐治亚理工学院物理系) ; Department of Physics, Tsinghua University(清华大学物理系)
AI总结 本研究提出了一种基于几何力学和自旋模型对偶性的原理性框架,用于发现多足运动中的新控制结构,实现了六足机器人非对称运动策略,使前进速度提高50%。
Comments 10 pages, 7 figures. arXiv admin note: text overlap with arXiv:2302.03019
GEBench: 图形用户界面生成模型的基准测试
机构 * StepFun ; South China University of Technology(南方科技大学) ; Peking University(北京大学) ; Tsinghua University(清华大学) ; Institute of Automation Chinese Academy of Sciences(中国科学院自动化研究所) ; The University of Chicago(芝加哥大学) ; Nanyang Technological University(南洋理工大学)
AI总结 GEBench提出一个评估GUI生成模型动态交互和时间一致性的基准测试,通过五维指标揭示模型在多步交互中的瓶颈问题。
Comments 23 pages, 5 figures, 4 tables
基于多个教师的贝叶斯语音合成器可以学习
机构 * Tsinghua University(清华大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
AI总结 BELLE通过贝叶斯推断提升语音合成的不确定性建模,以更少的数据实现更优的语音生成性能。
Comments Code is available at https://github.com/OpenTSLab/BELLE
VDRive:利用强化VLA和扩散策略实现端到端自动驾驶
机构 * Suzhou Automotive Research Institute of Tsinghua University(清华大学苏州汽车研究院)
AI总结 VDRive通过结合强化VLA和扩散策略,实现端到端自动驾驶,提升决策的可解释性和鲁棒性。
Comments WIP
MILR:通过测试时潜在推理改进多模态图像生成
机构 * University of Science and Technology of China(中国科学技术大学) ; State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室) ; Peking University(北京大学) ; Tsinghua University(清华大学) ; University of California, Los Angeles(加州大学洛杉矶分校)
AI总结 MILR通过测试时潜在推理提升多模态图像生成性能,实现跨模态推理和统一潜在空间优化。
Comments 21 pages,14 figures,9 tables
MoWM: 通过潜在到像素特征调制的混合世界模型实现具身规划
机构 * Tsinghua University(清华大学) ; Manifold AI ; Shanghai Jiao Tong University(上海交通大学)
AI总结 MoWM通过融合潜在世界模型与像素特征,提升具身规划中动作解码的精度与泛化能力。
关于并行文本生成的综述:从并行解码到扩散语言模型
机构 * Peking University(北京大学) ; University of Illinois Chicago(伊利诺伊大学香槟分校) ; Tsinghua University(清华大学) ; XPENG ; Alibaba Group(阿里巴巴集团) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
AI总结 本文综述了并行文本生成技术,分析了基于AR和非AR的方法,评估了其在速度、质量和效率上的权衡,并指出了未来研究方向。
多专家学习框架与状态空间模型用于光学和SAR图像配准
机构 * Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education of China(教育部智能感知与图像理解重点实验室) ; Hangzhou Institute of Technology, Xidian University(西安电子科技大学杭州学院) ; School of Computer Science, Xidian University(西安电子科技大学计算机学院) ; Department of Automation, Tsinghua University(清华大学自动化系)
AI总结 本文提出ME-SSM框架,通过多专家学习和状态空间模型提升光学与SAR图像配准的精度与效率。
DISPROTBENCH: 揭示蛋白质结构预测模型在内源性无序区域的功能限制
机构 * Institute for Clarity in Documentation(清晰文档研究所) ; Inria Paris-Rocquencourt(巴黎-罗克琴堡研究所) ; Rajiv Gandhi University(拉贾·甘地大学) ; Tsinghua University(清华大学) ; Palmer Research Laboratories(帕勒研究实验室)
AI总结 DISPROTBENCH通过引入功能不确定性敏感度度量,揭示蛋白质结构预测模型在内源性无序区域的功能限制及预测不确定性对下游任务的影响。
从人类反馈中鲁棒强化学习用于大语言模型微调
机构 * Department of Statistics, LSE(统计系,伦敦经济学院) ; Department of Mathematics, Tsinghua University(数学系,清华大学) ; School of Mathematics, University of Birmingham(数学学院,伯明翰大学) ; Department of Engineering Science, University of Oxford(工程科学系,牛津大学)
AI总结 本文提出了一种鲁棒的强化学习算法,用于改进大语言模型微调中从人类反馈学习奖励函数的性能,通过减少方差和改进后悔界,实验证明其在基准数据集上表现优异。
数据科学与技术迈向通用人工智能 Part I:分层数据管理
机构 * Tsinghua University(清华大学) ; ModelBest Inc.(ModelBest公司) ; Beijing Institute of Technology(北京理工大学) ; South China Agricultural University(华南农业大学)
AI总结 本文提出分层数据管理框架,通过数据与模型的协同进化提升大语言模型训练效率和性能。
Comments 16 pages, 3 figures, 7 tables
在离散潜在空间中进行下一步概念预测可使语言模型更强大
机构 * LUMIA Lab(LUMIA实验室) ; School of Artificial Intelligence(人工智能学院) ; Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Department of Electronic Engineering(电子工程系) ; College of Artificial Intelligence(人工智能学院) ; Tsinghua University(清华大学) ; Shanghai Innovation Institute(上海创新研究院)
AI总结 通过在离散潜在空间中进行下一步概念预测,ConceptLM在预训练任务中实现了更强大的语言模型性能。
FlattenGPT: Transformer中基于层扁平化的深度压缩
机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学) ; Tsinghua University(清华大学)
AI总结 FlattenGPT通过层扁平化技术实现Transformer模型的深度压缩,有效提升效率并保持性能,适用于多种模型类型和参数规模。
Comments Submitted to ICML 2026
WildReward: 从真实世界的人类交互中学习奖励模型
机构 * Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学)
AI总结 本文提出WildReward,通过真实世界用户交互直接学习奖励模型,无需偏好对,实现更优的校准和一致性,并在多种任务中取得显著提升。
注意间隙:通过意图-执行不匹配学习隐式阻抗
机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) ; Qiuzhen College, Tsinghua University(清华大学启晨学院) ; School of Mechanical and Aerospace Engineering, Jilin University(吉林大学机械与 aerospace 工程学院) ; EFORT Intelligent Robot Co., Ltd.(EFORT 智能机器人有限公司)
AI总结 通过意图-执行不匹配学习隐式阻抗,实现低成本硬件的鲁棒力感知与动态补偿。
Comments 14 pages, 9 figures, 5 tables
TFMLinker:通过基于图的上下文学习进行通用链接预测的通用链接预测器
机构 * Nankai University(南开大学) ; Tsinghua University(清华大学) ; Beihang University(北航)
AI总结 TFMLinker通过表格基础模型的上下文学习能力,在不同图上进行通用链接预测,无需数据集特定微调。
超越功能连接:基于fMRI的脑部疾病分类的时间序列建模
机构 * PolyU ; South China University of Technology(南方科技大学) ; Tsinghua University(清华大学) ; YMSC, Tsinghua University(清华大学交叉信息学院) ; Dept. of Biomedical Engineering, PolyU(PolyU 生物医学工程系) ; Sch. of Future Technology, South China University of Technology(南方科技大学未来技术学院) ; Dept. of Health Technology and Informatics, PolyU(PolyU 健康技术与信息学系) ; Dept. of Biomedical Engineering and the Dept. of Data Science and Artificial Intelligence, PolyU(PolyU 生物医学工程系和数据科学与人工智能系)
AI总结 本文提出DeCI框架,通过周期和漂移分解及通道独立性建模,提升fMRI脑部疾病分类的准确性和泛化能力。
Comments This paper has been accepted by IEEE Transactions on Medical Imaging
G-LNS:基于生成式大邻域搜索的LLM自动启发式设计
机构 * Software College, Northeastern University(东北大学软件学院) ; International Centre for Theoretical Physics Asia-Pacific, University of Chinese Academy of Sciences(中国科学院大学国际理论物理亚太中心) ; Taiji Laboratory for Gravitational Wave Universe, University of Chinese Academy of Sciences(中国科学院大学太极引力波宇宙实验室) ; Tsinghua University(清华大学)
AI总结 G-LNS通过生成式进化框架,结合LLM共同进化破坏与修复操作符,实现了大邻域搜索操作符的自动化设计,在复杂组合优化问题中表现出更强的性能和泛化能力。
基于博弈论的LLM启发式发现共演化
机构 * C2DL, Institute of Automation, Chinese Academy of Sciences(C2DL,自动化研究所,中国科学院) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) ; Tsinghua University(清华大学)
AI总结 本文提出ASRO框架,通过博弈论方法实现求解器与实例生成器的共演化,提升启发式发现的泛化能力和鲁棒性。