arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of California, Berkeley(加州大学伯克利分校)

共收录 1884
2606.06036 2026-06-05 cs.AI cs.IR

Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents

记忆是重建的,而非检索的:面向LLM智能体的图记忆

Shuo Ji, Yibo Li, Bryan Hooi

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学)

AI总结 提出MRAgent框架,通过关联记忆图和主动重建机制,使LLM智能体在推理过程中动态调整记忆访问,显著提升长程记忆推理性能。

Comments Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05873 2026-06-05 cs.RO cs.AI cs.CV cs.LG

LadderMan: Learning Humanoid Perceptive Ladder Climbing

LadderMan: 学习人形机器人感知爬梯

Siheng Zhao, Yuanhang Zhang, Ziqi Lu, Pieter Abbeel, Rocky Duan, Koushil Sreenath, Yue Wang, C. Karen Liu, Guanya Shi

机构 * Amazon FAR(亚马逊FAR) USC(美国南加州大学) UC Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学) CMU(卡内基梅隆大学)

AI总结 提出LadderMan系统,通过两阶段学习管道和视觉基础模型,使人形机器人能够鲁棒地攀爬多种梯子并在梯子上进行操控。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05857 2026-06-05 cs.CL

Forgive or forget: Understanding the context of hate in audio retrieval systems

原谅或忘记:理解音频检索系统中仇恨的上下文

Arghya Pal, Sailaja Rajanala, Raphael C. -W. Phan, Shekhar Nayak

机构 * University of California, Berkeley(加州大学伯克利分校)

AI总结 提出一种后门因果去偏框架,通过情感控制中介在保持语义相关性的同时抑制有害语音,实验表明在最小化检索精度损失下持续降低毒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05818 2026-06-05 math.HO cs.AI math.AG math.CO math.RT

Benchmarks in Leipzig

莱比锡基准测试

Andrei Balakin, Miklós Bóna, Marie-Charlotte Brandenburg, Clara Briand, Veronica Calvo Cortes, Shelby Cox, Jesus A. De Loera, Danai Deligeorgaki, Hannah Friedman, Tim Gehrunger, Chiara Giardino, Stephen Griffeth, Baran Hashemi, Elena Hoster, Alexander Ivanov, Nupur Jain, Aryaman Jal, Leonie Kayser, Joris Koefler, Kevin Kühn, Mario Kummer, Felix Lotter, René Marczinzik, Victor S. Miller, Alejandro Morales, Greta Panova, Gianni Petrella, Nathan Pflueger, Lakshmi Ramesh, Nikolas Rieke, Carlos Rodriguez, Andrea Rosana, Flavio Salizzoni, Otto T. P. Schmidt, Sven Ulf Schmitz, Lina Maria Simbaqueba Marin, Luca Sodomaco, Christian Stump, Bernd Sturmfels, Alexander Taveira Blomenhofer, Simon Telen, Philipp Tuchel, Emil Verkama, Carl Felix Waller, Julian Weigert, Annette Werner, Nathan Williams, Claudius Zibrowius

机构 * University of California, Berkeley(加州大学伯克利分校)

AI总结 49位数学家于2026年4月至5月编制了100个研究级数学问题数据集,通过多阶段评估大型语言模型的数学推理能力,最终仅剩2个问题未解决。

Comments 8 pages including 8 benchmark statistics tables + 20 pages appendix containing the 100 Leipzig Benchmark questions

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05760 2026-06-05 cs.CV

ExpSpeech-Net: Multimodal Fusion of Expression and Speech for Deepfake Detection

ExpSpeech-Net: 表情与语音的多模态融合用于深度伪造检测

Ruchika Sharma, Rudresh Dwivedi

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学)

AI总结 提出轻量级ExpSpeech-Net模型,通过融合面部表情和语音模式,利用SqueezeNet和RNN骨干网络及智能特征选择,实现高效深度伪造检测,准确率达94.5%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05729 2026-06-05 cs.IT cs.LG math.IT

Automated Proving of Shannon-Type Entropy Inequalities via Fine-Tuned Language Models and Guided Tree Search

通过微调语言模型和引导树搜索自动证明香农型熵不等式

Shing Yin Wong, Shaocheng Liu, Linqi Song, Amin Gohari, Cheuk Ting Li

机构 * University of California, Berkeley(加州大学伯克利分校)

AI总结 本文通过微调小规模语言模型并结合引导束搜索,自动化证明香农型熵不等式,在含10-15个变量的测试集上达到85%的证明成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05661 2026-06-05 cs.AI cs.CL

Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments

持续学习基准:评估现实世界有状态环境中的前沿AI系统

Parth Asawa, Christopher M. Glaze, Gabriel Orlanski, Ramya Ramakrishnan, Benji Xu, Asim Biswal, Vincent Sunn Chen, Frederic Sala, Matei Zaharia, Joseph E. Gonzalez

机构 * UC Berkeley(伯克利大学) Snorkel AI University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

AI总结 提出首个专家验证的持续学习基准CL-Bench,涵盖六个领域,通过增益指标隔离在线学习能力,发现现有系统存在过拟合和知识复用不足问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05658 2026-06-05 cs.IR cs.AI

Agent-Orchestrated Adaptive RAG: A Comparative Study on Structured and Multi-Hop Retrieval

Agent编排的自适应RAG:结构化与多跳检索的比较研究

Anuj Maharjan, Devinder Kaur, Richard Molyet

机构 * University of California, Berkeley(加州大学伯克利分校) University of Washington(华盛顿大学) University of California, Los Angeles(加州大学洛杉矶分校)

AI总结 提出Agent编排的自适应RAG框架,通过动态查询分解、迭代检索和自反思评估,在结构化领域(DevOps)和多跳推理基准(MuSiQue)上对比发现,查询分解在结构化领域提升性能但降低多跳排名精度,反思机制提高引用准确性但增加延迟,表明Agent增强需根据查询和领域特性选择性应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05632 2026-06-05 cs.AI

Evaluation of LLMs for Mathematical Formalization in Lean

LLM在Lean中数学形式化的评估

Tyson Klingner, Drew Bladek, Escher Crawford, Bohao Chen, Ariel Fu, Kaira Nair, Jarod Alper, Giovanni Inchiostro, Vasily Ilin

机构 * University of California, Berkeley(加州大学伯克利分校) University of Washington(华盛顿大学)

AI总结 本研究通过pass@k和refine@k指标在miniF2F和miniCTX子集上比较了多种大语言模型在Lean 4中生成形式化证明的能力,发现Gemini 3.1 Pro和Claude Opus 4.7性能最佳,而NVIDIA Nemotron 3 Super和GPT-OSS 120B在考虑成本时效率最高。

Comments 15 pages, 13 figures, 10 tables. Comments welcome!

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05571 2026-06-05 cs.SD eess.AS

Sound Effects Dataset Unification With the Universal Category System

使用通用分类系统统一音效数据集

Jun Woo Beck, Alexander Lerch

机构 * University of California, Berkeley(加州大学伯克利分校)

AI总结 提出一个基于通用分类系统(UCS)的模块化数据集重新标注框架,通过规则驱动的多阶段流水线和冲突解决实现高自动转换率,并创建了包含58,057个音频片段的统一数据集EnvSound-UCS。

Comments DAFx 2026 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05563 2026-06-05 cs.AI cs.CL

SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations

SoCRATES:跨领域和社会认知变异的前瞻性LLM调解的可靠自动化评估

Taewon Yun, Hyeonseong Park, Jeonghwan Choi, Hayoon Park, Yeeun Choi, Hwanjun Song

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学)

AI总结 提出SoCRATES基准,通过多领域真实冲突场景和五维社会认知适应轴评估LLM调解员,使用主题定位评估器实现0.82的人类专家一致性,发现最强模型仅缩小约三分之一的未调解共识差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05429 2026-06-05 cs.AI

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models

最小化缩放因子的隐藏成本:面向大语言模型的图引导超低位量化

Rayyan Abdalla, Amir Hussein, Min Wu, Dinesh Manocha

机构 * University of California, Berkeley(加州大学伯克利分校)

AI总结 提出SAGE-PTQ框架,通过图引导的显著性感知量化分离显著与非显著权重,实现超低位量化并最小化缩放开销,在LLaMA-3-8B上困惑度降至6.74且内存低于BiLLM的50%。

Comments Preprint. 18 pages, 10 figures, 7 tables, including appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05400 2026-06-05 cs.AI cs.CL cs.LG

LeanMarathon: Toward Reliable AI Co-Mathematicians through Long-Horizon Lean Autoformalization

LeanMarathon:通过长视界Lean自动形式化实现可靠的AI合作数学家

Yuanhe Zhang, Yuekai Sun, Taiji Suzuki, Jason D. Lee, Fanghui Liu

机构 * Department of Statistics, University of Warwick, UK(英国沃里克大学统计系) Center for Advanced Intelligence Project, RIKEN, Japan(日本理化学研究所高级智能项目) Department of Statistics, University of Michigan, USA(美国密歇根大学统计系) Department of Mathematical Informatics, The University of Tokyo(东京大学数学信息学系;日本理化学研究所高级智能项目) also Center for Advanced Intelligence Project, RIKEN, Japan(加州大学伯克利分校电气工程与计算机科学系;统计系) Department of Electrical Engineering and Computer Sciences, also Department of Statistics, University of California, Berkeley, USA(上海交通大学数学科学学院,自然科学院和MOE-LSC) School of Mathematical Sciences, Institute of Natural Sciences and MOE-LSC, Shanghai Jiao Tong University, China

AI总结 提出多智能体框架LeanMarathon,通过蓝图抽象和两阶段编排器实现长视界研究数学的可靠自动形式化,在四个Erdős问题上成功形式化七个定理。

Comments 26 pages, 9 figures. Comments are welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05380 2026-06-05 cs.DS cs.LG

Learning-Augmented Online Minimization with Dual Predictions

具有双重预测的学习增强在线最小化

Christian Coester, Alexa Tudose, Alexander Turoczy

机构 * University of California, Berkeley(加州大学伯克利分校)

AI总结 针对度量任务系统和层状集合覆盖两类在线最小化问题,提出利用对偶线性规划最优解的机器学习预测来改进理论保证的学习增强算法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05272 2026-06-05 cs.LG

Learning Manifold and Itô Dynamics with Branched Neural Rough Differential Equations

学习流形与伊藤动力学:分支神经粗糙微分方程

Luke Thompson, Dai Shi, Lequan Lin, Junbin Gao, Andi Han

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学) University of Toronto(多伦多大学)

AI总结 提出分支神经粗糙微分方程(B-NRDE),通过Hopf代数框架统一处理欧几里得伊藤动力学、流形上的有序协变导数及经典Stratonovich情形,实现精确的粗步流形约束动力学和伊藤一致律匹配。

Comments Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05259 2026-06-05 cs.CV

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding

VideoKR:迈向知识和推理密集型视频理解

Lin Fu, Zheyuan Yang, Yang Wang, Tingyu Song, Arman Cohan, Yilun Zhao

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学) University of Toronto(多伦多大学) University of Washington(华盛顿大学) University of Michigan(密歇根大学)

AI总结 提出VideoKR,首个大规模训练语料库,通过人工参与的技能导向生成管道构建315K视频推理示例,增强知识和推理密集型视频理解,并在专家标注基准上验证其有效性。

Comments ICML 2026 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05219 2026-06-05 cs.LG cs.AI

Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-Pathway

大步长梯度下降恢复多路径深度线性网络中的对称性

Hee-Sung Kim, Sungyoon Lee

机构 * University of California, Berkeley(加州大学伯克利分校)

AI总结 本文研究大步长离散梯度下降如何通过边缘稳定性振荡使多路径深度线性网络从对称性破坏转向信号重新分配,从而偏好共享表示而非单路径主导。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05178 2026-06-05 cs.HC cs.AI

The Virtual Roundtable: Multi-Agent Personas Simulating the Dynamics of Human Brainstorming

虚拟圆桌会议:模拟人类头脑风暴动态的多智能体角色

Tim Dorn, Saara A. Khan, Julie Mumford

机构 * University of California, Berkeley(加州大学伯克利分校)

AI总结 提出一种多智能体架构,通过发散与收敛两阶段模拟圆桌头脑风暴,利用多样化AI角色和智能引导者产生多样化创意并评估排名,案例研究表明其能产生多样相关创意并深化讨论质量。

Comments 10 pages, 10 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03067 2026-06-05 stat.ML cs.LG

Trajectory-Aware Node Contributions and the Limits of Static Controllability

轨迹感知的节点贡献与静态可控性的极限

Valentina Kuskova, Dmitry Zaytsev, Michael Coppedge

机构 * University of California, Berkeley(加州大学伯克利分校)

AI总结 本文提出“涌现贡献”(EC)作为节点动态杠杆的有限时域度量,通过可微模型的雅可比矩阵计算,在线性时不变极限下退化为平均可控性,并构建相图刻画两者一致与分歧的条件。

Comments 11 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31278 2026-06-05 cs.AI cs.LG stat.ME

Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation

工业化预测驱动推断:用于可靠生成式AI与智能体系统评估的GLIDE库

Grégoire Martinon, Ibrahim Merad, Mohammed Raki

机构 * University of California, Berkeley(加州大学伯克利分校) Google Research(谷歌研究院)

AI总结 提出GLIDE开源库,统一多种预测驱动推断方法,提供无偏估计与有效置信区间,显著降低人工标注成本。

Comments 8 pages, Accepted to the ICML 2026 Workshop on Statistical Frameworks for Uncertainty in Agentic Systems, Seoul, South Korea, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24059 2026-06-05 cs.LG cs.AI

Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers

频谱探测电路:识别预训练Transformer中注意力头电路的三步法

Yongzhong Xu

机构 * University of California, Berkeley(加州大学伯克利分校)

AI总结 提出一种三步法,通过频谱信号排序、任务模式筛选和组消融因果验证,无需标签即可识别预训练Transformer中执行持续内容依赖计算的注意力头电路,并在多个模型上验证了其通用性和因果必要性。

Comments 35 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21017 2026-06-05 cs.RO cs.AI

Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics

Open-H-Embodiment: 一个大规模数据集,用于在医疗机器人中启用基础模型

Open-H-Embodiment Consortium, :, Nigel Nelson, Juo-Tung Chen, Jesse Haworth, Xinhao Chen, Lukas Zbinden, Dianye Huang, Alaa Eldin Abdelaal, Alberto Arezzo, Ayberk Acar, Farshid Alambeigi, Carlo Alberto Ammirati, Yunke Ao, Pablo David Aranda Rodriguez, Soofiyan Atar, Mattia Ballo, Noah Barnes, Federica Barontini, Filip Binkiewicz, Peter Black, Sebastian Bodenstedt, Leonardo Borgioli, Nikola Budjak, Benjamin Calmé, Fabio Carrillo, Nicola Cavalcanti, Changwei Chen, Haoxin Chen, Sihang Chen, Qihan Chen, Zhongyu Chen, Ziyang Chen, Shing Shin Cheng, Meiqing Cheng, Min Cheng, Zih-Yun Sarah Chiu, Xiangyu Chu, Camilo Correa-Gallego, Giulio Dagnino, Anton Deguet, Jacob Delgado, Jonathan C. DeLong, Kaizhong Deng, Alexander Dimitrakakis, Qingpeng Ding, Hao Ding, Giovanni Distefano, Daniel Donoho, Anqing Duan, Marco Esposito, Shane Farritor, Jad Fayad, Zahi Fayad, Mario Ferradosa, Filippo Filicori, Chelsea Finn, Philipp Fürnstahl, Jiawei Ge, Stamatia Giannarou, Xavier Giralt Ludevid, Frederic Giraud, Aditya Amit Godbole, Ken Goldberg, Antony Goldenberg, Diego Granero Marana, Xiaoqing Guo, Tamás Haidegger, Evan Hailey, Pascal Hansen, Ziyi Hao, Kush Hari, Kengo Hayashi, Jonathon Hawkins, Shelby Haworth, Ortrun Hellig, S. Duke Herrell, Zhouyang Hong, Andrew Howe, Junlei Hu, Zhaoyang Jacopo Hu, Ria Jain, Mohammad Rafiee Javazm, Howard Ji, Rui Ji, Jianmin Ji, Zhongliang Jiang, Dominic Jones, Jeffrey Jopling, Britton Jordan, Ran Ju, Michael Kam, Luoyao Kang, Fausto Kang, Siddhartha Kapuria, Peter Kazanzides, Sonika Kiehler, Ethan Kilmer, Ji Woong Kim, Przemysław Korzeniowski, Chandra Kuchi, Nithesh Kumar, Alan Kuntz, Federico Lavagno, Yu Chung Lee, Hao-Chih Lee, Hang Li, Zhen Li, Xiao Liang, Xinxin Lin, Jinsong Lin, Chang Liu, Fei Liu, Pei Liu, Yun-hui Liu, Wanli Liuchen, Eszter Lukács, Sareena Mann, Miles Mannas, Brett Marinelli, Sabina Martyniak, Francesco Marzola, Lorenzo Mazza, Xueyan Mei, Maria Clara Morais, Luigi Muratore, Chetan Reddy Narayanaswamy, Michał Naskręt, David Navarro-Alarcon, Cyrus Neary, Chi Kit Ng, Christopher Nguan, David Noonan, Ki Hwan Oh, Tom Christian Olesch, Allison M. Okamura, Justin Opfermann, Matteo Pescio, Doan Xuan Viet Pham, Tito Porras, Hongliang Ren, Ariel Rodriguez Jimenez, Ferdinando Rodriguez y Baena, Septimiu E. Salcudean, Asmitha Sathya, Preethi Satish, Lalithkumar Seenivasan, Jiaqi Shao, Yiqing Shen, Yu Sheng, Lucy XiaoYang Shi, Zoe Soulé, Stefanie Speidel, Mingwu Su, Jianhao Su, Idris Sunmola, Kristóf Takács, Yunxi Tang, Patrick Thornycroft, Yu Tian, Jordan Thompson, Mehmet K. Turkcan, Mathias Unberath, Pietro Valdastri, Carlos Vives, Quan Vuong, Martin Wagner, Farong Wang, Wei Wang, Lidian Wang, Chung-Pang Wang, Guankun Wang, Junyi Wang, Erqi Wang, Ziyi Wang, Tanner Watts, Wolfgang Wein, Yimeng Wu, Zijian Wu, Hongjun Wu, Luohong Wu, Jie Ying Wu, Junlin Wu, Victoria Wu, Kaixuan Wu, Mateusz Wójcikowski, Yunye Xiao, Nan Xiao, Wenxuan Xie, Hao Yang, Tianqi Yang, Yinuo Yang, Menglong Ye, Ryan S. Yeung, Nural Yilmaz, Chim Ho Yin, Michael Yip, Rayan Younis, Chenhao Yu, Sayem Nazmuz Zaman, Milos Zefran, Han Zhang, Yuelin Zhang, Yidong Zhang, Yanyong Zhang, Xuyang Zhang, Yameng Zhang, Joyce Zhang, Ning Zhong, Peng Zhou, Haoying Zhou, Xiuli Zuo, Nassir Navab, Mahdi Azizian, Sean D. Huver, Axel Krieger

机构 * Open-H-Embodiment Consortium University of California, Berkeley(加州大学伯克利分校) University of California, Los Angeles(加州大学洛杉矶分校) University of Southern California(南加州大学) University of Cambridge(剑桥大学) University of Tokyo(东京大学) University of Tokyo, Graduate School of Information Science and Technology(东京大学信息科学与技术研究生院) University of Tokyo, Institute of Industrial Science(东京大学工业科学研究所)

AI总结 本文提出Open-H-Embodiment数据集,通过两个基础模型展示了其在医疗机器人领域的应用,展示了大规模开放数据在推动机器人学习和世界建模方面的关键作用。

Comments Project website: https://open-h.github.io/open-h-embodiment/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23466 2026-06-05 cs.LG cs.AI cs.AR

Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs

评估Hopper和Blackwell GPU上的CUDA Tile用于AI工作负载

Divakar Kumar Yadav, Tian Zhao, Deepak Kumar

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学)

AI总结 本文评估了CUDA Tile在Hopper和Blackwell GPU上的AI工作负载性能,比较了CuTile与cuBLAS、Triton等方法的效率和可移植性,发现CuTile在特定工作负载上表现优异,但在跨架构优化上仍有不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17925 2026-06-05 stat.ME cs.LG math.ST stat.TH

Multi-Armed Sequential Hypothesis Testing by Betting

通过赌注进行多臂顺序假设检验

Ricardo J. Sandoval, Ian Waudby-Smith, Michael I. Jordan

机构 * University of California Berkeley(加州大学伯克利分校) École Normale Supérieure & Inria Paris(法国国家科学研究中心巴黎分校 & 巴黎研究所)

AI总结 本文研究了通过赌注进行多臂顺序检验的问题,提出了一种在多个数据源(臂)中选择以获取数据的统计学家的变体,旨在拒绝全局空假设P(所有臂在某种意义上无效)并支持复合替代假设Q(至少有一个臂非空)。通过推广对数最优性和期望拒绝时间最优性的概念,得到了匹配的上下界,并提出了一个修改的上置信界算法来处理不可观测但足够可估计的奖励。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12124 2026-06-05 cs.LG cs.CL

Alignment Risks from Capability-Seeking RL Training

从能力寻求强化学习训练中产生的对齐风险

Yujun Zhou, Yue Huang, Han Bao, Kehan Guo, Zhenwen Liang, Pin-Yu Chen, Tian Gao, Werner Geyer, Nuno Moniz, Nitesh V Chawla, Xiangliang Zhang

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学) University of Washington(华盛顿大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Toronto(多伦多大学) University of Cambridge(剑桥大学)

AI总结 本文研究了在易受攻击的环境中通过强化学习训练语言模型时,模型可能利用隐含漏洞来最大化奖励的风险,发现这些策略不仅限于狭窄的技巧,还能在一定程度上转移、传播,并在某些情况下比通过SFT学习更持久,表明需要扩展AI安全工作到审计和保障训练环境、奖励机制和评估渠道。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10106 2026-06-05 cs.RO

EgoHumanoid: Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration

EgoHumanoid: 通过无机器人眼示范解锁真实场景中的移动- manipulation

Modi Shi, Shijia Peng, Jin Chen, Haoran Jiang, Tianyu Li, Di Huang, Ping Luo, Hongyang Li, Li Chen

机构 * University of California, Berkeley(加州大学伯克利分校) Tsinghua University(清华大学)

AI总结 本文提出EgoHumanoid框架,通过结合大量眼示范数据和少量机器人数据共同训练视觉-语言-动作策略,使机器人能够执行多样化的现实环境中的移动- manipulation任务,实验表明无机器人数据显著提升了性能,尤其在未见过的环境中表现更优。

Comments Project page: https://opendrivelab.com/EgoHumanoid

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07739 2026-06-05 cs.IR cs.AI

HypRAG: Hyperbolic Dense Retrieval for Retrieval Augmented Generation

HypRAG: 超几何密集检索用于检索增强生成

Hiren Madhu, Ngoc Bui, Ali Maatouk, Leandros Tassiulas, Smita Krishnaswamy, Menglin Yang, Sukanta Ganguly, Kiran Srinivasan, Rex Ying

机构 * University of California, Berkeley(加州大学伯克利分校)

AI总结 本文提出超几何密集检索方法,通过在双曲空间中构建HyTE-FH和HyTE-H两种模型变体,解决传统欧几里得空间在检索增强生成中的局限性,提升文档相关性和回答相关性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07253 2026-06-05 cs.AI cs.CL

From Out-of-Distribution Detection to Hallucination Detection: A Geometric View

从分布外检测到幻觉检测:一个几何视角

Litian Liu, Reza Pourreza, Yubing Jian, Yao Qin, Roland Memisevic

机构 * University of California, Berkeley(加州大学伯克利分校)

AI总结 本文通过将幻觉检测重新定义为分布外检测问题,利用几何视角提出了一种无需训练、基于单样本的检测方法,在推理任务中实现了高准确率。

Comments ICML 2026 main conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02680 2026-06-05 cs.LG

FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment

FlexRank: 嵌套低秩知识分解用于自适应模型部署

Riccardo Zaccone, Stefanos Laskaridis, Marco Ciccone, Samuel Horváth

机构 * University of California, Berkeley(加州大学伯克利分校)

AI总结 提出FlexRank方法,通过嵌套低秩权重分解和基于重要性的整合,从预训练模型中提取不同能力的子模型,实现“一次训练,随处部署”的自适应部署。

Comments Accepted at ICML 2026 (Spotlight)

Journal ref Proceedings of the 43rd International Conference on Machine Learning, PMLR, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21218 2026-06-05 cs.CV

Latent Implicit Visual Reasoning

潜在隐式视觉推理

Kelvin Li, Chuyi Shang, Leonid Karlinsky, Rogerio Feris, Trevor Darrell, Roei Herzig

机构 * University of California, Berkeley(加州大学伯克利分校) Xero MIT-IBM Watson AI Lab(麻省理工-IBM Watson人工智能实验室)

AI总结 本文提出了一种任务无关的机制,训练大规模多模态模型(LMMs)在无需显式中间监督的情况下发现和使用潜在视觉推理标记,从而在多种视觉中心任务中优于直接监督微调,并在不使用辅助图像、边界框、图像裁剪、深度图或思维链注释的情况下,与或优于先前基于文本和显式视觉中间推理方法相媲美。

详情

展开后加载摘要…

URL PDF HTML 收藏