arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Carnegie Mellon University(卡内基梅隆大学)

共收录 2338
2508.04660 2026-05-12 cs.CL

Composing Policy Gradients and Prompt Optimization for Language Model Programs

将策略梯度与提示优化组合用于语言模型程序

Noah Ziems, Dilara Soylu, Lakshya A Agrawal, Isaac Miller, Liheng Lai, Chen Qian, Kaiqiang Song, Meng Jiang, Dan Klein, Matei Zaharia, Karel D'Oosterlinck, Christopher Potts, Omar Khattab

机构 * University of Notre Dame(诺特大学) Stanford University(斯坦福大学) UC Berkeley(伯克利大学) Anyscale CMU(卡内基梅隆大学) Zoom, Inc.(Zoom公司) Contextual AI MIT(麻省理工学院)

AI总结 本文研究了如何将GRPO与自动提示优化结合,以提升多提示程序的性能,实验表明这种组合在分类、多跳搜索和隐私保护委托任务中平均提升准确率11%。

Comments ACM CAIS 2026. Lakshya*, Dilara*, and Noah* contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09995 2026-05-12 cs.CL

Annotations Mitigate Post-Training Mode Collapse

标注缓解训练后模式崩溃

Jacob Mitchell Springer, Madhu Advani, Lukas Aichberger, Arwen Bradley, Eran Malach, Omid Saremi, Sinead Williamson, Preetum Nakkiran, Etai Littwin, Aditi Raghunathan

机构 * Carnegie Mellon University(卡内基梅隆大学) Apple(苹果公司) Johannes Kepler University Linz(林茨约翰尼斯·开普勒大学)

AI总结 本文提出标注锚定训练方法,通过在预训练中引入语义标注,减少训练后模型的语义模式崩溃,提升模型多样性。

Comments 21 pages, 8 figures, 11 tables. Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09915 2026-05-12 cs.CL cs.AI cs.CY

Position: Academic Conferences are Potentially Facing Denominator Gaming Caused by Fully Automated Scientific Agents

位置:学术会议可能面临由完全自动化科学代理引发的分母游戏威胁

Rong Shan, Te Gao, Hang Zheng, Yunjia Xi, Jiachen Zhu, Zeyu Zheng, Yong Yu, Weinan Zhang, Jianghao Lin

机构 * Shanghai Jiao Tong University(上海交通大学) Central South University(中南大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文探讨了由完全自动化科学代理引发的分母游戏威胁,分析了其对学术会议评审系统的影响,并提出系统性政策改革作为缓解措施。

Comments Accepted by ICML'26 Position Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09539 2026-05-12 cs.CL

TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems

TacoMAS: 在基于大语言模型的多智能体系统中拓扑与能力的测试时间共演

Chen Xu, Yicheng Hu, Ruizi Wang, Xinyu Lin, Wenjie Wang, Dongrui Liu, Fuli Feng

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Shanghai AI Lab(上海人工智能实验室)

AI总结 TacoMAS通过联合调整拓扑和能力,以不同时间尺度优化多智能体系统,实验表明其在四个基准测试中优于20个基线模型,平均提升13.3%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09294 2026-05-12 cs.LG cs.AI

Towards Effective Theory of LLMs: A Representation Learning Approach

面向大语言模型的有效理论:一种表征学习方法

Muhammed Ustaomeroglu, Guannan Qu

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出RET框架,通过学习宏观状态而非微观细节描述大语言模型计算,验证其在可解释性中的实用性。

Comments Project webpage: https://ustaomeroglu.github.io/RET/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20172 2026-05-12 cs.LG math.ST stat.ML stat.TH

Cover meets Robbins while Betting on Bounded Data: $\ln n$ Regret and Almost Sure $\ln\ln n$ Regret

覆盖与罗宾斯在有界数据上的博弈:$\ln n$的遗憾和几乎确定的$\ln\ln n$遗憾

Shubhada Agrawal, Aaditya Ramdas

机构 * Indian Institute of Science(印度科学研究院) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出一种混合投注策略,结合了罗宾斯和覆盖的见解,实现对随机数据的适应性和对抗数据的保护,展示了$\ln\ln n$的几乎确定遗憾和$\ln n$的最坏情况遗憾。

Comments Improved a regret bound. New regret bound for a classical mixture

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19835 2026-05-12 cs.LG cs.AI

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts

专家再利用:移动专家架构的计算高效前沿转移

Chaitanya Dwivedi, Binxuan Huang, Himanshu Gupta, Pratik Jayarao, Neeraj Varshney, Bing Yin

机构 * Amazon Stores Foundation AI(亚马逊商店基金会人工智能) Carnegie Mellon University(卡内基梅隆大学) Anthropic

AI总结 本文提出专家再利用方法,通过持续预训练逐步扩展MoE模型容量,降低训练成本,实验显示在7B-13B参数规模下,模型验证损失与基线相当,节省32%的GPU小时。

Comments 9 Pages in main paper, 29 Pages total

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06473 2026-05-12 cs.LG

MICA: Multivariate Infini Compressive Attention for Time Series Forecasting

MICA:多变量无限压缩注意力用于时间序列预测

Willa Potosnak, Nina Żukowska, Michał Wiliński, Dan Howarth, Ignacy Stępka, Mononito Goswami, Artur Dubrawski

机构 * Carnegie Mellon University(卡内基梅隆大学) Amazon(亚马逊)

AI总结 MICA通过引入跨通道注意力机制,提升多变量时间序列预测的效率与准确性,显著降低预测误差,优于传统Transformer模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11181 2026-05-12 cs.CL

Code Mixologist : A Practitioner's Guide to Building Code-Mixed LLMs

代码调酒师:构建代码混合LLM的从业者指南

Himanshu Gupta, Pratik Jayarao, Chaitanya Dwivedi, Neeraj Varshney

机构 * Arizona State University(亚利桑那州立大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文探讨了代码混合和切换在大语言模型中的挑战,提出统一的分类体系和实用指南,涵盖数据、建模和评估方法,分析现有评估实践并讨论安全问题。

Comments 8 pages main paper, 13 pages total

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18880 2026-05-12 cs.CL cs.AI cs.CY

Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction

LLMs能否估计学生困难?人类-人工智能难度对齐用于项目难度预测的 proficiency 模拟

Ming Li, Han Chen, Yunze Xiao, Jian Chen, Hong Jiao, Tianyi Zhou

机构 * University of Maryland(马里兰大学) Carnegie Mellon University(卡内基梅隆大学) University at Buffalo(布法罗大学) MBZUAI

AI总结 本文研究了LLMs在估计学生困难方面的表现,发现模型规模扩大并不总能提高准确性,且模型倾向于形成机器共识而非与人类对齐,揭示了当前模型在自动难度预测中的挑战。

Comments ACL2026, camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13332 2026-05-12 cs.AI cs.CL

Explicit Reasoning Makes Better Judges: A Systematic Study on Accuracy, Efficiency, and Robustness

显式推理使评判更可靠:对准确性、效率和鲁棒性的系统研究

Pratik Jayarao, Himanshu Gupta, Neeraj Varshney, Chaitanya Dwivedi

机构 * Arizona State University(亚利桑那州立大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文通过对比显式推理与非显式推理模型在RewardBench任务中的表现,发现显式推理模型在准确性、效率和鲁棒性上均优于非显式模型,且在多语言环境下也表现出优势。

Comments Accepted in 2025 NeurIPS Foundations of Reasoning in Language Models Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12708 2026-05-12 cs.CL

AgentReview: Exploring Peer Review Dynamics with LLM Agents

AgentReview: 探索基于LLM代理的同行评审动态

Yiqiao Jin, Qinlin Zhao, Yiyang Wang, Hao Chen, Kaijie Zhu, Yijia Xiao, Jindong Wang

机构 * Georgia Institute of Technology(佐治亚理工学院) University of Science and Technology of China(中国科学技术大学) Carnegie Mellon University(卡内基梅隆大学) University of California, Santa Barbara(加州大学圣巴巴拉分校) University of California, Los Angeles(加州大学洛杉矶分校) William & Mary(威廉与玛丽学院)

AI总结 本文提出AgentReview框架,通过LLM模拟同行评审过程,揭示评审者偏见导致论文决策37.1%的变异,结合社会学理论提升评审机制设计。

Comments Accepted at EMNLP 2024 Main Track (Oral). https://agentreview.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.00476 2026-05-12 math.OC cs.LG stat.ML

The Mixing method: low-rank coordinate descent for semidefinite programming with diagonal constraints

混合方法:带有对角约束的结构半正定规划的低秩坐标下降法

Po-Wei Wang, Wei-Cheng Chang, J. Zico Kolter

机构 * Carnegie Mellon University(卡内基梅隆大学) Bosch Center for Artificial Intelligence(博世人工智能中心)

AI总结 本文提出了一种低秩坐标下降方法,用于求解带对角约束的结构半正定规划问题。该方法在优化性能上显著优于现有方法,并证明其在随机初始化下以局部线性速率收敛到全局最优解。

Comments The proof has been updated to match the version presented in the 2021 thesis: https://ml.cmu.edu/research/phd-dissertation-pdfs/thesis_poweiw.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09243 2026-05-12 cs.AI q-bio.NC

How Much is Brain Data Worth for Machine Learning?

脑数据对机器学习有多大价值?

Lane Lewis, Zhixin Wang, David Schwab, Xaq Pitkow

机构 * Neuroscience Institute, Carnegie Mellon University, Pittsburgh, PA, USA(卡内基梅隆大学神经科学研究所) Department of Machine Learning, Carnegie Mellon University, Pittsburgh, PA, USA(卡内基梅隆大学机器学习系) NSF AI Institute for Artificial and Natural Intelligence (ARNI)(国家科学基金会人工智能与自然智能研究所) Carnegie Mellon University, Pittsburgh, PA, USA(卡内基梅隆大学) CUNY Graduate Center, New York, NY, USA(纽约市立大学研究生中心)

AI总结 本文通过数学建模探讨脑数据与任务数据在机器学习中的价值与交换率,分析不同条件下脑数据对模型性能的提升作用。

Comments 9 pages main text, 5 figures, 34 pages of appendix with detailed proofs

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09218 2026-05-12 cs.CV cs.AI cs.LG cs.RO

Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models

Flame3D: 无需3D特定训练的3D场景零样本组合推理

Sagar Bharadwaj, Ziyong Ma, Anurag Ghosh, Srinivasan Seshan, Anthony Rowe

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 Flame3D通过可组合的空间工具和训练自由框架,实现3D场景的零样本组合推理,支持动态空间合成和外部数据整合,展示了在ScanQA和Compose3D上的竞争力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08703 2026-05-12 cs.AI cs.CL cs.CV cs.LG

RewardHarness: Self-Evolving Agentic Post-Training

RewardHarness: 自我进化代理的后训练

Yuxuan Zhang, Penghui Du, Bo Li, Cong Wei, Junwen Miao, Huaisong Zhang, Songcheng Cai, Yubo Wang, Dongfu Jiang, Yuyu Zhang, Ping Nie, Wenhu Chen, Changqian Yu, Kelsey R. Allen

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute(向量研究所) Kolors Team, Kuaishou Technology(快手团队) Carnegie Mellon University(卡内基梅隆大学) University of Waterloo(滑铁卢大学) Etude AI Tsinghua University(清华大学) Georgia Institute of Technology(佐治亚理工学院)

AI总结 RewardHarness通过自我进化代理框架,利用少量人类偏好示例迭代优化工具库,实现高效图像编辑评估,准确率达47.4%,超越GPT-5 5.3个百分点。

Comments Project page: https://rewardharness.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08556 2026-05-12 cs.LG

Can Revealed Preferences Clarify LLM Alignment and Steering?

揭示偏好能否澄清大语言模型对齐与引导?

Khurram Yamin, Jingjing Tang, Eric Horvitz, Bryan Wilder

机构 * Carnegie Mellon University(卡内基梅隆大学) Microsoft Research(微软研究院)

AI总结 本文提出一种经验方法,通过估计LLM在决策任务中的隐含偏好,评估模型是否一致地追求目标,是否能描述其目标,以及提示是否能可靠引导策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05682 2026-05-12 cs.HC cs.AI cs.CY

PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI

PersonaTeaming: 支持基于人设的生成AI红队测试

Wesley Hanwen Deng, Mingxi Yan, Sunnie S. Y. Kim, Akshita Jha, Lauren Wilcox, Kenneth Holstein, Motahhare Eslami, Leon A. Gatys

机构 * Carnegie Mellon University(卡内基梅隆大学) Apple(苹果公司)

AI总结 本文提出PersonaTeaming方法,通过整合人设提升红队测试的自动化与人机协作能力,实验显示其在攻击成功率和提示多样性方面优于现有方法,并通过用户界面促进人机协同创新。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00445 2026-05-12 cs.LG

The Power of Order: Fooling LLMs with Adversarial Table Permutations

顺序的力量:用对抗性表格排列欺骗LLMs

Xinshuai Dong, Haifeng Chen, Xuyuan Liu, Shengyu Chen, Haoyu Wang, Shaoan Xie, Kun Zhang, Zhengzhang Chen

机构 * CMU(卡内基梅隆大学) NEC Labs(NEC 实验室) Dartmouth College(达特茅斯学院) MBZUAI

AI总结 研究揭示LLMs对表格布局的脆弱性,提出对抗性表格排列攻击,证明其对多种LLM性能的破坏,凸显结构数据处理的潜在弱点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06286 2026-05-12 cs.AI

When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs

当代理人言辞与行为不一致时:从LLMs中验证提取的信念

Khurram Yamin, Jingjing Tang, Santiago Cortes-Gomez, Amit Sharma, Eric Horvitz, Bryan Wilder

机构 * Carnegie Mellon University(卡内基梅隆大学) Microsoft Research(微软研究院)

AI总结 本文提出一种决策理论框架,通过提取概率判断和决策并检验其一致性,验证LLMs的信念是否合理,发现最强模型的不一致程度较小。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01015 2026-05-12 cs.CL cs.CY

Large Language Models as Students Who Think Aloud: Overly Coherent, Verbose, and Confident

大型语言模型作为思考出声的学生:过于连贯、啰嗦且自信

Conrad Borchers, Jill-Jênn Vie, Roger Azevedo

机构 * Carnegie Mellon University(卡内基梅隆大学) Soda Team, Inria Saclay(Soda团队,法国国家科学研究中心萨克雷分部) University of Central Florida(中央佛罗里达大学)

AI总结 本文评估LLM在模拟学生思考过程中的表现,发现其推理过于连贯、啰嗦且缺乏变异性,揭示了使用生成式AI设计适应性系统时的认知局限。

Comments Manuscript under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02012 2026-05-12 cs.CV cs.LG

Improved Mean Flows: On the Challenges of Fastforward Generative Models

改进的均值流:关于快速前向生成模型挑战的研究

Zhengyang Geng, Yiyang Lu, Zongze Wu, Eli Shechtman, J. Zico Kolter, Kaiming He

机构 * CMU(卡内基梅隆大学) MIT(麻省理工学院) Adobe(Adobe公司) THU(清华大学)

AI总结 本文改进了MeanFlow框架,通过重新参数化训练目标和指导机制,提升了训练稳定性与灵活性,实现了在ImageNet上的1.72 FID成绩。

Comments Technical report. Code at https://github.com/Lyy-iiis/imeanflow

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06096 2026-05-12 stat.ML cs.AI cs.LG stat.ME

Post-detection inference for sequential changepoint localization

检测后的推断用于序列变化点定位

Aytijhya Saha, Aaditya Ramdas

机构 * Massachusetts Institute of Technology(麻省理工学院) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出了一种通用框架,用于在序列变化点检测后构建置信集,不依赖特定假设,理论上有保证,且在实践中应用广泛。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08348 2026-05-12 cs.CL

How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits

电路能告诉我们多少?测量语言模型电路的一致性和特异性

Michael Li, Nishant Subramani

机构 * Language Technologies Institute, Carnegie Mellon University, Pittsburgh, Pennsylvania, USA(语言技术研究所,卡内基梅隆大学,匹兹堡,宾夕法尼亚州,美国)

AI总结 研究通过分析六个任务和七个模型的电路重用情况,发现任务内电路重用高且共享组件对性能至关重要,但电路不具任务特异性,这引发了对电路支持针对性理解和干预的质疑。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08326 2026-05-12 cs.LG cs.AI

LLM Advertisement based on Neuron Auctions

基于神经拍卖的大型语言模型广告

Peiran Yun, Wenxin Xu, Jiayuan Liu, Yihang Zhang, Liang Zeng, Lingkai Kong, Tonghan Wang

机构 * Tsinghua University(清华大学) Carnegie Mellon University(卡内基梅隆大学) Harvard University(哈佛大学)

AI总结 本文提出神经拍卖机制,通过在LLM内部表示中进行拍卖,解决广告嵌入中的收益、平台收入与用户体验平衡问题,实现策略-proof且优化平台收益。

Comments 17 pages, 9 figures, including appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08250 2026-05-12 cs.CV cs.AI

Why Do DiT Editors Drift? Plug-and-Play Low Frequency Alignment in VAE Latent Space

为何DiT编辑器会漂移?VAE潜在空间中的插件式低频对齐

Xiaoce Wang, Sifan Zhou, Kaifei Wang, Leli Xu, Xuerui Qiu, Tao He, Ming Li

机构 * Tsinghua University(清华大学) Carnegie Mellon University(卡内基梅隆大学) Peking University(北京大学) CASIA University of Electronic Science and Technology of China(电子科技大学) Guangming Laboratory(光明实验室)

AI总结 本文从潜在空间频率角度研究DiT编辑器漂移问题,提出无需训练的VAE-LFA方法,通过低频对齐抑制语义漂移并保持高频细节,适用于白盒和黑盒DiT编辑器。

Comments 9 pages main paper, 12 figures, 25 pages in total

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23238 2026-05-12 cs.CR cs.AI

Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models

明目张胆:面向可检测性的推理模型抗蒸馏

Max Hartman, Vidhata Jayaraman, Moulik Choraria, Yash Savani, Lav R. Varshney

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Carnegie Mellon University(卡内基梅隆大学) Stony Brook University(石溪大学)

AI总结 本文提出一种面向可检测性的抗蒸馏方法,通过Stackelberg博弈框架显式编码可检测性约束,以稀疏扰动替代全迹中毒,提升教师模型性能并降低防御可见性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04333 2026-05-12 cs.LG cs.AI

What Does Flow Matching Bring To TD Learning?

流匹配为时差学习带来了什么?

Bhavya Agrawalla, Michal Nauman, Aviral Kumar

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Warsaw(华沙大学)

AI总结 本文探讨了流匹配在时差学习中的作用,指出其通过测试时恢复和增强特征学习机制提升性能,优于传统批评者。

Comments Added code link, updated acknowledgements

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21677 2026-05-12 cs.LG cs.SE

Prophecy: Inferring Formal Properties from Neuron Activations

预言:从神经元激活推断形式属性

Divya Gopinath, Corina S. Pasareanu, Muhammad Usman

机构 * NASA Ames(NASA阿姆斯研究中心) KBR(KBR公司) CMU(卡内基梅隆大学)

AI总结 Prophecy通过分析神经网络内部层的神经元激活状态,自动推断前馈网络的形式属性,适用于不同模型和输出属性的验证与监控。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08060 2026-05-11 cs.CL cs.AI cs.GT cs.MA

The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents

记忆诅咒:扩展回忆如何在LLM代理中侵蚀合作意图

Jiayuan Liu, Tianqin Li, Shiyi Du, Xin Luo, Haoxuan Zeng, Emanuel Tewolde, Tai Sing Lee, Tonghan Wang, Carl Kingsford, Vincent Conitzer

机构 * Carnegie Mellon University(卡内基梅隆大学) Foundations of Cooperative AI Lab (FOCAL)(合作人工智能基础实验室) University of Michigan(密歇根大学) Harvard University(哈佛大学)

AI总结 研究发现扩展回忆会系统性地削弱多代理社会困境中的合作意图,通过三种分析揭示记忆内容对合作的影响,证明记忆是影响多代理行为的主动因素。

详情

展开后加载摘要…

URL PDF HTML 收藏