arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Carnegie Mellon University(卡内基梅隆大学)

共收录 2342
2505.22650 2026-02-16 cs.LG

On Learning Verifiers and Implications to Chain-of-Thought Reasoning

关于学习验证器及其对思维链推理的影响

Maria-Florina Balcan, Avrim Blum, Zhiyuan Li, Dravyansh Sharma

机构 * Carnegie Mellon University(卡内基梅隆大学) Toyota Technological Institute at Chicago(芝加哥丰田技术研究所) Northwestern University(西北大学)

AI总结 本文提出了一种形式化的PAC学习框架,用于学习可靠的验证器以评估自然语言思维链推理的正确性,并提供了样本复杂性上界和下界结果。

Comments 26 pages, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18138 2026-02-16 cs.LG

B3C: A Minimalist Approach to Offline Multi-Agent Reinforcement Learning

B3C: 一种针对离线多智能体强化学习的极简方法

Woojun Kim, Katia Sycara

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 B3C通过引入批评者剪裁和非线性价值分解,有效解决离线多智能体强化学习中的过估计问题,提升性能。

Comments Accepted at the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12271 2026-02-13 cs.CV cs.LG

MonarchRT: Efficient Attention for Real-Time Video Generation

MonarchRT:实时视频生成中的高效注意力机制

Krish Agarwal, Zhuoming Chen, Cheng Luo, Yongqi Chen, Haizhong Zheng, Xun Huang, Atri Rudra, Beidi Chen

机构 * Carnegie Mellon University(卡内基梅隆大学) University at Buffalo(布法罗大学) Morpheus AI

AI总结 MonarchRT通过结构化注意力参数化方法,实现高效实时视频生成,达到95%的注意力稀疏性,优于现有内核性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10078 2026-02-13 cs.SD cs.LG

Improving Speech Emotion Recognition with Mutual Information Regularized Generative Model

通过互信息正则化生成模型改进语音情感识别

Chung-Soo Ahn, Rajib Rana, Sunil Sivadas, Carlos Busso, Jagath C. Rajapakse

机构 * College of Computing and Data Science at Nanyang Technological University(南洋理工大学计算机与数据科学学院) University of Southern Queensland(南方昆士兰大学) NCS Group(NCS集团) Language Technologies Institute, School of Computer Science, Carnegie Mellon University(计算机科学学院语言技术研究所,卡内基梅隆大学)

AI总结 本文提出了一种基于互信息正则化的生成模型,通过跨模态对齐和特征合成提升语音情感识别的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11609 2026-02-13 cs.AI q-bio.GN

scPilot: Large Language Model Reasoning Toward Automated Single-Cell Analysis and Discovery

scPilot:面向自动化单细胞分析与发现的大型语言模型推理

Yiming Gao, Zhen Wang, Jefferson Chen, Mark Antkowiak, Mengzhou Hu, JungHo Kong, Dexter Pratt, Jieyuan Liu, Enze Ma, Zhiting Hu, Eric P. Xing

机构 * UC San Diego(UC圣地亚哥大学) MBZUAI CMU(卡内基梅隆大学) Texas A&M(德克萨斯A&M大学)

AI总结 scPilot通过组学原生推理提升单细胞分析的准确性与可解释性,实现自动化细胞类型注释和轨迹重建。

Comments Accepted at NeurIPS 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11374 2026-02-13 cs.LG cs.AI

Retrieval-Aware Distillation for Transformer-SSM Hybrids

面向检索的蒸馏:Transformer-SSM混合模型

Aviv Bick, Eric P. Xing, Albert Gu

机构 * Carnegie Mellon University(卡内基梅隆大学) MBZUAI(穆斯林人工智能研究院) Cartesia AI(Cartesia人工智能)

AI总结 本文提出检索感知蒸馏方法,通过保留关键注意力头并蒸馏其余头部,实现更高效的Transformer-SSM混合模型,显著降低内存消耗并提升检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11342 2026-02-13 cs.HC cs.AI

Situated, Dynamic, and Subjective: Envisioning the Design of Theory-of-Mind-Enabled Everyday AI with Industry Practitioners

情境化、动态化和主观化:与行业从业者共同展望具有理论思维能力的日常人工智能设计

Qiaosi Wang, Jini Kim, Avanita Sharma, Alicia, Lee, Jodi Forlizzi, Hong Shen

机构 * Human-Computer Interaction Institute(人机交互研究所) School of Design(设计学院) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文通过与行业从业者协同设计,提出具有理论思维能力的日常AI应具备情境化、动态化和主观化三大设计原则,以支持持续的人机交互。

Comments 16 pages, preprint for ACM CHI 2026 Conference

Journal ref Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI '26), April 13--17, 2026, Barcelona, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11242 2026-02-13 cs.CV

ReTracing: An Archaeological Approach Through Body, Machine, and Generative Systems

ReTracing:通过身体、机器和生成系统进行考古学研究

Yitong Wang, Yue Yao

机构 * Department of Machine Learning(机器学习系) Carnegie Mellon University(卡内基梅隆大学) SIPA Technology Policy and Innovation(技术政策与创新研究所) Columbia University(哥伦比亚大学)

AI总结 ReTracing通过AI、人类和机器人互动,探索生成系统如何通过编舞运动反映社会文化偏见。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11176 2026-02-13 cs.CL cs.AI cs.CE cs.LG

Evaluating Few-Shot Temporal Reasoning of LLMs for Human Activity Prediction in Smart Environments

评估LLM在智能环境中的少样本时间推理用于人类活动预测

Maral Doctorarastoo, Katherine A. Flanigan, Mario Bergés, Christopher McComb

机构 * organization= Department of Civil \& Environmental Engineering, Carnegie Mellon University , addressline= 5000 Forbes Ave , city= Pittsburgh , state= PA , postcode= 15213 , country= USA organization= Department of Mechanical Engineering, Carnegie Mellon University , addressline= 5000 Forbes Ave , city= Pittsburgh , state= PA , postcode= 15213 , country= USA

AI总结 本文研究了预训练语言模型在智能环境中的少样本时间推理能力,通过评估其在人类活动预测中的表现,发现其在低数据环境下具备强大的时间理解能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09013 2026-02-13 cs.RO cs.CV

Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction

通过3D手-物体轨迹重建从RGB人类视频中学习灵巧操作策略

Hongyi Chen, Tony Dong, Tiancheng Wu, Liquan Wang, Yash Jangir, Yaru Niu, Yufei Ye, Homanga Bharadhwaj, Zackory Erickson, Jeffrey Ichnowski

机构 * Carnegie Mellon University(卡内基梅隆大学) Georgia Institute of Technology(佐治亚理工学院) Stanford University(斯坦福大学)

AI总结 VIDEOMANIP通过3D手-物体轨迹重建从RGB视频直接学习灵巧操作策略,实现无需设备的高效训练和高成功率抓取

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15926 2026-02-13 cs.LG cs.CL cs.CY

DSO: Direct Steering Optimization for Bias Mitigation

DSO:直接转向优化以减轻偏差

Lucas Monteiro Paes, Nivedha Sivakumar, Yinong Oliver Wang, Masha Fedzechkina, Barry-John Theobald, Luca Zappella, Nicholas Apostoloff

机构 * Apple(苹果公司) Carnegie Mellon University(卡内基梅隆大学)

AI总结 DSO通过强化学习优化激活转换,在VLMs和LLMs中实现公平性与能力的最佳平衡,提供推理时的偏见控制能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23276 2026-02-13 cs.CL

A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results

宴会派对基准:多模态数据集和比较评估结果

Thai-Binh Nguyen, Katerina Zmolikova, Pingchuan Ma, Ngoc Quan Pham, Christian Fuegen, Alexander Waibel

机构 * Karlsruhe Institute of Technology, Germany(卡尔斯鲁厄理工学院) Carnegie Mellon University, USA(卡内基梅隆大学) Interactive-AI LLC(Interactive-AI公司) Meta AI, UK(Meta AI)

AI总结 本文提出多模态上下文感知识别任务,通过结合音频和视觉线索提升重叠对话识别性能,展示多模态在解决宴会派对问题中的关键作用。

Comments Accepted at ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03728 2026-02-13 cs.SD cs.LG eess.AS eess.SP

Lightweight and Generalizable Acoustic Scene Representations via Contrastive Fine-Tuning and Distillation

通过对比微调和蒸馏实现轻量且通用的声学场景表示

Kuang Yuan, Yang Gao, Xilin Li, Xinhao Mei, Syavosh Zadissa, Tarun Pruthi, Saeed Bagheri Sereshki

机构 * Meta Reality Labs(Meta现实实验室) Carnegie Mellon University(卡内基梅隆大学)

AI总结 ContrastASC通过对比微调和蒸馏实现轻量且通用的声学场景表示,提升少样本适应能力的同时保持闭合集性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17061 2026-02-13 cs.MA cs.AI cs.IR

Parallelism Meets Adaptiveness: Scalable Documents Understanding in Multi-Agent LLM Systems

并行与适应性:多智能体大语言模型系统中的可扩展文档理解

Chengxuan Xia, Qianye Wu, Sixuan Tian, Yilun Hao

机构 * University of California, Santa Cruz(加州大学圣克ruz分校) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出一种多智能体大语言模型系统协调框架,通过动态任务路由、双向反馈和并行评估机制,提升文档理解的适应性和效率。

Comments Accepted at AAAI 2026 Workshop on WoMAPF, Camera ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10793 2026-02-13 cs.SD cs.HC cs.LG eess.AS

SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures

SonicSieve:利用声学微结构在智能手机上实现定向语音提取

Kuang Yuan, Yifeng Wang, Xiyuxing Zhang, Chengyi Shen, Swarun Kumar, Justin Chan

机构 * Carnegie Mellon University(卡内基梅隆大学) Tsinghua University(清华大学) Zhejiang University(浙江大学)

AI总结 SonicSieve通过生物启发式声学微结构在智能手机上实现定向语音提取,提升信号质量并优于传统麦克风阵列。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11123 2026-02-12 cs.LG cond-mat.mtrl-sci

From Natural Language to Materials Discovery:The Materials Knowledge Navigation Agent

从自然语言到材料发现:材料知识导航代理

Genmao Zhuang, Amir Barati Farimani

机构 * Department of Materials Science and Engineering, Carnegie Mellon University(材料科学与工程系,卡内基梅隆大学) Department of Mechanical Engineering, Carnegie Mellon University(机械工程系,卡内基梅隆大学)

AI总结 材料知识导航代理通过自然语言处理技术,实现了材料发现的自动化和智能化,能够从文献和数据库中提取关键信息,提出新的材料候选方案。

Comments 22 pages,5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16376 2026-02-12 cs.CL

Polymer-Agent: Large Language Model Agent for Polymer Design

聚合物代理:用于聚合物设计的大型语言模型代理

Vani Nigam, Achuth Chandrasekhar, Amir Barati Farimani

机构 * Department of Materials Science and Engineering, Carnegie Mellon University(材料科学与工程系,卡内基梅隆大学) Department of Mechanical Engineering, Carnegie Mellon University(机械工程系,卡内基梅隆大学)

AI总结 本文提出基于LLM的聚合物代理,通过结构-性质预测和生成技术,助力实验室研究人员高效发现新型聚合物结构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15922 2026-02-12 cs.CL

Aligning Dialogue Agents with Global Feedback via Large Language Model Multimodal Reward Decomposition

通过大语言模型多模态奖励分解对齐对话代理

Dong Won Lee, Hae Won Park, Cynthia Breazeal, Louis-Philippe Morency

机构 * MIT(麻省理工学院) CMU(卡内基梅隆大学)

AI总结 本文提出了一种基于大语言模型的多模态奖励分解方法,通过分解会话级反馈来提升对话生成质量,无需人工反馈。

Comments 9 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10544 2026-02-12 cs.LG cs.NA math.NA

Bridging the Compression-Precision Paradox: A Hybrid Architecture for Clinical EEG Report Generation with Guaranteed Measurement Accuracy

弥合压缩-精度悖论:一种用于临床EEG报告生成的混合架构,保证测量精度

Wuyang Zhang, Zhen Luo, Chuqiao Gu, Jianming Ma, Yebo Cao, Wangming Yuan, Yinzhi Jin

机构 * Northeastern University(东北大学) Carnegie Mellon University(卡内基梅隆大学) George Mason University(乔治·马歇尔大学)

AI总结 本研究提出了一种混合架构,通过信号处理和跨模态翻译实现临床EEG报告生成,保证测量精度,减少误报并提升检测速度。

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10469 2026-02-12 cs.GT cs.LG math.OC

Online Generalized-mean Welfare Maximization: Achieving Near-Optimal Regret from Samples

在线广义均值福利最大化:从样本中实现近最优的遗憾

Zongjun Yang, Rachitesh Kumar, Christian Kroer

机构 * Columbia University(哥伦比亚大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出了一种在线广义均值福利最大化算法,在无需分布知识的情况下实现近最优的遗憾率,通过重新求解范式应对非平稳性和分布偏移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15975 2026-02-12 cs.RO

Fast Task Planning with Neuro-Symbolic Relaxation

快速任务规划与神经符号放松

Qiwei Du, Bowen Li, Yi Du, Shaoshu Su, Taimeng Fu, Zitong Zhan, Zhipeng Zhao, Chen Wang

机构 * Spatial AI & Robotics Lab, Department of Computer Science and Engineering, University at Buffalo(空间人工智能与机器人实验室,计算机科学与工程系,布法罗大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 Flax通过结合神经重要性预测与符号扩展,实现快速可靠的复杂环境任务规划。

Comments 8 pages, 6 figures

Journal ref IEEE Robotics and Automation Letters (RA-L), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11739 2026-02-12 cs.CL cs.AI

ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without Training

ZeroTuning: 解锁初始标记的潜力以无需训练的方式提升大语言模型

Feijiang Han, Xiaodong Yu, Jianheng Tang, Delip Rao, Weihua Du, Lyle Ungar

机构 * University of Pennsylvania(宾夕法尼亚大学) AMD(AMD公司) Peking University(北京大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 ZeroTuning通过仅调整初始标记的注意力实现无需训练的大语言模型性能提升,适用于多种任务和场景。

Comments ICLR 2026 Accepted Version: proofread, introduction rewritten, additional experiments and appendix material added

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03462 2026-02-12 cs.RO cs.SY eess.SY

Multi-Momentum Observer Contact Estimation for Bipedal Robots

双动量观测器接触估计用于双足机器人

J. Joe Payne, Daniel A. Hagen, Denis Garagić, Aaron M. Johnson

机构 * Department of Mechanical Engineering, Carnegie Mellon University(机械工程系,卡内基梅隆大学) Palladyne AI Corporation(Palladyne AI公司)

AI总结 本文提出了一种基于动量观测器的双足机器人接触模式估计方法,通过多动态模型和马尔可夫融合实现高精度接触检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.04196 2026-02-12 cs.RO

Robotic Depowdering for Additive Manufacturing Via Pose Tracking

通过姿态跟踪的机器人去粉用于增材制造

Zhenwei Liu, Junyi Geng, Xikai Dai, Tomasz Swierzewski, Kenji Shimada

机构 * Department of Mechanical Engineering, Carnegie Mellon University(机械工程系,卡内基梅隆大学) Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学)

AI总结 本文提出了一种基于视觉的机器人去粉系统,通过姿态跟踪实时去除3D打印部件表面的未熔合粉末,无需预处理即可适应不同形状的部件。

Comments Github link: https://github.com/zhenweil/Robotic-Depowdering-for-Additive-Manufacturing-Via-Pose-Tracking Video link: https://www.youtube.com/watch?v=AUIkyULAhqM

Journal ref 2022 IEEE Robotics and Automation Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22499 2026-02-11 cs.CV cs.AI stat.AP

Scalable Dynamic Origin-Destination Demand Estimation Enhanced by High-Resolution Satellite Imagery Data

基于高分辨率卫星影像数据的可扩展动态OD需求估计

Jiachao Liu, Pablo Guarda, Koichiro Niinuma, Sean Qian

机构 * Department of Civil and Environmental Engineering, Carnegie Mellon University, Pittsburgh, PA(土木与环境工程系,卡内基梅隆大学) Fujitsu Research of America, Pittsburgh, PA(富士通美国研究院) Heinz College of Information Systems and Public Policy, Carnegie Mellon University, Pittsburgh, PA(Heinz信息学院)

AI总结 本研究提出基于高分辨率卫星影像数据的动态OD需求估计框架,通过结合传统交通数据提升估计精度,尤其在无本地传感器的链接上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09040 2026-02-11 eess.AS cs.AI cs.LG cs.SD

Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures

用于联合嵌入预测架构中自监督语音表示学习的软聚类锚点

Georgios Ioannides, Adrian Kieback, Judah Goldfeder, Linsey Pang, Aman Chadha, Aaron Elkins, Yann LeCun, Ravid Shwartz-Ziv

机构 * Carnegie Mellon University(卡内基梅隆大学) New York University(纽约大学) James Silberrad Brown Center for AI(詹姆斯·西伯拉德·布朗人工智能中心) Columbia University(哥伦比亚大学) Northeastern University(东北大学) Stanford University(斯坦福大学) Amazon GenAI(亚马逊生成人工智能)

AI总结 GMM-Anchored JEPA通过软聚类锚点提升语音表示学习,实现ASR、情感识别和槽填充的性能提升。

Comments 15 pages, 5 figures. Code: github.com/gioannides/clustering-anchored-jepa

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07087 2026-02-11 cs.LG

When Should We Introduce Safety Interventions During Pretraining?

在预训练过程中何时引入安全干预

Dylan Sam, Sachin Goyal, Pratyush Maini, Alexander Robey, J. Zico Kolter

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文研究了在预训练过程中何时引入安全干预以提高模型鲁棒性,发现20%-60%的预训练阶段引入干预效果最佳,且早期干预能更清晰地区分安全与有害内容。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08613 2026-02-11 cs.CV eess.IV

Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework

评估说话面部生成中的身份泄露:指标与评估框架

Dogucan Yaman, Fevziye Irem Eyiokur, Hazım Kemal Ekenel, Alexander Waibel

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) KIT Campus Transfer GmbH (KCT)(KIT校园转移公司) Istanbul Technical University(伊斯坦布尔技术大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出了一种评估说话面部生成中身份泄露的系统方法,通过三种测试设置和衍生指标量化泄露问题,并探讨参考图像选择对泄露的影响。

Comments Accepted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11362 2026-02-11 cs.LG cs.CV

PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits

PersonaX: 多模态数据集与LLM推断行为特征

Loka Li, Wong Yu Kang, Minghao Fu, Guangyi Chen, Zhenhao Chen, Gongxu Luo, Yuewen Sun, Salman Khan, Peter Spirtes, Kun Zhang

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) Carnegie Mellon University(卡内基梅隆大学) University of California San Diego(加州大学圣地亚哥分校) Australian National University(澳大利亚国立大学)

AI总结 PersonaX通过多模态数据集结合LLM推断行为特征,推动多模态特征分析与因果推理发展。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09093 2026-02-11 stat.ML cs.LG math.OC

Sharp High-Probability Rates for Nonlinear SGD under Heavy-Tailed Noise via Symmetrization

非线性SGD在重尾噪声下的高概率收敛速率研究

Aleksandar Armacki, Dragana Bajovic, Dusan Jakovetic, Soummya Kar

机构 * École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院) University of Novi Sad(诺维萨德大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出非线性SGD在重尾噪声下的高概率收敛速率分析,通过对称化方法改进非对称噪声处理,实现更优的收敛性能。

Comments 43 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏