arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Carnegie Mellon University(卡内基梅隆大学)

共收录 2349
2601.06189 2026-01-13 cs.AI cs.LG

Rational Synthesizers or Heuristic Followers? Analyzing LLMs in RAG-based Question-Answering

理性合成器还是启发式追随者?分析基于检索增强生成的问答系统

Atharv Naphade

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文通过GroupQA数据集分析LLM在基于检索增强生成的问答系统中如何整合冲突证据,发现模型倾向于优先证据且解释不忠实,表明LLM作为脆弱的启发式追随者。

Comments 13 pages, 9 figures, ACL ARR submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06101 2026-01-13 cs.CY cs.AI

How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures

如何评估人工智能素养:自我报告与基于客观的评估之间的不一致

Shan Zhang, Ruiwei Xiao, Anthony F. Botelho, Guanze Liao, Thomas K. F. Chiu, John Stamper, Kenneth R. Koedinger

机构 * University of Florida(佛罗里达大学) Carnegie Mellon University(卡内基梅隆大学) National Tsing Hua University(国立清华大学) The Chinese University of Hong Kong(香港中文大学)

AI总结 本研究开发并评估了教师人工智能素养的自我报告和基于客观的测量方法,揭示了两者在不同教师群体中的显著差异,并提出了适用于专业发展的诊断工具。

Comments 16 pages, 6 figures, LAK2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12216 2026-01-13 cs.SE cs.AI cs.CL

Training Versatile Coding Agents in Synthetic Environments

在合成环境中训练多功能编码代理

Yiqi Zhu, Apurva Gandhi, Graham Neubig

机构 * Tsinghua University(清华大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 SWE-Playground通过从头生成项目和任务,训练多功能编码代理,能够处理更广泛的编码任务并提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16976 2026-01-13 cs.LG math.ST stat.ML stat.TH

Gradient descent for deep equilibrium single-index models

深度均衡单索引模型的梯度下降

Sanjit Dandapanthula, Aaditya Ramdas

机构 * Carnegie Mellon University, Department of Statistics(卡内基梅隆大学统计系) Carnegie Mellon University, Department of Statistics and Machine Learning(卡内基梅隆大学统计与机器学习系)

AI总结 本文研究了深度均衡单索引模型的梯度下降动态,证明了线性收敛性和守恒定律,通过实验验证了理论结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11964 2026-01-13 cs.RO

E2-BKI: Evidential Ellipsoidal Bayesian Kernel Inference for Uncertainty-aware Gaussian Semantic Mapping

E2-BKI: 基于证据的椭圆贝叶斯核推断的不确定性感知高斯语义映射

Junyoung Kim, Minsik Jeon, Jihong Min, Kiho Kwak, Junwon Seo

机构 * Agency for Defense Development(国防发展局) Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)

AI总结 E2-BKI通过证据深度学习和几何对齐核函数,实现不确定性感知的高斯语义映射,提升复杂户外环境中的映射鲁棒性和效率。

Comments Accepted to IEEE RA-L. Our project website can be found at https://kjyoung.github.io/Homepage/#/Projects/E2-BKI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11692 2026-01-13 cs.CV

ORB-SfMLearner: ORB-Guided Self-supervised Visual Odometry with Selective Online Adaptation

ORB-SfMLearner: 基于ORB的自监督视觉里程计与选择性在线适应

Yanlin Jin, Rui-Yang Ju, Haojun Liu, Yuzhong Zhong

机构 * College of Electrical Engineering, Sichuan University(四川大学电气工程学院) Rice University(里士满大学) Graduate Institute of Networking and Multimedia, National Taiwan University(台湾大学网络与多媒体研究所) Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所)

AI总结 ORB-SfMLearner通过引入ORB特征和交叉注意力机制,提升视觉里程计的精度与通用性,实现更稳健的自身运动估计。

Comments ICRA 2025; Project page: https://www.neiljin.site/projects/orbsfm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12143 2026-01-12 cs.LG cs.CL stat.ML

Simple Mechanisms for Representing, Indexing and Manipulating Concepts

简单机制用于表示、索引和操作概念

Yuanzhi Li, Raghu Meka, Rina Panigrahy, Kulin Shah

机构 * Carnegie Mellon University(卡内基梅隆大学) Google Research(谷歌研究) UT Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出通过多项式零集和矩统计量来表征概念,并利用签名发现概念的共同结构和层次关系。

Comments 29 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05473 2026-01-12 cs.CL cs.HC

Towards Valid Student Simulation with Large Language Models

迈向基于大语言模型的可信学生模拟

Zhihao Yuan, Yunze Xiao, Ming Li, Weihao Xuan, Richard Tong, Mona Diab, Tom Mitchell

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Maryland(马里兰大学) The University of Tokyo(东京大学) NEOLAF

AI总结 本文提出基于大语言模型的学生模拟框架,通过知识状态规范解决能力悖论问题,强调知识忠实度对教育应用的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05431 2026-01-12 cs.LG

Prediction of Fault Slip Tendency in CO${_2}$ Storage using Data-space Inversion

利用数据空间反演预测二氧化碳储存中的断层滑动倾向

Xiaowen He, Su Jiang, Louis J. Durlofsky

机构 * Department of Energy Science and Engineering, Stanford University(能源科学与工程系,斯坦福大学) Department of Civil and Environmental Engineering, Carnegie Mellon University(土木与环境工程系,卡内基梅隆大学)

AI总结 本研究提出一种基于变分自编码器的数据空间反演方法,用于预测二氧化碳储存中的压力、应力、应变场及断层滑动倾向,同时减少关键地质参数的不确定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05174 2026-01-09 cs.LG cs.AI

FaST: Efficient and Effective Long-Horizon Forecasting for Large-Scale Spatial-Temporal Graphs via Mixture-of-Experts

FaST: 为大规模时空图实现高效且有效的长周期预测

Yiji Zhao, Zihao Zhong, Ao Wang, Haomin Wen, Ming Jin, Yuxuan Liang, Huaiyu Wan, Hao Wu

机构 * Yunnan University(云南大学) Carnegie Mellon University(卡内基梅隆大学) Griffith University(格里菲斯大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Beijing Jiaotong University(北京交通大学)

AI总结 FaST通过异构性感知混合专家框架,实现大规模时空图的高效长周期预测,提升预测精度与计算效率。

Comments Accepted to KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04458 2026-01-09 cs.LG

Using Large Language Models to Detect Socially Shared Regulation of Collaborative Learning

利用大语言模型检测协作学习中的社会共享调节

Jiayi Zhang, Conrad Borchers, Clayton Cohn, Namrata Srivastava, Caitlin Snyder, Siyuan Guo, Ashwin T S, Naveeduddin Mohammed, Haley Noh, Gautam Biswas

机构 * University of Pennsylvania(宾夕法尼亚大学) Carnegie Mellon University(卡内基梅隆大学) Vanderbilt University(范德堡大学) University of Detroit Mercy(底特律默克大学) New York University(纽约大学)

AI总结 本文利用大语言模型检测协作学习中的社会共享调节行为,通过嵌入方法提升学习分析的预测能力。

Comments Short research paper accepted at Learning Analytics and Knowledge (LAK '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23941 2026-01-09 cs.CL cs.CY

Disentangling Learning from Judgment: Representation Learning for Open Response Analytics

分离学习与判断:面向开放性回答的表示学习

Conrad Borchers, Manit Patel, Seiyon M. Lee, Anthony F. Botelho

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Florida(佛罗里达大学)

AI总结 本文提出一种分离学习与判断的分析框架,通过建模教师先验和内容嵌入,提升开放性回答评分的准确性与可审计性。

Comments Short research paper accepted at Learning Analytics and Knowledge (LAK '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04055 2026-01-08 cs.CL

Modular Prompt Optimization: Optimizing Structured Prompts with Section-Local Textual Gradients

模块化提示优化:通过部分局部文本梯度优化结构化提示

Prith Sharma, Austin Z. Henley

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 模块化提示优化通过局部文本梯度优化结构化提示,提升小型开源LLM的推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04054 2026-01-08 cs.LG

LinkD: AutoRegressive Diffusion Model for Mechanical Linkage Synthesis

LinkD:用于机械连杆合成的自回归扩散模型

Yayati Jadhav, Amir Barati Farimani

机构 * Department of Mechanical Engineering, Carnegie Mellon University, Pittsburgh, PA, USA(机械工程系,卡内基梅隆大学,匹兹堡,PA,USA)

AI总结 LinkD通过自回归扩散模型实现机械连杆的高效逆向设计,结合因果Transformer和DDPM,支持动态生成和自适应修正,适用于大规模复杂机械系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03444 2026-01-08 cs.CL cs.AI cs.HC

Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale

评分尺度对LLM作为裁判的影响:人类与LLM的对齐在0-5评分尺度上最高

Weiyue Li, Minda Zhao, Weixuan Dong, Jiahui Cai, Yuze Wei, Michael Pocress, Yi Li, Wanyan Yuan, Xiaoyue Wang, Ruoyu Hou, Kaiyuan Lou, Wenqi Zeng, Yutong Yang, Yilun Du, Mengyu Wang

机构 * Harvard University(哈佛大学) CMU(卡内基梅隆大学) Stanford University(斯坦福大学) UC San Diego(圣地亚哥大学)

AI总结 本文研究了评分尺度对LLM作为裁判一致性的影响,发现0-5评分尺度在人类与LLM之间产生最强的一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03265 2026-01-08 cs.CL cs.CR cs.LG

Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models

Jailbreak-Zero:大型语言模型安全评估的帕累托最优渗透测试路径

Kai Hu, Abhinav Aggarwal, Mehran Khodabandeh, David Zhang, Eric Hsin, Li Chen, Ankit Jain, Matt Fredrikson, Akash Bharadwaj

机构 * Meta Superintelligence Labs(Meta超智能实验室) Carnegie Mellon University(卡内基梅隆大学)

AI总结 Jailbreak-Zero通过生成多样化对抗性提示并微调攻击模型,实现了LLM安全评估的帕累托最优,提高了攻击成功率并减少了人工干预需求。

Comments Socially Responsible and Trustworthy Foundation Models at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11251 2026-01-08 cs.AI cs.CC math.HO q-bio.NC

Explaining Necessary Truths

解释必要的真理

Gülce Kardeş, Simon DeDeo

机构 * Department of Computer Science, University of Colorado, Boulder, CO 80309, USA(计算机科学系,科罗拉多大学,布洛克,科罗拉多80309,美国) the Santa Fe Institute, Santa Fe, NM 87501, USA(圣菲研究所,圣菲,新墨西哥87501,美国) Carnegie Mellon University, Pittsburgh, PA 15213, USA(卡内基梅隆大学,匹兹堡,宾夕法尼亚15213,美国)

AI总结 本文提出基于计算复杂性的框架,解释逻辑必然真理的成因,并通过模拟人类解决SAT谜题验证理论。

Comments 7 pages, in review

Journal ref Proceedings of the Annual Meeting of the Cognitive Science Society, 47 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03168 2026-01-07 cs.CL cs.LG

Can Embedding Similarity Predict Cross-Lingual Transfer? A Systematic Study on African Languages

嵌入相似性能否预测跨语言迁移?针对非洲语言的系统研究

Tewodros Kederalah Idris, Prasenjit Mitra, Roald Eiselen

机构 * Carnegie Mellon University Africa(卡内基梅隆大学非洲分校) North-West University(北开普敦大学)

AI总结 本文通过系统评估五种嵌入相似性度量,发现余弦间隙和检索度量能有效预测跨语言迁移成功,同时揭示了模型特定分析的重要性。

Comments 13 pages, 1 figure, 19 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02906 2026-01-07 cs.CL

Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration

语音基础模型中的线性脚本表示可实现零样本转写

Ryan Soh-Eun Shim, Kwanghee Choi, Kalvin Chang, Ming-Hao Hsu, Florian Eichin, Zhizheng Wu, Alane Suhr, Michael A. Hedderich, David Harwath, David R. Mortensen, Barbara Plank

机构 * LMU Munich(慕尼黑大学) Munich Center for Machine Learning(慕尼黑机器学习中心) University of Texas at Austin(德克萨斯大学奥斯汀分校) Carnegie Mellon University(卡内基梅隆大学) University of California, Berkeley(加州大学伯克利分校) Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

AI总结 通过在语音基础模型中引入线性脚本表示,实现对语音识别输出脚本的零样本控制,提升多语言转写性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02499 2026-01-07 cs.LG

Polynomial Convergence of Riemannian Diffusion Models

黎曼扩散模型的多项式收敛性

Xingyu Xu, Ziyi Zhang, Yorie Nakahira, Guannan Qu, Yuejie Chi

机构 * CMU Department of Electrical and Computer Engineering, Carnegie Mellon University(卡内基梅隆大学电气与计算机工程系) Yale Department of Statistics and Data Science, Yale University(耶鲁大学统计与数据科学系)

AI总结 本文提出黎曼扩散模型,证明在$ L_2 $-准确分数估计下,多项式小步长可保证总变差距离下的小采样误差,无需数据分布的平滑性或正性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02391 2026-01-07 cs.CL cs.SD eess.AS

WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables

WearVox: 一款面向可穿戴设备的多通道语音助手基准测试平台

Zhaojiang Lin, Yong Xu, Kai Sun, Jing Zheng, Yin Huang, Surya Teja Appini, Krish Narang, Renjie Tao, Ishan Kapil Jain, Siddhant Arora, Ruizhi Li, Yiteng Huang, Kaushik Patnaik, Wenfang Xu, Suwon Shon, Yue Liu, Ahmed A Aly, Anuj Kumar, Florian Metze, Xin Luna Dong

机构 * Meta Carnegie Mellon University(卡内基梅隆大学)

AI总结 WearVox是一个面向可穿戴设备的多通道语音助手基准测试平台,通过真实场景下的多通道音频数据评估语音助手性能,揭示了环境噪声对模型表现的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16883 2026-01-07 eess.SY cs.LG cs.SY

Myopically Verifiable Probabilistic Certificates for Safe Control and Learning

近视验证的概率证书用于安全控制与学习

Zhuoyuan Wang, Haoming Jing, Christian Kurniawan, Albert Chern, Yorie Nakahira

机构 * Department of Electrical and Computer Engineering, Carnegie Mellon University(电气与计算机工程系,卡内基梅隆大学) Information Systems Technology and Design pillar, Singapore University of Technology and Design(信息系统技术与设计支柱,新加坡科技设计大学) Department of Computer Science and Engineering, University of California San Diego(计算机科学与工程系,加州大学圣地亚哥分校)

AI总结 本文提出了一种概率不变性技术,用于设计近视验证的概率证书,以确保安全控制与学习的长期安全。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01798 2026-01-06 cs.CV cs.AI

VerLM: Explaining Face Verification Using Natural Language

通过自然语言解释面部验证

Syed Abdul Hannan, Hazim Bukhari, Thomas Cantalapiedra, Eman Ansar, Massa Baali, Rita Singh, Bhiksha Raj

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出一种视觉-语言模型用于面部验证,通过自然语言解释决策过程,提升验证的透明度和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01364 2026-01-06 cs.CV

Unsupervised SE(3) Disentanglement for in situ Macromolecular Morphology Identification from Cryo-Electron Tomography

无监督SE(3)解耦用于从冷冻电镜成像中识别原位大分子形态

Mostofa Rafid Uddin, Mahek Vora, Qifeng Wu, Muyuan Chen, Min Xu

机构 * Carnegie Mellon University(卡内基梅隆大学) Indian Institute of Technology (IIT)(印度理工学院) Division of CryoEM and Bioimaging SSRL SLAC National Accelerator Laboratory Stanford University(冷冻电镜与生物成像部SLAC国家加速器实验室斯坦福大学)

AI总结 本文提出了一种无监督的SE(3)解耦框架,用于从冷冻电镜成像中识别原位大分子形态,通过分离SE(3)变换与形态内容,提高了噪声数据下的形态识别效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23323 2026-01-06 cs.LG

LLM Interpretability with Identifiable Temporal-Instantaneous Representation

基于可识别的时间-瞬时表示的LLM可解释性

Xiangchen Song, Jiaqi Sun, Zijian Li, Yujia Zheng, Kun Zhang

机构 * Carnegie Mellon University(卡内基梅隆大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

AI总结 本文提出了一种针对LLM高维概念空间的可识别时间因果表示学习框架,通过结合SAE技术,提升了LLM的可解释性。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02698 2026-01-06 cs.LG cs.AI math.OC

Beyond Expectations: Learning with Stochastic Dominance Made Practical

超越期望:使随机支配成为可能的学习

Shicong Cen, Jincheng Mei, Hanjun Dai, Dale Schuurmans, Yuejie Chi, Bo Dai

机构 * Carnegie Mellon University(卡内基梅隆大学) Google Research(谷歌研究) Google(谷歌) University of Alberta(阿尔伯塔大学) Carnegie Mellon(卡内基梅隆大学) Georgia Tech(佐治亚理工学院)

AI总结 本文提出了一种基于随机支配的通用学习框架,通过推广随机支配概念并开发高效计算方法,实现了在多种学习任务中更优的风险权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09716 2026-01-06 cs.CV cs.AI

HCVP: Leveraging Hierarchical Contrastive Visual Prompt for Domain Generalization

HCVP:利用层次对比视觉提示进行领域泛化

Guanglin Zhou, Zhongyi Han, Shiming Chen, Biwei Huang, Liming Zhu, Tongliang Liu, Lina Yao, Kun Zhang

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) University of New South Wales(新南威尔士大学) Carnegie Mellon University(卡内基梅隆大学) University of California, San Diego(加州大学圣地亚哥分校) CSIRO’s Data61(CSIRO数据61) University of Sydney(悉尼大学) Macquarie University(麦考瑞大学)

AI总结 HCVP通过层次对比视觉提示方法提升领域泛化能力,有效分离不变特征与特定特征,优于现有算法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00160 2026-01-05 cs.SD eess.AS

IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition

IKFST: IOO和KOO算法用于加速和精确的基于WFST的端到端自动语音识别

Zhuoran Zhuang, Ye Chen, Chao Luo, Tian-Hao Zhang, Xuewei Zhang, Jian Ma, Jiatong Shi, Wei Zhang

机构 * Fliggy Alibaba(阿里巴巴飞猪) University of Science and Technology Beijing(北京科技大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 IKFST通过IOO和KOO算法优化WFST解码,提升端到端语音识别的效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.25034 2026-01-01 cs.LG cs.AI cs.CV cs.NE

Generative Classifiers Avoid Shortcut Solutions

生成分类器避免捷径解法

Alexander C. Li, Ananya Kumar, Deepak Pathak

机构 * Carnegie Mellon University(卡内基梅隆大学) Stanford University(斯坦福大学)

AI总结 生成分类器通过建模所有特征避免捷径解法,提升在分布偏移下的性能。

Comments ICLR 2025. Code: https://github.com/alexlioralexli/generative-classifiers

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17784 2026-01-01 cs.LG

AdditiveLLM: Large Language Models Predict Defects in Additive Manufacturing

AdditiveLLM:大型语言模型预测增材制造缺陷

Peter Pak, Amir Barati Farimani

机构 * Department of Mechanical Engineering, Carnegie Mellon University, Pittsburgh, PA, USA(机械工程系,卡内基梅隆大学,匹兹堡,宾夕法尼亚州,美国)

AI总结 AdditiveLLM通过自然语言输入预测增材制造缺陷,准确率达93%,简化了工艺参数优化过程。

Journal ref Additive Manufacturing Letters, Vol. 14, July 2025, Article 100292

详情

展开后加载摘要…

URL PDF HTML 收藏