arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Carnegie Mellon University(卡内基梅隆大学)

2026-05-14 至 2026-05-14 共收录 16
2605.13706 2026-05-14 cs.CR cs.AI cs.CY cs.NI

Identifying AI Web Scrapers Using Canary Tokens

通过信标令牌识别人工智能网络爬虫

Steven Seiden, Triss Ren, Caroline Zhang, Taein Kim, Enze Liu, Emily Wenger

机构 * Duke University(杜克大学) University of Pittsburgh(匹兹堡大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出一种自动识别与大型语言模型相关的网络爬虫的方法,通过动态网站和信标令牌验证爬虫身份,实验表明能可靠识别多个未公开的爬虫。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24649 2026-05-14 cs.CV

MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows

MedOpenClaw 和 MedFlowBench:在完整研究流程中审计医疗代理

Weixiang Shen, Chengzhi Shen, Yanzhu Hu, Che Liu, Junde Wu, Jiayuan Zhu, Xiao Han, Zongyue Li, Jingpei Wu, Min Xu, Daguang Xu, Yueming Jin, Benedikt Wiestler, Daniel Rueckert, Jiazhen Pan

机构 * Technical University of Munich(慕尼黑技术大学) TUM University Hospital(TUM大学医院) LMU Munich(慕尼黑大学) Imperial College London(伦敦帝国理工学院) University of Oxford(牛津大学) Carnegie Mellon University(卡内基梅隆大学) NVIDIA(NVIDIA公司) National University of Singapore(新加坡国立大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

AI总结 本文提出 MedFlowBench 和 MedOpenClaw,用于评估医疗影像代理在完整研究流程中生成可审计证据的能力,发现仅依赖答案评分不够,需结合正确证据验证。

Comments 33 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15280 2026-05-14 cs.HC cs.AI

LLM-based Multimodal Feedback Produces Equivalent Learning and Better Student Perceptions than Educator Feedback

基于大语言模型的多模态反馈在学习效果和学生感知上与教师反馈相当且更优

Chloe Qianhui Zhao, Jie Cao, Jionghao Lin, Kenneth R. Koedinger

机构 * Carnegie Mellon University(卡内基梅隆大学) The University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) The University of Hong Kong(香港大学)

AI总结 本文提出一种实时AI辅助的多模态反馈系统,通过整合结构化文本解释与动态多媒体资源,实现学习效果与学生感知的提升,证明AI反馈在清晰度、具体性等方面优于传统教师反馈。

Comments 11 pages, to be published at the 16th International Learning Analytics & Knowledge Conference (LAK '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12937 2026-05-14 cs.CV cs.AI cs.HC

AuraMask: An Extensible Pipeline for Developing Aesthetic Anti-Facial Recognition Image Filters

AuraMask:一种可扩展的开发美观反面部识别图像滤镜的管道

Jacob Lagogiannis, William Agnew, Rosa I. Arriaga, Sauvik Das

机构 * Franklin and Marshall College(弗兰克林与马歇尔学院) Carnegie Mellon University(卡内基梅隆大学) Georgia Institute of Technology(佐治亚理工学院)

AI总结 本文提出AuraMask,一种生成有效且美观的反面部识别图像滤镜的方法,通过40个仿Instagram滤镜的实例,验证其在对抗有效性与用户接受度上的优势。

Comments 21 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12904 2026-05-14 cs.LG

VIP-COP: Context Optimization for Tabular Foundation Models

VIP-COP:表格基础模型的上下文优化

Yilong Chen, Xueying Ding, Leman Akoglu

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 VIP-COP通过显式选择机制优化上下文,提升表格基础模型性能,具备快速、预算感知、黑盒兼容和可解释性等优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12798 2026-05-14 cs.LG cs.AI cs.CL

Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer

通过数据中介转移的视角看涌现与潜意识偏差

Baris Askin, Muhammed Ustaomeroglu, Anupam Nayak, Gauri Joshi, Guannan Qu, Carlee Joe-Wong

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 研究探讨了通过数据中介转移现象理解模型在狭窄有害数据集上微调时产生的偏差问题,分析了微调数据结构、预训练分布和训练渠道之间的相互作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12788 2026-05-14 cs.LG cs.CY

From Heuristics to Analytics: Forecasting Effort and Progress in Online Learning

从启发式到分析:在线学习中的努力和进度预测

Eric S. Qiu, Danielle R. Thomas, Boyuan Guo, Vincent Aleven, Conrad Borchers

机构 * Cornell University(康奈尔大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文通过ITS日志数据,利用监督学习方法预测学生每周练习时间和掌握新技能情况,展示基于特征的模型在降低MAE方面优于启发式基线,同时分析了预测特征的重要性,为 tutor-learner 目标设定提供支持。

Comments Accepted as full paper to the 19th International Conference on Educational Data Mining (EDM 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12745 2026-05-14 cs.HC cs.AI

What Do You Think I Think? Accounting for Human Beliefs Using Second-Order Theory of Mind

我思考什么?利用二级理论 of mind 账户人类信念

Patrick Callaghan, Reid Simmons, Henny Admoni

机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)

AI总结 本文提出一种二级理论 of mind 框架,使智能体能检测并反馈人类错误信念及认知偏差,从而提升交互信息量。

Comments To appear in the proceedings of The 2026 Cognitive Science Society Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12733 2026-05-14 cs.LG cs.AI stat.ML

From Generalist to Specialist Representation

从通用到专业表征

Yujia Zheng, Fan Feng, Yuke Li, Shaoan Xie, Kevin Murphy, Kun Zhang

机构 * CMU(卡内基梅隆大学) UIUC(伊利诺伊大学香槟分校) UCSD(加州大学圣地亚哥分校) MBZUAI(穆斯林人工智能研究所) UMD(马里兰大学) UBC(不列颠哥伦比亚大学)

AI总结 研究在非参数设置下如何通过非监督方式识别时间步间任务结构,并利用稀疏正则化分离任务相关潜在表征,为从通用到专业模型的理论保障提供新思路。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12684 2026-05-14 cs.CV cs.AI cs.HC

Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?

视觉审美基准:前沿模型能评判美吗?

Yichen Feng, Yuetai Li, Chunjiang Liu, Yuanyuan Chen, Fengqing Jiang, Yue Huang, Hang Hua, Zhengqing Yuan, Kaiyuan Zheng, Luyao Niu, Bhaskar Ramasubramanian, Basel Alomair, Xiangliang Zhang, Misha Sra, Zichen Chen, Radha Poovendran, Zhangchen Xu

机构 * Bake AI University of Washington(华盛顿大学) University of California, Santa Barbara(加州大学圣巴巴拉分校) Stanford University(斯坦福大学) University of Notre Dame(诺丁汉大学) Carnegie Mellon University(卡内基梅隆大学) MIT-IBM Watson AI Lab(麻省理工-IBM沃森人工智能实验室) Western Washington University(西雅图华盛顿大学) King Abdulaziz City for Science and Technology(国王阿卜杜勒阿齐兹科技城)

AI总结 本文提出视觉审美基准(VAB),通过比较选择评估审美,发现前沿模型在判断最佳和最差图像时表现逊于人类专家,表明需进一步改进。

Comments Project page: https://vab.bakelab.ai. Code: https://github.com/BakeLab/Visual-Aesthetic-Benchmark. Dataset: https://huggingface.co/datasets/BakeLab/Visual-Aesthetic-Benchmark

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10906 2026-05-14 cs.LG cs.AI

DataMaster: Data-Centric Autonomous AI Research

DataMaster: 以数据为中心的自主AI研究

Yaxin Du, Xiyuan Yang, Zhifan Zhou, Wanxu Liu, Zixing Lei, Zimeng Chen, Fenyi Liu, Haotian Wu, Yuzhu Cai, Zexi Liu, Xinyu Zhu, WenHao Wang, Linfeng Zhang, Chen Qian, Siheng Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Carnegie Mellon University(卡内基梅隆大学) Zhejiang University(浙江大学) Beijing University of Aeronautics and Astronautics(北京航空航天大学)

AI总结 DataMaster通过自主数据工程优化数据侧,提升下游性能,无需改变学习算法。其包含数据树、共享数据池和全局记忆三大组件,有效解决开放搜索空间和延迟验证问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22474 2026-05-14 cs.RO cs.LG

When to Act, Ask, or Learn: Uncertainty-Aware Policy Steering

何时行动、提问或学习:不确定性感知的策略引导

Jessie Yuan, Yilin Wu, Andrea Bajcsy

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出不确定性感知策略引导框架,通过联合推理任务语义不确定性和低层动作可行性,选择执行高置信度动作、通过自然语言查询澄清任务歧义或请求动作干预修正低层策略,提升机器人部署性能。

Comments To appear in Robotics: Science and Systems 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11202 2026-05-14 cs.LO cs.AI

interwhen: A Generalizable Framework for Steering Reasoning Models with Test-time Verification

interwhen:一种可泛化的引导推理模型进行测试时验证的框架

Vishak K Bhat, Prateek Chanda, Vijval Ekbote, Ashmit Khandelwal, Maitreyi Swaroop, Vineeth N. Balasubramanian, Subbarao Kambhampati, Nagarajan Natarajan, Amit Sharma

机构 * Microsoft Research(微软研究院) IIT Bombay(印度理工学院班加罗尔分校) Carnegie Mellon University(卡内基梅隆大学) Arizona State University(亚利桑那州立大学)

AI总结 interwhen提出一种单轨迹验证框架,通过反馈中间推理轨迹引导模型行为,解决验证状态获取和 verifier 缺乏的问题,提升推理代理的任务完成和政策合规性。

Comments 56 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22792 2026-05-14 eess.AS cs.CL cs.SD

CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR

CALM:多说话人ASR的联合上下文语音-语言建模

Muhammad Shakeel, Yosuke Fukumoto, Chikara Maeda, Chyi-Jiunn Lin, Shinji Watanabe

机构 * Honda Research Institute Japan Co., Ltd.(本田研究院日本株式会社) Carnegie Mellon University(卡内基梅隆大学)

AI总结 CALM通过端到端框架整合目标说话人条件和上下文偏差,降低多说话人语音识别的词错误率和字符错误率,验证了跨语言的联合语音-语言建模有效性。

Comments Accepted to IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20474 2026-05-14 eess.AS cs.CL cs.SD

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder

统一语音识别、分离和ASR的多说话人编码器

Muhammad Shakeel, Yui Sudo, Yifan Peng, Chyi-Jiunn Lin, Shinji Watanabe

机构 * Honda Research Institute Japan, Japan(本田研究院日本) Carnegie Mellon University, USA(卡内基梅隆大学)

AI总结 本文提出统一多说话人编码器(UME),通过共享语音基础编码器联合学习说话人识别、语音分离和多说话人ASR任务的表示,提升重叠语音数据性能。

Comments Accepted to IEEE ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11556 2026-05-14 cs.CL cs.AI cs.MA

Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMs

多智能体大语言模型中分布式信息下的集体推理系统性失效

Yuxuan Li, Aoi Naito, Hirokazu Shirado

机构 * School of Computer Science, Carnegie Mellon University, Pittsburgh, USA(计算机科学学院,卡内基梅隆大学,匹兹堡,美国) School of Environment and Society, Institute of Science Tokyo, Tokyo, Japan(环境与社会学院,东京科学研究所,东京,日本)

AI总结 研究发现多智能体大语言模型在分布式信息下仅能获得30.1%的准确率,而单智能体在完整信息下可达80.7%。系统性失效源于智能体无法识别和应对潜在的信息不对称,导致关键分布式事实未被探索。通过轻量级结构化通信协议可显著提升集体推理能力。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏