arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Harvard University(哈佛大学)

共收录 140
2512.15134 2026-06-12 cs.LG cs.AI cs.CL 版本更新

From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?

从孤立到纠缠:可解释性方法何时识别和解缠已知概念?

Aaron Mueller, Andrew Lee, Shruti Joshi, Ekdeep Singh Lubana, Dhanya Sridhar, Patrik Reizinger

机构 * Boston University(波士顿大学) Harvard University(哈佛大学) Mila – Quebec AI Institute(魁北克AI研究所) Goodfire(Goodfire公司)

AI总结 本文提出多概念评估框架,研究稀疏自编码器和探针等方法是否真正解缠概念,发现特征通常只对单一概念敏感,但概念分布在多个特征上,且干预特征常影响多个概念,表明相关性指标不足以证明干预选择性。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19652 2026-06-12 cs.CV 版本更新

Navigating Gigapixel Pathology Images with Large Multimodal Models

利用大型多模态模型导航千兆像素病理图像

Thomas A. Buckley, Kian R. Weihrauch, Katherine Latham, Andrew Z. Zhou, Padmini A. Manrai, Arjun K. Manrai

机构 * Department of Biomedical Informatics, Harvard Medical School(哈佛医学院生物医学信息学系) Department of Pathology, Massachusetts General Hospital(麻省总医院病理学系) Department of Pathology and Laboratory Medicine, Brown University(布朗大学病理学与实验室医学系)

AI总结 提出GIANT方法,无需训练即可让通用多模态模型自主导航WSI,通过迭代选择多放大倍数裁剪并聚合证据,在MultiPathQA基准上实现SOTA。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16380 2026-06-12 cs.CL cs.AI cs.CY cs.HC cs.LG 版本更新

MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes

MoReBench:评估语言模型中的程序性和多元道德推理,超越结果

Yu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko, Paul de Font-Reaulx, Raphaël Millière, Paula Rodriguez, Chen Bo Calvin Zhang, Ziwen Han, Udari Madhushani Sehwag, Yash Maurya, Christina Q Knight, Harry R. Lloyd, Florence Bacus, Conor Downey, Mantas Mazeika, Bing Liu, Yejin Choi, Mitchell L Gordon, Sydney Levine

机构 * University of Washington(华盛顿大学) New York University(纽约大学) Scale AI Harvard University(哈佛大学) University of Michigan(密歇根大学) UNC Chapel Hill(北卡罗来纳大学教堂山分校) Center for AI Safety(人工智能安全中心) Stanford University(斯坦福大学) MIT(麻省理工学院) University of Oxford(牛津大学)

AI总结 提出MoReBench基准,包含1000个道德场景和超过2.3万条标准,用于评估语言模型在道德推理中的程序性推理能力,发现现有基准无法预测模型表现,且模型对特定道德框架存在偏好。

Comments 46 pages, 8 figures, 10 tables. Published in ICLR 2026. Accepted at CHAI workshop and SPP 2026 (non-archival)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00903 2026-06-12 stat.AP cs.CL cs.LG 版本更新

Causal Inference with Generative Artificial Intelligence: Application to Texts as Treatments

基于生成式人工智能的因果推断:以文本作为处理变量

Kosuke Imai, Kentaro Nakamura

机构 * Harvard University(哈佛大学) John F. Kennedy School of Government(约翰·F·肯尼迪政府学院)

AI总结 提出利用生成式AI(如大语言模型)生成处理变量并利用其内部表示进行因果效应估计,避免从数据中学习因果表示,提高估计准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04021 2026-06-12 cs.DC cs.AI cs.LG cs.PF 版本更新

Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning

Prism: 通过GPU内存气球实现经济高效的多LLM服务

Shan Yu, Yifan Qiao, Mingyuan Ma, Yangmin Li, Shuo Yang, Xinyuan Tong, Yang Wang, Zhiqiang Xie, Yuwei An, Shiyi Cao, Ke Bao, Deepak Vij, Xiaoning Ding, Yichen Wang, Qingda Lu, Zhong Wang, Gao Gao, Harry Xu, Junyi Shu, Jiarong Xing, Ying Sheng

机构 * UCLA(加州大学洛杉矶分校) UC Berkeley(伯克利加州大学) Harvard University(哈佛大学) CMU(卡内基梅隆大学) University of Edinburgh(爱丁堡大学) Intel(英特尔) Stanford University(斯坦福大学) LMSYS(灵州市系统实验室) ByteDance(字节跳动) Alibaba Cloud(阿里云) Tsinghua University(清华大学) Novita AI Rice University(里士满大学)

AI总结 针对多LLM服务中资源效率低下的问题,提出基于内存气球的内存中心化LLM协同服务框架Prism,统一空间与时间共享,已在10K+ GPU生产环境部署。

Comments OSDI'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05573 2026-06-11 cs.LG 版本更新

Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View

为什么深度在可并行化序列模型中重要:一个李代数视角

Gyuryang Heo, Timothy Ngotiaoco, Kazuki Irie, Samuel J. Gershman, Bernardo L. Sabatini

机构 * Howard Hughes Medical Institute, Department of Neurobiology, Harvard Medical School(霍华德·休斯医学研究所,哈佛医学院神经生物学系) Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(自然与人工智能研究学院,哈佛大学) Department of Psychology and Center for Brain Science, Harvard University(心理学系和脑科学中心,哈佛大学)

AI总结 从李代数控制视角,研究可并行化序列模型(如Transformer变体和状态空间模型)的表达能力与深度关系,证明误差随深度增加呈指数下降。

Comments v2: Format update; split former Theorem 3.4 into Theorem 3.4 and Corollary 3.5 for clarity; corrected an indexing error affecting Corollary 3.6, Proposition 3.7, and Figure 2

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09809 2026-06-10 cs.AI 版本更新

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

评估卡:AI评估报告的解释层

Avijit Ghosh, Anka Reuel, Jenny Chim, Wm. Matthew Kennedy, Srishti Yadav, Jennifer Mickel, Yanan Long, Andrew Tran, Anastassia Kornilova, Damian Stachura, Kevin Klyman, Felix Friedrich, Jeba Sania, Jan Batzner, Anoop Mishra, Eliya Habba, Yixiong Hao, Nathan Heath, Shalaleh Rismani, Usman Gohar, Andrea Loehr, David Manheim, Ruchira Dhar, Sree Harsha Nelaturu, Aarush Sinha, Leshem Choshen, Drishti Sharma, Ishan Khire, Amit Saha, Subramanyam Sahoo, Michael Hardy, Michael Alexander Riegler, Kabir Manghnani, Michelle Lin, Yanan Jiang, Yilin Huang, Asaf Yehudai, Jessica Ji, Aris Hofmann, Mubashara Akhtar, Max Lamparth, Nuno Moniz, Yacine Jernite, Stella Biderman, Zeerak Talat, Sanmi Koyejo, Mykel Kochenderfer, Irene Solaiman

机构 * Hugging Face Stanford University(斯坦福大学) Queen Mary University of London(伦敦玛丽女王大学) University of Copenhagen(哥本哈根大学) Trustible EleutherAI TU Darmstadt(达姆施塔特工业大学) Weizenbaum Institute & Technical University of Munich(魏森鲍姆研究所与慕尼黑工业大学) Harvard University(哈佛大学) The Hebrew University of Jerusalem(耶路撒冷希伯来大学) Iowa State University(爱荷华州立大学) IBM Research(IBM研究院) University of Chicago(芝加哥大学) Independent(独立) Berkeley AI Safety Institute (BASIS)(伯克利人工智能安全研究所) Simula University of Edinburgh(爱丁堡大学) ETH Zurich & ETH AI Center(苏黎世联邦理工学院与ETH AI中心) Oxford Internet Institute(牛津互联网研究所) Amherst College(阿默斯特学院) University of Nebraska(内布拉斯加大学) Syntony Research McGill University(麦吉尔大学) Evals Consensus Israel Institute of Technology(以色列理工学院) IOL.Learn & Zuse Institute Berlin(IOL.Learn与柏林祖泽研究所) Georgia Institute of Technology(佐治亚理工学院) Quebec AI Institute, Université de Montréal(魁北克人工智能研究所,蒙特利尔大学) University of Notre Dame(圣母大学) Georgetown University(乔治城大学) DHBW Stuttgart(斯图加特双元制大学) Massachusetts Institute of Technology(麻省理工学院)

AI总结 针对AI评估报告不一致的问题,提出EvalCards作为统一记录层,通过结构化模式、四种解释信号和监控工具,覆盖5816个模型和635个基准,揭示报告实践中的系统性差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17251 2026-06-10 stat.ML cs.LG 版本更新

Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization

线性回归中的风险比较:隐式正则化主导显式正则化

Jingfeng Wu, Peter L. Bartlett, Sham M. Kakade, Jason D. Lee, Bin Yu

机构 * University of California, Berkeley(加州大学伯克利分校) Alphabetical order Harvard University(哈佛大学) Google DeepMind(谷歌DeepMind)

AI总结 本文通过实例比较线性回归中梯度下降、岭回归和随机梯度下降的有限样本风险,发现梯度下降优于岭回归,但与随机梯度下降不可比,且在某些问题中梯度下降可能更差。

Comments Accepted for presentation at the Conference on Learning Theory (COLT) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01616 2026-06-09 cs.LG cs.AI cs.CY cs.NI 版本更新

Learning Behavioral Signals from Encrypted Smartphone Network Traffic

从加密智能手机网络流量中学习行为信号

Rameen Mahmood, Omar El Shahawy, Souptik Barua, Zachary Beattie, Jeffrey Kaye, Xuhai "Orson'' Xu, Chao-Yi Wu, Danny Yuxing Huang

机构 * New York University(纽约大学) NYU Langone Health(NYU Langone健康) NYU Grossman School of Medicine(NYU Grossman医学院) Oregon Health & Science University(俄勒冈健康与科学大学) Columbia University(哥伦比亚大学) Harvard Medical School(哈佛医学院)

AI总结 本文利用基于Transformer的模型从加密网络流量中学习行为表征,结合用户特定适配器,并通过稀疏表示和广义估计方程分析,发现压力、孤独感和睡眠障碍分别与个体间差异、个体内波动及两者组合相关,且学习到的表征优于传统手工特征。

Comments 19 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06407 2026-06-09 cs.CV cs.IR cs.LG eess.IV 版本更新

A Vision-language Framework for Comparative Reasoning in Radiology

放射学中比较推理的视觉语言框架

Tengfei Zhang, Ziheng Zhao, Xiaoman Zhang, Lisong Dai, Pengcheng Qiu, Ya Zhang, Yanfeng Wang, Weidi Xie

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) Department of Biomedical Informatics, Harvard Medical School(哈佛医学院生物医学信息学系) Department of Radiology, Renmin Hospital of Wuhan University(武汉大学仁民医院放射科) Shanghai Sixth People’s Hospital Affiliated to Shanghai Jiao Tong University(上海交通大学附属第六人民医院)

AI总结 提出一个实体感知的跨图像推理框架,通过构建大规模比较影像数据集MedReCo-DB和开发MedReCo及MedReCo-VLM模型,实现了参考病例检索和时间比较解读,显著提升了放射学比较推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24942 2026-06-09 cs.LG cs.AI 版本更新

Riemannian-Manifold Steering: Geometry-Aware Generative Autoencoders for Label-Free Steering

黎曼流形操控:用于无标签操控的几何感知生成自编码器

Narmeen Oozeer, Shivam Raval, Philip Quirke, Manikandan Ravikiran, Jeff Phillips, Shriyash Upadhyay, Amirali Abdullah

机构 * Martian Harvard University(哈佛大学) Thoughtworks University of Utah(犹他大学)

AI总结 提出将语言模型操控重新定义为激活空间上的黎曼测地线计算,通过基于输出空间Hellinger距离学习的编码器实现无标签、无拓扑先验的流形操控。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09787 2026-06-09 astro-ph.IM astro-ph.GA cs.LG 版本更新

Learning What's Real: Disentangling Signal and Measurement Artifacts in Multi-Sensor Data, with Applications to Astrophysics

学习真实内容:在多传感器数据中分离信号和测量伪影,应用于天体物理学

Pablo Mercader-Perez, Carolina Cuesta-Lazaro, Daniel Muthukrishna, Jeroen Audenaert, V. Ashley Villar, David W. Hogg, Marc Huertas-Company, William T. Freeman

机构 * Massachusetts Institute of Technology(麻省理工学院) Flatiron Institute, Simons Foundation(Flatiron研究所,Simons基金会) Institute for Advanced Studies(高级研究 institute) Harvard University(哈佛大学) New York University(纽约大学) Instituto de Astrofísica de Canarias(加那利大天文台)

AI总结 本文提出一种深度学习框架,通过重叠观测、双编码器架构和反事实生成目标,分离多传感器数据中的信号与伪影,提升天体物理学研究的准确性。

Comments Accepted at the 2nd Workshop on Foundation Models for Science at ICLR 2026. 10 pages, 7 figures (main text), plus appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15327 2026-06-09 cs.LG cs.AI cs.CL stat.ML 版本更新

Prescriptive Scaling Reveals the Evolution of Language Model Capabilities

规范性缩放揭示语言模型能力的演变

Hanlin Zhang, Jikai Jin, Vasilis Syrgkanis, Sham Kakade

机构 * Harvard University(哈佛大学) Stanford University(斯坦福大学)

AI总结 通过大规模观测评估和分位数回归,提出规范性缩放定律,将预训练计算预算映射到下游准确率,并验证其时间稳定性,引入平衡I-最优采样算法降低评估成本。

Comments ICML 2026 Oral. Blog Post: https://jkjin.com/prescriptive-scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07355 2026-06-09 cs.AI cs.CV cs.LG 版本更新

A Geometric Unification of Concept Learning with Concept Cones

概念学习与概念锥的几何统一

Alexandre Rocchi, Thomas Fel, Gianni Franchi

机构 * AMIAD Kempner Institute, Harvard University(哈佛大学凯普勒研究所)

AI总结 通过共享几何框架(概念锥)统一监督式概念瓶颈模型与无监督稀疏自编码器,提出包含关系度量评估概念对齐,并发现稀疏性与扩展因子的最佳平衡点。

Comments 33 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08908 2026-06-09 math.ST cs.LG econ.TH stat.TH 版本更新

Statistical Decision Theory with Counterfactual Loss

具有反事实损失的统计决策理论

Benedikt Koch, Kosuke Imai

机构 * Harvard University(哈佛大学)

AI总结 针对经典统计决策理论忽略反事实信息的问题,提出在强可忽略性下反事实风险可识别当且仅当损失函数在潜在结果上可加,并证明可加反事实损失能捕捉决策难度,通过符号线性逆规划无需数据即可判断可识别性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01916 2026-06-09 cs.LG 版本更新

Causal Representation Learning from Network Data

从网络数据中进行因果表示学习

Jifan Zhang, Michelle M. Li, Elena Zheleva

机构 * Department of Statistics and Data Science, Northwestern University(统计与数据科学系,西北大学) Department of Biomedical Informatics, Harvard University(生物医学信息学系,哈佛大学) Department of Computer Science, University of Illinois Chicago(计算机科学系,伊利诺伊大学芝加哥分校)

AI总结 提出GraCE-VAE,利用图神经网络编码器整合生物网络和通路信息,在软干预下识别潜在因果图与干预目标,实验证明利用结构化生物上下文可提升干预结果预测。

Comments 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06469 2026-06-09 cs.CL 版本更新

ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?

ClinicalBench: 大型语言模型能在临床预测中击败传统机器学习模型吗?

Canyu Chen, Jian Yu, Shan Chen, Che Liu, Zhongwei Wan, Shuang Zhou, Yuan Luo, Rui Zhang, Danielle Bitterman, Fei Wang, Kai Shu

机构 * Department of Computer Science Northwestern University Evanston USA(计算机科学系西北大学艾文斯顿美国) Department of Computer Science University of Texas at Austin Austin USA(计算机科学系德克萨斯大学奥斯汀美国) Boston Children's Hospital, Harvard Medical School Boston USA(波士顿儿童医院哈佛医学院波士顿美国) Department of Computer Science Imperial College London London UK(计算机科学系伦敦帝国学院伦敦英国) Department of Computer Science Ohio State University Columbus USA(计算机科学系俄亥俄州立大学哥伦布美国) Massachusetts General Hospital, Harvard Medical School Boston USA(麻省总医院哈佛医学院波士顿美国) Department of Preventive Medicine, Feinberg School of Medicine Northwestern University Chicago USA(预防医学系费因伯格医学院西北大学芝加哥美国) Division of Computational Health Sciences, Department of Surgery University of Minnesota Minneapolis USA(计算健康科学部外科部明尼苏达大学明尼阿波利斯美国) Department of Population Health Sciences, Weill Cornell Medicine Cornell University New York USA(流行病学与公共卫生系韦尔·科恩医学中心康奈尔大学纽约美国) Department of Computer Science Emory University Atlanta USA(计算机科学系埃默里大学亚特兰大美国) Northwestern University(西北大学) University of Texas at Austin(德克萨斯大学奥斯汀) Boston Children's Hospital, Harvard Medical School(波士顿儿童医院哈佛医学院) Imperial College London(伦敦帝国学院) Ohio State University(俄亥俄州立大学) Massachusetts General Hospital, Harvard Medical School(麻省总医院哈佛医学院) University of Minnesota(明尼苏达大学) Cornell University(康奈尔大学) Emory University(埃默里大学)

AI总结 构建ClinicalBench基准,通过三个临床预测任务比较14个通用和8个医学LLM与11个传统ML模型,发现LLM在临床预测上仍无法超越传统ML模型。

Comments Accepted to Proceedings of KDD 2026. The first two authors contributed equally. 12 pages for main paper, 62 pages including appendix. Project website: https://clinicalbench.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01718 2026-06-09 stat.ML cs.LG math.ST stat.TH 版本更新

Entropic Optimal Transport Eigenmaps for Nonlinear Alignment and Joint Embedding of High-Dimensional Datasets

熵最优传输特征映射用于高维数据集的非线性对齐与联合嵌入

Boris Landa, Yuval Kluger, Rong Ma

机构 * Department of Electrical and Computer Engineering, Yale University(耶鲁大学电气与计算机工程系) Department of Biostatistics, Harvard University(哈佛大学生物统计学系) Program in Applied Mathematics, Yale University(耶鲁大学应用数学项目) Interdepartmental Program in Computational Biology and Bioinformatics, Yale University(耶鲁大学计算生物学与生物信息学跨学科项目) Department of Pathology, Yale University School of Medicine(耶鲁大学医学院病理学系)

AI总结 提出熵最优传输特征映射方法,通过EOT计划矩阵的奇异向量对齐和联合嵌入两个数据集,具有理论保证,在生成模型下证明其收敛性,并在模拟和真实生物数据中展示优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.10277 2026-06-09 cs.CL cs.LG cs.SI 版本更新

Measuring a hate speech spectrum with faceted Rasch item response theory and perspective-aware, explainable-by-design deep learning

使用分面Rasch项目反应理论和可解释性设计的深度学习测量仇恨言论谱系

Chris J. Kennedy, Geoff Bacon, Alexander Sahn, Claudia von Vacano

机构 * Center for Precision Psychiatry, Mass General Hospital Department of Psychiatry, Harvard Medical School(精准精神病学中心,麻省总医院精神病科,哈佛医学院) D-Lab University of California, Berkeley(加州大学伯克利分校D实验室)

AI总结 提出结合监督深度学习与分面Rasch项目反应理论的方法,将仇恨言论分解为10个有序标签,通过IRT模型转化为区间测量值并调整标注者视角,在RoBERTa模型上提升准确性,实现连续谱系测量与可解释性。

Comments 7 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10102 2026-06-08 cs.RO cs.SY eess.SY 版本更新

A Human-Sensitive Controller: Adapting to Human Musculoskeletal Disorder-Related Constraints via Reinforcement Learning

一种人类敏感控制器:通过强化学习适应人类肌肉骨骼疾病相关约束

Vitor Martins, Sara M. Cerqueira, Mercedes Balcells, Elazer R Edelman, Cristina P. Santos

机构 * Fundação para a Ciência e Tecnologia(葡萄牙科学与技术基金会) Centro de Microssistemas Eletromecânicos da Universidade do Minho(University of Minho微机电系统中心) Massachusetts Institute of Technology(麻省理工学院) Brigham and Women’s Hospital, Harvard Medical School(哈佛医学院布莱尔妇女医院) GEVAB, IQS School of Engineering(GEVAB,IQS工程学院) LABBELS-Associate Laboratory, University of Minho(University of Minho关联实验室)

AI总结 提出基于强化学习的人类敏感机器人控制策略,使用Q学习和深度Q网络优化协作机器人的人机工效,在保持零疼痛风险下平均缩短38%任务完成时间。

详情

展开后加载摘要…

URL PDF HTML 收藏