arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Columbia University(哥伦比亚大学)

2026-06-30 至 2026-06-30 共收录 8
2606.30491 2026-06-30 cs.CL cs.AI

SIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation

SIMAX: 一种可扩展且可解释的多保真度标注医患对话模拟框架

Zhuhan Bao, Rui Yang, Bohao Yang, Zhiyi Liu, Sicheng Shu, Ruio Heerschap, Le Li, Doris Yang, Elisabeth Bond, Haoyuan Wang, Nicoleta Economou-Zavlanos, Joshua M. Biro, Matthew McDermott, Nan Liu, Anand Chowdhury, Kai Sun, Kathryn Pollak, Ed Hammond, Chuan Hong

机构 * Department of Biostatistics and Bioinformatics, Duke University School of Medicine(杜克大学医学学院生物统计学与生物信息学系) Duke-NUS AI + Medical Sciences Initiative, Duke-NUS Medical School(杜克-新加坡国立大学医学科学院AI+医学科学计划) Centre for Biomedical Data Science, Duke-NUS Medical School(杜克-新加坡国立大学医学学院生物医学数据科学中心) Department of Statistical Science, Duke University(杜克大学统计科学系) Leiden University Medical Centre(莱顿大学医学中心) Department of Mathematics, University of Texas at Austin(德克萨斯大学奥斯汀分校数学系) Department of Internal Medicine, Yale School of Medicine(耶鲁医学院内科学系) Department of Biostatistics, Epidemiology and Informatics, Perelman School of Medicine, University of Pennsylvania(宾夕法尼亚大学佩尔曼医学院生物统计学、流行病学与信息学系) The Graduate Group in Applied Mathematics and Computational Science, School of Arts and Sciences, University of Pennsylvania(宾夕法尼亚大学艺术与科学学院应用数学与计算科学联合组) Medstar Health National Center for Human Factors in Healthcare, Washington, DC, USA(Medstar健康国家人因工程中心,华盛顿特区,美国) Department of Biomedical Informatics, Columbia University(哥伦比亚大学生物医学信息学系) Cancer Prevention and Control, Duke Cancer Institute, Durham, NC, USA(杜克癌症研究所癌症预防与控制部,达勒姆,北卡罗来纳州,美国) Department of Population Health Sciences, Duke University School of Medicine(杜克大学医学学院流行病学与公共卫生系) Division of Rheumatology and Immunology, Duke University School of Medicine(杜克大学医学学院风湿病学与免疫学系) Pre-hospital and Emergency Research Centre, Health Services Research and Population Health, Duke-NUS Medical School(杜克-新加坡国立大学医学学院院前急救与应急研究中心,健康服务研究与人口健康) NUS Artificial Intelligence Institute, National University of Singapore(新加坡国立大学人工智能研究所) Division of Pulmonary, Allergy and Critical Care Medicine, Duke University School of Medicine(杜克大学医学学院呼吸科、过敏科与危重医学系) Duke Center for Health Informatics, Duke University(杜克大学健康信息学中心) Duke Clinical Research Institute, Durham, NC, USA(杜克临床研究中心,达勒姆,北卡罗来纳州,美国)

AI总结 提出SIMAX框架,通过预定义场景、角色和沟通行为生成可控医患对话,自动评估显示语音自然度和转录保真度良好,可用于开发和验证沟通编码系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29049 2026-06-30 cs.LG

MOSAIC: Orchestrating Collaborative Knowledge Tracing with Hierarchical Semantic Alignment

MOSAIC: 通过层次语义对齐编排协作知识追踪

Xinjin Li, Mengyue Wang, Yuzhen Lin, Pengbin Feng, Ziqi Sha, Yeyang Zhou, Yu Ma

机构 * Columbia University(哥伦比亚大学) University of California, Berkeley(加州大学伯克利分校) School of Information Systems and Management, Carnegie Mellon University(信息系统与管理学院,卡内基梅隆大学) Department of Mathematics, University of Southern California(数学系,南加州大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Computer Science Department, UC San Diego(计算机科学系,UCSD)

AI总结 提出MOSAIC框架,利用冻结LLM生成动态嵌入和层次预测提示,结合跨粒度一致性目标,在协作知识追踪中实现多粒度掌握估计,在多个数据集上取得SOTA。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28573 2026-06-30 cs.LG math.ST stat.TH

Replica Symmetry Breaking and Algorithmic Thresholds in Empirical Risk Minimization under Multi-Index Model

多指标模型下经验风险最小化的副本对称破缺与算法阈值

Andrea Montanari, Kangjie Zhou

机构 * Department of Mathematics and Department of Statistics, Stanford University(数学系和统计系,斯坦福大学) Department of Statistics, Columbia University(统计系,哥伦比亚大学)

AI总结 研究高维非凸经验风险最小化中多项式时间算法可达的优化区域,提出增量近似消息传递算法并刻画其训练误差及泛化误差,在渐近分析中证明算法性能最优。

Comments 80 pages; 3 pdf figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08691 2026-06-30 cs.LG stat.ME 新提交

Hierarchical Projection for Adaptive Knowledge Transfer

自适应知识迁移的分层投影

Samhita Pal, Tian Gu

机构 * Vanderbilt University Medical Center(范德比尔特大学医学中心) Columbia University(哥伦比亚大学)

AI总结 提出ProjectionTL框架,通过分层贝叶斯建模与自适应投影实现源选择与特征选择,缓解负迁移,提升跨域学习的准确性、稳定性和可解释性。

Comments We found a mistake in the proof that needs to be revised

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09160 2026-06-30 cs.LG

Objective-Specific Privileged Bases via Full-Prefix Matryoshka Learning

面向目标的特权基通过全前缀马特罗什卡学习

Arghamitra Talukder, Philippe Chlenski, Itsik Pe'er

机构 * Computer Science, Columbia University(哥伦比亚大学计算机科学系)

AI总结 本文研究马特罗什卡表示学习如何诱导与任务对齐的特权基,证明全前缀MRL能高效恢复有序主方向,且坐标大小反映信息量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13932 2026-06-30 cs.SE cs.AI

Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action

代码推理用于软件工程任务:调查与呼吁行动

Saurabh Pujar, Ira Ceka, Irene Manotas, Gail Kaiser, Baishakhi Ray, Shyam Ramji

机构 * IBM Columbia University(哥伦比亚大学)

AI总结 本文调查代码推理技术,探讨其在软件工程任务中的影响,提出未来研究方向。

Comments Published in Transactions on Machine Learning Research (06/2026) 40 pages, 8 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23220 2026-06-30 cs.CL cs.LG

Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders

模型方向,而非词语:使用稀疏自编码器的机制性主题模型

Carolina Zheng, Nicolas Beltran-Velez, Sweta Karlekar, Claudia Shi, Achille Nazaret, Asif Mallik, Amir Feder, David M. Blei

机构 * Columbia University(哥伦比亚大学) Google Research(谷歌研究院) Independent(独立研究者)

AI总结 本文提出机制性主题模型(MTMs),利用稀疏自编码器学习可解释特征,以揭示深层概念主题,并通过topic judge评估框架验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08775 2026-06-30 cs.HC cs.AI

What Should We Engineer in Prompts? Training Humans in Requirement-Driven LLM Use

提示中应工程什么?基于需求驱动的LLM使用训练

Qianou Ma, Weirui Peng, Chenyang Yang, Hua Shen, Kenneth Koedinger, Tongshuang Wu

机构 * Carnegie Mellon University(卡内基梅隆大学) Columbia University(哥伦比亚大学) University of Washington(华盛顿大学)

AI总结 本文提出ROPE框架,通过评估和训练套件帮助用户生成清晰需求,提升LLM应用构建效果。

Comments 15 pages; TOCHI 2025

Journal ref TOCHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏