arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Columbia University(哥伦比亚大学)

2026-06-12 至 2026-06-12 共收录 3
2603.25450 2026-06-12 cs.AI 版本更新

Cross-Model Disagreement as a Label-Free Correctness Signal

跨模型分歧作为无标签正确性信号

Matt Gorbett, Suman Jana

机构 * Independent Researcher(独立研究者) Department of Computer Science Columbia University(计算机科学系哥伦比亚大学)

AI总结 提出跨模型分歧作为无标签正确性指标,通过验证模型对生成模型答案的困惑度或熵来检测错误,无需训练或标签,在多个基准上优于模型内不确定性方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22090 2026-06-12 cs.RO 版本更新

ReactEMG Stroke: Healthy-to-Stroke Few-shot Adaptation for sEMG-Based Intent Detection

ReactEMG 中风:基于表面肌电图的意图检测的健康到中风少样本适应

Runsheng Wang, Katelyn Lee, Xinyue Zhu, Lauren Winterbottom, Dawn M. Nilsen, Joel Stein, Matei Ciocarlie

机构 * Department of Mechanical Engineering, Columbia University in the City of New York(哥伦比亚大学纽约市机械工程系) Department of Computer Science, Columbia University in the City of New York(哥伦比亚大学纽约市计算机科学系) Department of Rehabilitation and Regenerative Medicine, Columbia University Irving Medical Center(哥伦比亚大学伊文思医疗中心康复与再生医学系)

AI总结 提出一种健康到中风的适应流程,利用大规模健康受试者sEMG预训练模型,仅用少量中风患者数据微调,显著提升意图检测准确率和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20208 2026-06-12 cs.CL 版本更新

From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation

从基准到技能:LLM评估的低秩因子

Aviya Maimon, Amir DN Cohen, Gal Vishne, Shauli Ravfogel, Reut Tsarfaty

机构 * Bar-Ilan University(巴伊兰大学) OriginAI Data Science Institute Columbia University(哥伦比亚大学数据科学学院) Center for Data Science New York University(纽约大学数据科学中心)

AI总结 通过因子分析发现LLM基准性能矩阵本质低秩,揭示任务冗余,提出基于潜在技能空间的评估框架,用于识别冗余任务、用小任务子集建模新模型和按技能轮廓选模型。

详情

展开后加载摘要…

URL PDF HTML 收藏