Tracing Agentic Failure from the Flow of Success
从成功流中追溯智能体失败
机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; Microsoft Research(微软研究院)
AI总结 研究基于大语言模型的智能体系统失败归因问题,提出OAT方法将其转化为单类学习,用神经控制微分方程建模成功轨迹动态模式,实验表明该方法比基线快且F1分数更高,是诊断智能体系统失败的有效方向。
高校专区
从成功流中追溯智能体失败
机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; Microsoft Research(微软研究院)
AI总结 研究基于大语言模型的智能体系统失败归因问题,提出OAT方法将其转化为单类学习,用神经控制微分方程建模成功轨迹动态模式,实验表明该方法比基线快且F1分数更高,是诊断智能体系统失败的有效方向。
高维高斯均值估计在可实现污染下的研究
机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; University of California, San Diego(加州大学圣迭戈分校)
AI总结 研究在可实现ε污染模型下高斯分布均值估计问题,提出信息-计算差距,显示算法需更多样本或指数时间,同时给出近似匹配下界的方法。
首先,不伤害:迈向临床安全的大语言模型
机构 * Harvard Combined Dermatology Program(哈佛联合皮肤科项目) ; Department of Dermatology, Mass General Brigham(麻省总医院皮肤科) ; Harvard Medical School(哈佛医学院) ; Stanford Center for Biomedical Informatics Research(斯坦福生物医学信息学研究中心) ; Stanford University(斯坦福大学) ; Division of Hospital Medicine, Department of Medicine, Stanford University School of Medicine(斯坦福大学医学院医院医学科) ; Department of Medicine, Cambridge Health Alliance(剑桥健康联盟医学科) ; Beth Israel Deaconess Hospital–Plymouth(贝塞斯达德acons医院-普利茅斯) ; Department of Medicine, University of California, San Francisco(加州大学旧金山分校医学科) ; Department of Neurology, Stanford University School of Medicine(斯坦福大学医学院神经科) ; Department of Medicine, Beth Israel Deaconess Medical Center(贝塞斯达德acons医学中心医学科) ; Division of Cardiology, Department of Medicine, Cambridge Health Alliance(剑桥健康联盟心脏病科) ; Department of Cardiovascular Medicine, Summa Health System(Summa健康系统心血管医学科) ; Division of Allergy, Pulmonary, and Critical Care Medicine, Department of Medicine, University of Wisconsin-Madison(威斯康星大学麦迪逊分校医学科过敏、呼吸科和危重医学科) ; Division of Pulmonary and Critical Care Medicine, Department of Medicine, Massachusetts General Hospital(麻省总医院呼吸科和危重医学科) ; Center for Immunology and Inflammatory Diseases, Department of Medicine, Massachusetts General Hospital(麻省总医院免疫和炎症疾病中心) ; Broad Institute of MIT and Harvard(MIT和哈佛Broad研究所) ; Division of Pulmonary, Critical Care, and Sleep Medicine, Cambridge Health Alliance(剑桥健康联盟呼吸科、危重医学科和睡眠医学科)
AI总结 提出NOHARM基准,包含1100个初级到专科咨询案例,评估28个LLM的医疗建议安全性,发现高达22.6%的案例存在严重危害风险,其中遗漏错误占80%以上。
坐标缺失或损坏情况下的线性回归
机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; University of California, San Diego(加州大学圣地亚哥分校) ; University of California, Davis(加州大学戴维斯分校)
AI总结 研究高斯协变量下多变量线性回归在数据可能被擦除或损坏时的情况,通过建立新信息论下界刻画误差,得出缺失数据与损坏数据设置中最优误差匹配,知道损坏位置无普遍优势的结论。