arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Stanford University(斯坦福大学)

共收录 2237
2602.14200 2026-07-01 cs.LG

TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Time-Series Reasoning

TS-Haystack:一种用于长上下文时间序列推理的多任务检索基准

Nicolas Zumarraga, Thomas Kaar, Ning Wang, William Tennien, Alpay Hasanli, Max Rosenblattl, Fan Wu, Kevin Riehl, Maxwell A. Xu, Markus Kreft, Kevin O'Sullivan, Elgar Fleisch, Paul Schmiedmayer, Robert Jakob, Patrick Langer

机构 * Agentic Systems Lab, ETH Zurich(1 非常规系统实验室,苏黎世联邦理工学院) Stanford University(2 斯坦福大学) Traffic Engineering Group, Institute for Transport Planning and Systems, ETH Zurich(3 交通工程组,交通规划与系统研究所,苏黎世联邦理工学院) University of Illinois Urbana-Champaign(4 印第安纳大学厄巴纳-香槟分校) Google(5 谷歌) Centre for Digital Health Interventions, ETH Zurich(6 数字健康干预中心,苏黎世联邦理工学院) Centre for Digital Health Interventions, University of St. Gallen(7 数字健康干预中心,圣加尔登大学)

AI总结 本文提出TS-Haystack基准,评估长上下文时间序列推理能力,发现现有TSLMs存在长上下文退化问题,代理检索框架在9/10任务中表现优异。

Comments Workshop version of this paper published at ICLR TSALM 2026. Benchmark generation code and datasets: https://github.com/AI-X-Labs/TS-Haystack

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01070 2026-07-01 cs.CL 版本更新

What If We Allocate Test-Time Compute Adaptively?

如果我们在测试时动态分配计算资源呢?

Ahsan Bilal, Ahmed Mohsin, Muhammad Umer, Ali Subhan, Hassan Rizwan, Ayesha Mohsin, Dean Hougen

机构 * Stanford University(斯坦福大学) University of Oklahoma(俄克拉荷马大学) University of California Riverside(加州大学河滨分校) National University of Sciences and Technology(国立科技大学)

AI总结 本文提出了一种基于验证器的自适应框架,通过迭代轨迹生成与选择提升推理效率,在多个数据集上优于传统测试时缩放方法。

Comments International Conference on Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18767 2026-07-01 cs.LO cs.LG 版本更新

Nazrin: An Atomic Neural Proof Automation Tactic in Lean 4

Nazrin: Lean 4中的原子神经证明自动化策略

Leni Aniva, Iori Oikawa, David Dill, Clark Barrett

机构 * Stanford University(斯坦福大学) Northeastern University(东北大学)

AI总结 提出原子策略集、转置原子化算法、ExprGraph数据结构和基于图神经网络的Nazrin证明器,通过仅调度原子策略克服现有证明代理的挑战,并在消费级硬件上训练和评估。

Comments 16 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30645 2026-06-30 cs.RO cs.AI cs.GR cs.SY eess.SY

VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes

VLK:从重建场景中的合成交互学习人形机器人全身操控

Yen-Jen Wang, Jiaman Li, Sirui Chen, Takara E. Truong, Pei Xu, Pieter Abbeel, Rocky Duan, Koushil Sreenath, Angjoo Kanazawa, Carmelo Sferrazza, Guanya Shi, Karen Liu

机构 * Amazon FAR(亚马逊FAR) UC Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 提出VLK方法,通过3D高斯泼溅重建室内场景并合成视觉-语言-运动学数据,训练人形机器人全身操控策略,实现从仿真到真实的迁移。

Comments 19 pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30457 2026-06-30 cs.RO

Behavior Prompting Policy: Demonstrations as Prompts for Manipulation

行为提示策略:将演示作为操作提示

Austin Patel, Ben Pekarek, Joel Enrique Castro Hernandez, Shuran Song

机构 * Stanford University(斯坦福大学) University of California, Berkeley(加州大学伯克利分校)

AI总结 提出行为提示策略(BPP),通过单次人类演示(行为提示)使机器人推理时执行新任务,结合手持操作接口iPhUMI收集多样化数据,在DrawAnything和LIBERO-Gen基准上验证零样本适应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29959 2026-06-30 cs.IR cs.CL

Know Before You Fetch: Calibrated Retrieval-Budget Allocation for Retrieval-Augmented Generation

知悉再取:面向检索增强生成的可校准检索预算分配

Zhe Dong, Fang Qin, Manish Shah, Yicheng Wang

机构 * University of Maine at Presque Isle(缅因大学普雷斯基尔分校) Stanford University(斯坦福大学) Independent Researcher(独立研究者)

AI总结 提出自适应RAG作为可校准检索预算分配,通过校准序列对数概率和前缀logit不确定性为正确性概率,实现分级上下文选择、选择性弃权和显式延迟/令牌权衡,显著提升校准质量并优化检索预算。

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29724 2026-06-30 cs.LG

Simplifying Flow Matching Transformations with Low-Rank Mixture Models

简化低秩混合模型的流匹配变换

Liam A. Kruse, Houjun Liu, Alexandros E. Tzikas, Mansur M. Arief, Mykel J. Kochenderfer

机构 * Stanford Intelligent Systems Laboratory in the Department of Aeronautics and Astronautics at Stanford University(斯坦福大学航空航天系智能系统实验室) Industrial and Systems Engineering Department at King Fahd University of Petroleum and Minerals(科威特石油矿物大学工业与系统工程系)

AI总结 提出用概率主成分分析混合模型作为归一化流的潜变量密度,通过KL散度对齐数据分布,简化流变换,加速训练并提升生成质量。

Comments Accepted at CoDIT 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29636 2026-06-30 quant-ph cs.LG

Lie Group Diffusion Models for Hardware-Aware Quantum Circuit Synthesis

李群扩散模型用于硬件感知的量子电路合成

Jyotirmai Singh

机构 * Stanford University(斯坦福大学)

AI总结 提出李群扩散模型,结合SU(2)流形几何和硬件约束,通过电路骨架选择器和扩散模型生成量子电路,在哈密顿模拟任务中优于基线,并支持定制化约束。

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29522 2026-06-30 cs.LG cs.CL

Do Models Read What They Write? Causal Registers in Scratchpad Reasoning

模型是否读取它们写的内容?草稿推理中的因果寄存器

Benjamin Shih, John Winnicki, Eric Darve

机构 * Institute for Computational and Mathematical Engineering(计算与数学工程研究所) Stanford University(斯坦福大学)

AI总结 通过状态跟踪任务,发现训练模型在草稿中写入中间状态后,模型会因果性地使用这些状态进行后续计算,而非仅作为文本输出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29034 2026-06-30 cs.CL cs.AI cs.IR cs.LG

The strength of clinical evidence is recoverable from language model representations but not from their stated grades

临床证据强度可从语言模型表示中恢复,但无法从其陈述的等级中恢复

Soroosh Tayebi Arasteh

机构 * Lab for AI in Medicine(医学人工智能实验室) RWTH Aachen University(亚琛工业大学) University Hospital RWTH Aachen(亚琛大学医院) Stanford University(斯坦福大学)

AI总结 本研究通过构建临床声明数据集,测试22个开源大语言模型,发现模型内部表示可编码证据强度(中位AUROC 71.8%),但模型陈述的等级接近随机(低于估计器25-27个百分点),且该信号主要来自词汇特征,不随规模提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28573 2026-06-30 cs.LG math.ST stat.TH

Replica Symmetry Breaking and Algorithmic Thresholds in Empirical Risk Minimization under Multi-Index Model

多指标模型下经验风险最小化的副本对称破缺与算法阈值

Andrea Montanari, Kangjie Zhou

机构 * Department of Mathematics and Department of Statistics, Stanford University(数学系和统计系,斯坦福大学) Department of Statistics, Columbia University(统计系,哥伦比亚大学)

AI总结 研究高维非凸经验风险最小化中多项式时间算法可达的优化区域,提出增量近似消息传递算法并刻画其训练误差及泛化误差,在渐近分析中证明算法性能最优。

Comments 80 pages; 3 pdf figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28480 2026-06-30 cs.SE cs.AI

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents

TUA-Bench:通用终端使用代理的基准测试

Shoufa Chen, Luyuan Wang, Xuan Yang, Zhiheng Liu, Yuren Cong, Yuanfeng Ji, Feiyan Zhou, Xiaohui Zhang, Fanny Yang, Belinda Zeng

机构 * Meta AI Duke University(杜克大学) Stanford University(斯坦福大学)

AI总结 提出TUA-Bench,包含120个跨五类任务的终端代理基准,涵盖日常数字活动与科学工程工作流,通过执行评分评估,发现最强模型Claude Opus 4.8仅达65.8%,揭示通用终端代理的显著差距。

Comments Website: https://www.tuabench.ai

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28393 2026-06-30 cs.CV

Transition-Aware best-of-N sampling for Longitudinal Chest X-ray Reports

面向纵向胸部X光报告的过渡感知最佳N采样

Halil Ibrahim Gulluk, Max Van Puyvelde, Wim Van Criekinge, Olivier Gevaert

机构 * Department of Electrical Engineering, Stanford University, Stanford, CA, USA(电气工程系,斯坦福大学,斯坦福,CA,美国) Department of Biomedical Data Science, Stanford University School of Medicine, Stanford, CA, USA(生物医学数据科学系,斯坦福大学医学院,斯坦福,CA,美国) Department of Mathematical Modelling, Statistics & Bioinformatics, Ghent University, Ghent, Belgium(数学建模、统计与生物信息学系,根特大学,根特,比利时)

AI总结 提出首个无需训练的过渡感知最佳N采样方案,通过集合距离编码前后变化并基于余弦距离评分,在纵向胸部X光报告生成中显著优于随机选择。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00302 2026-06-30 stat.ML cs.LG

ERICA: Quantifying Replicability of Cluster Analysis

ERICA: 量化聚类分析的可复现性

Siamak K. Sorooshyari, Manuel A. Rivas, Robert Tibshirani

机构 * Stanford University(斯坦福大学)

AI总结 提出ERICA框架,通过迭代聚类分配计算统计量,量化数据集中的聚类结构是否可复现,并应用于合成数据和乳腺癌基因表达数据,发现合成数据可复现而部分真实数据存在不可复现性。

Comments Updated writing, added link to GitHub code in the Conclusion and Discussion section

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10109 2026-06-30 cs.AI cs.HC cs.LG

LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals

基于自我报告的LLM代理能够实现通用个体模拟

Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, Michael S. Bernstein

机构 * Computer Science Department, Stanford University(斯坦福大学计算机科学系) Department of Communication Studies, Northwestern University(西北大学传播学系) Department of Communication, University of Washington(华盛顿大学传播学系) Google DeepMind(谷歌DeepMind) Department of Sociology, Stanford University(斯坦福大学社会学系) Sciences Po(巴黎政治学院)

AI总结 本文研究了基于自我报告数据的LLM代理在模拟个体行为方面的有效性,通过不同数据源构建代理并验证其在多种任务中的准确性和跨群体公平性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26815 2026-06-30 cs.CL cs.AI cs.IR

Sustainable Hybrid Document-Routed Retrieval for Financial RAG: Resolving the Robustness-Precision Trade-off

通过混合文档路由检索解决金融RAG中的鲁棒性与精度权衡

Zhiyuan Cheng, Longying Lai, Yue Liu

机构 * organization= School of Engineering, Stanford University , city= Stanford , state= CA , country= USA organization= Simon Business School, University of Rochester , city= Rochester , state= NY , country= USA organization= Accounting \& Information Systems, Rutgers University , city= Newark , state= NJ , country= USA

AI总结 本文提出混合文档路由检索方法,通过结合语义文件路由与分块检索,解决金融文档问答中鲁棒性与精度的权衡问题,提升检索性能。

Comments 26 pages, 4 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02459 2026-06-30 cs.CV

ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing

ReSpace:基于文本的自回归3D室内场景合成与编辑

Martin JJ. Bucher, Iro Armeni

机构 * Stanford University(斯坦福大学)

AI总结 ReSpace通过自回归方法实现文本驱动的3D室内场景合成与编辑,具备明确房间边界和资产无关部署能力,支持通过自然语言添加、移除和交换物体,实验在物体添加和完整场景合成质量上超越现有方法。

Comments 23 pages, 17 figures, 11 tables (incl. appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17753 2026-06-30 cs.CY cs.AI

The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems

2025人工智能代理指数:记录已部署代理式人工智能系统的技术和安全特性

Leon Staufer, Kevin Feng, Kevin Wei, Luke Bailey, Yawen Duan, Mick Yang, A. Pinar Ozisik, Stephen Casper, Noam Kolt

机构 * University of Cambridge(剑桥大学) University of Washington(华盛顿大学) Harvard Law School(哈佛法学院) Stanford University(斯坦福大学) Concordia AI(康科迪亚AI) University of Pennsylvania(宾夕法尼亚大学) Massachusetts Institute of Technology(麻省理工学院) Hebrew University of Jerusalem(耶路撒冷希伯来大学)

AI总结 本文提出2025人工智能代理指数,记录30种先进代理式AI系统的起源、设计、能力、生态系统及安全特性,揭示代理发展中的趋势和开发者透明度问题。

Comments To be publishesd at ACM FAccT 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06133 2026-06-30 math.OC cs.AI cs.LG

LLM Serving Optimization with Variable Prefill and Decode Lengths

具有可变预填和解码长度的LLM服务优化

Meixuan Wang, Yinyu Ye, Zijie Zhou

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Department of Management Science and Engineering, Stanford University(斯坦福大学管理科学与工程系) Department of Industrial Engineering and Decision Analytics, HKUST(香港科技大学工业工程与决策分析系)

AI总结 研究在固定KV缓存内存预算下,异构提示和响应长度的离线调度问题,提出Sorted-F算法以平衡批次大小与下游解码成本,实现常数因子近似保证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06972 2026-06-30 cs.CL

Categorize Early, Integrate Late: Divergent Processing Strategies in Automatic Speech Recognition

早分类,晚整合:自动语音识别中的发散处理策略

Nathan Roll, Pranav Bhalerao, Martijn Bartelds, Arjun Pawar, Yuka Tatsumi, Tolulope Ogunremi, Chen Shani, Calbert Graham, Meghan Sumner, Dan Jurafsky

机构 * Stanford University(斯坦福大学)

AI总结 研究通过架构指纹法分析Transformer和Conformer在语音识别中的处理策略,发现Conformer早分类早判别,而Transformer晚整合,揭示了不同架构对表示学习的影响。

Comments 3 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19633 2026-06-30 cs.AI

OptiMUS-0.3: Using Large Language Models to Model and Solve Optimization Problems at Scale

OptiMUS-0.3:利用大语言模型在大规模范围内建模和解决优化问题

Ali AhmadiTeshnizi, Wenzhi Gao, Herman Brunborg, Shayan Talaei, Connor Lawless, Madeleine Udell

机构 * School of Management Science and Engineering, Stanford University(斯坦福大学管理科学与工程学院) Institute for Computational and Mathematical Engineering, Stanford University(斯坦福大学计算与数学工程研究所)

AI总结 本文提出OptiMUS-0.3系统,通过大语言模型自动建模和解决线性规划问题,提升优化工具的实用性与效率,实验证明其在易实例和实际案例中表现优异。

Comments This paper documents OptiMUS-0.3, improving on OptiMUS-0.1 (arXiv:2310.06116) and OptiMUS-0.2 (arXiv:2402.10172). arXiv admin note: text overlap with arXiv:2402.10172

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07971 2026-06-30 cs.HC cs.AI

SPHERE: An Evaluation Card for Human-AI Systems

SPHERE:人类-人工智能系统评估卡

Qianou Ma, Dora Zhao, Xinran Zhao, Chenglei Si, Chenyang Yang, Ryan Louie, Ehud Reiter, Diyi Yang, Tongshuang Wu

机构 * Carnegie Mellon University(卡内基梅隆大学) Stanford University(斯坦福大学) University of Aberdeen(阿伯丁大学)

AI总结 本文提出SPHERE评估卡,从五个维度评估人类-人工智能系统,通过审查39个系统揭示当前评估实践及改进方向,并提出三项提升评估有效性和严谨性的建议。

Journal ref ACL Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18864 2026-06-30 cs.AI cs.CL cs.HC cs.LG physics.soc-ph q-bio.OT

Accelerating scientific discovery with Co-Scientist

用Co-Scientist加速科学发现

Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Artiom Myaskovsky, Grzegorz Glowaty, Felix Weissenberger, Alessio Orlandi, Dan Popovici, Anil Palepu, Keran Rong, Ryutaro Tanno, Khaled Saab, Fan Zhang, Jacob Blum, Andrew Carroll, Kavita Kulkarni, Nenad Tomasev, Dina Zverinski, Ivor Rendulic, Elahe Vedadi, Florian Hasler, Luka Rimanic, Marina Boia, Ivan Budiselic, Ben Feinstein, Mathias Bellaiche, Tom Sheffer, Jan Freyberg, Jeremy Ratcliff, Ottavia Bertolli, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Yuan Guan, Vikram Dhillon, Eeshit Dhaval Vaishnav, Byron Lee, Tiago R D Costa, José R Penadés, Gary Peltz, Yossi Matias, James Manyika, Demis Hassabis, Yunhan Xu, Pushmeet Kohli, Annalisa Pawlosky, Alan Karthikesalingam, Vivek Natarajan

机构 * Google Cloud AI Research(谷歌云人工智能研究) Google DeepMind(谷歌DeepMind) Google Research(谷歌研究) Stanford University School of Medicine(斯坦福大学医学院) Houston Methodist(休斯顿卫理公会医院) Sequome Fleming Initiative and Imperial College London(弗莱明倡议与伦敦帝国理工学院)

AI总结 Co-Scientist是一种基于Gemini的多智能体AI系统,通过异步任务框架和竞赛进化过程,提升科学假设生成质量,应用于药物再利用、新靶点发现和抗菌机制解释,验证了其加速科学发现的能力。

Comments 157 pages in total (main 42 pages, supplementary information 115 pages), 4 main figures, 1 main table, 6 extended data figures, 2 extended data tables, 9 supplementary figures, 4 supplementary tables, 37 main references, 117 supplementary references. Nature (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13944 2026-06-30 math.ST cs.LG stat.ME stat.ML stat.TH

Generalization error of min-norm interpolators in transfer learning

迁移学习中最小范数插值器的泛化误差

Yanke Song, Kenneth Gu, Sohom Bhattacharya, Pragya Sur

机构 * Department of Statistics, Harvard University(哈佛大学统计系) Department of Statistics, Stanford University(斯坦福大学统计系) Department of Statistics, University of Florida(佛罗里达大学统计系)

AI总结 本文研究了迁移学习中池化最小l2范数插值器的泛化误差,分析了协变量偏移和模型偏移下的偏差和方差,揭示了在低信噪比时增加数据可能有害,高信噪比时转移学习的收益条件,并扩展了通用设计下的结果。

Comments 149 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.08174 2026-06-30 econ.EM cs.LG stat.ME

Policy design in experiments with unknown interference

具有未知干扰的实验政策设计

Davide Viviano, Jess Rudder

机构 * Harvard University(哈佛大学) Oregon State University(俄勒冈州立大学) Stanford University(斯坦福大学) University of Chicago(芝加哥大学)

AI总结 本文研究了存在外溢效应的政策估计与推断实验设计。通过在集群对间变化随机化来估计治疗概率变化的边际效应,提出政策最优性检验,并设计多波实验估计福利最大化治疗规则。

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.08832 2026-06-30 math.ST cs.LG stat.ML stat.TH

Universality of empirical risk minimization

经验风险最小化的普遍性

Andrea Montanari, Basil Saeed

机构 * Department of Electrical Engineering, Stanford University(斯坦福大学电气工程系) Department of Statistics and Department of Mathematics, Stanford University(斯坦福大学统计学系和数学系)

AI总结 研究了经验风险最小化在高维统计和学习理论中的普遍性,证明了在特定条件下最小值仅依赖于数据分布的均值和协方差。

Comments 90 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27752 2026-06-29 cs.LG 新提交

PerturbCellRL: Verifier-Guided Reinforcement Learning for Single-Cell Perturbation Prediction

PerturbCellRL:用于单细胞扰动预测的验证器引导强化学习

Dongxia Wu, Mingyu Li, Yuhui Zhang, Anurendra Kumar, Emma Lundberg, Serena Yeung-Levy, Emily B. Fox

机构 * Stanford University(斯坦福大学) Peking University(北京大学)

AI总结 提出PerturbCellRL框架,通过强化学习后训练预训练的单细胞转录组生成器,利用细胞级验证器奖励确保生物一致性,在多个基准上提升预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27575 2026-06-29 cs.CV 新提交

Perceptual 3D Simulation With Physical World Modeling

基于物理世界建模的感知3D仿真

Wanhee Lee, Klemen Kotar, Rahul Mysore Venkatesh, Jared Watrous, Daniel L. K. Yamins

机构 * Stanford University(斯坦福大学)

AI总结 提出P3Sim系统,结合学习型物理世界模型、几何条件模块和持久场景记忆,在部分观测和不完整3D变换信号下模拟未来场景状态,实现多任务泛化。

Comments Published as a conference paper at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27484 2026-06-29 cs.CV 新提交

Fine-tuning a multimodal large language model for clinician-grade autism behavioral scoring from short home videos

微调多模态大语言模型用于从短视频中对自闭症行为进行临床级评分

Mohammadmahdi Honarmand, Parnian Azizian, Aaron Kline, Kae Nurge, Zerin Nasrin Tumpa, Saimourya Surabhi, Kaitlyn Dunlap, Yang Qian, Ali Kargarandehkordi, Sameer Neupane, Peter Washington, Dennis P. Wall

机构 * Stanford University(斯坦福大学) University of Hawaii at Manoa(夏威夷大学马诺阿分校) University of California San Francisco(加利福尼亚大学旧金山分校)

AI总结 微调Gemini 2.5 Pro模型,从400个临床评分的家庭视频中提取30个行为特征,在99名儿童上实现与临床医生评分的一致性提升40%,并零样本达到53%的自闭症诊断F1提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28309 2026-06-29 stat.ML cs.LG math.ST stat.TH 新提交

Surprises in Proper Positive-Only Learning

纯正样本学习中的意外

Shai Ben-David, Farnam Mansouri, Anay Mehrotra, Manolis Zampetakis

机构 * University of Waterloo and Vector Institute(滑铁卢大学和向量研究所) Stanford University(斯坦福大学) Yale University(耶鲁大学)

AI总结 本文研究了纯正样本二分类问题,提出了“均匀外部可分离性”这一新组合条件,并证明概念类可被恰当学习当且仅当VC维有限且满足该条件,揭示了与标准PAC学习截然不同的丰富图景。

详情

展开后加载摘要…

URL PDF HTML 收藏