arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Stanford University(斯坦福大学)

共收录 2246
2502.19312 2026-04-20 cs.LG cs.AI cs.CL cs.HC stat.ML

FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users

FSPO:少样本优化合成偏好以个性化真实用户

Anikait Singh, Sheryl Hsu, Kyle Hsu, Eric Mitchell, Stefano Ermon, Tatsunori Hashimoto, Archit Sharma, Chelsea Finn

机构 * Stanford University(斯坦福大学) OpenAI Google DeepMind(谷歌DeepMind)

AI总结 FSPO通过少样本优化合成偏好实现LLM个性化,利用用户描述理性化提升奖励建模和指令遵循,生成100万合成偏好数据,在三个领域实现87%的Alpaca Eval胜率。

Comments Website: https://fewshot-preference-optimization.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14718 2026-04-17 cs.AI cond-mat.dis-nn hep-th

The Agentification of Scientific Research: A Physicist's Perspective

科学研究的代理化:一位物理学家的观点

Xiao-Liang Qi

机构 * Leinweber Institute for Theoretical Physics, Stanford University(莱因韦伯理论物理研究所,斯坦福大学)

AI总结 本文探讨了AI革命对科学研究的影响,强调AI不仅是工具而是合作伙伴,改变科研效率与协作结构,主张持续学习与多样性思想对原创发现的重要性。

Comments 14 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14451 2026-04-17 astro-ph.CO cs.AI cs.CV physics.data-an

FAIR Universe Weak Lensing ML Uncertainty Challenge: Handling Uncertainties and Distribution Shifts for Precision Cosmology

FAIR Universe弱引力透镜ML不确定性挑战:处理不确定性和分布偏移以实现精确宇宙学

Biwei Dai, Po-Wen Chang, Wahid Bhimji, Paolo Calafiura, Ragansu Chakkappai, Yuan-Tang Chou, Sascha Diefenbacher, Jordan Dudley, Ibrahim Elsharkawy, Steven Farrell, Isabelle Guyon, Chris Harris, Elham E Khoda, Benjamin Nachman, David Rousseau, Uroš Seljak, Ihsan Ullah, Yulei Zhang

机构 * Lawrence Berkeley National Laboratory(伯克利劳伦斯国家实验室) Université Paris-Saclay, CNRS/IN2P3, IJCLab(巴黎萨克雷大学,CNRS/IN2P3,IJCLab) ChaLearn University of Washington(华盛顿大学) University of California, Berkeley(加州大学伯克利分校) University of Toronto(多伦多大学) Stanford University(斯坦福大学) SLAC National Accelerator Laboratory(SLAC国家加速器实验室) University of California, San Diego(圣地亚哥大学)

AI总结 本文提出首个弱引力透镜基准数据集,通过处理有限训练集和分布偏移,推动ML方法在弱引力透镜分析中的应用,促进系统误差处理和数据效率提升。

Comments Whitepaper for the FAIR Universe Weak Lensing ML Uncertainty Challenge Competition. More info is available at our GitHub repository https://github.com/FAIR-Universe/Cosmology_Challenge. 13 pages, 5 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14241 2026-04-17 q-bio.BM cond-mat.stat-mech cs.LG q-bio.QM

Polyformer: a generative framework for thermodynamic modeling of polymeric molecules

Polyformer:一种用于聚合物分子热力学建模的生成框架

Alessio Valentini, David Pekker, Chungwen Liang, Todd Martinez, Swagatam Mukhopadhyay

机构 * PsiDagger Department of Physics and Astronomy, University of Pittsburgh(物理与天文学系,匹兹堡大学) Department of Chemistry and The PULSE Institute, Stanford University(化学系和PULSE研究所,斯坦福大学) SLAC National Accelerator Laboratory(SLAC国家加速器实验室)

AI总结 Polyformer 是一种生成框架,用于热力学建模聚合物分子。该框架同时解决分子折叠、构象集合及温度变化对构象集合的影响问题,通过蛋白质域的实验验证显示出与分子动力学轨迹的良好一致性。

Comments 9+epsilon pages+references+appendix, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02585 2026-04-17 cs.AI cs.CL

Mitigating LLM biases toward spurious social contexts using direct preference optimization

通过直接偏好优化减轻LLM对虚假社会情境的偏见

Hyunji Nam, Dorottya Demszky

机构 * Stanford University(斯坦福大学)

AI总结 本文研究了LLM对虚假社会情境的鲁棒性,提出Debiasing-DPO方法,通过自监督训练和监督微调减少偏见并提升预测准确性。

Comments 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21025 2026-04-17 cs.CV

CaptionQA: Is Your Caption as Useful as the Image Itself?

CaptionQA: 你的描述是否和它所代表的图像一样有用?

Shijia Yang, Yunong Liu, Bohan Zhai, Ximeng Sun, Zicheng Liu, Emad Barsoum, Manling Li, Chenfeng Xu

机构 * Advanced Micro Devices, Inc.(先进微器件公司) Stanford University(斯坦福大学) Independent Researcher(独立研究者) Northwestern University(西北大学) UT Austin(德克萨斯大学奥斯汀分校)

AI总结 CaptionQA通过评估生成描述在下游任务中的实用性,揭示了图像与描述之间的效用差距,涵盖四个领域,包含25个顶级和69个子类别,提供33,027个密集标注的多选问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09179 2026-04-17 cs.LG cs.AI

A Confounding Factors-Inhibition Adversarial Learning Framework for Multi-site fMRI Mental Disorder Identification

多站点fMRI精神障碍识别的混淆因素抑制对抗学习框架

Xin Wen, Shijie Guo, Wenbo Ning, Rui Cao, Yan Niu, Bin Wan, Peng Wei, Xiaobo Liu, Jie Xiang

机构 * School of Software, Taiyuan University of Technology(太原科技大学软件学院) School of Computer Science (Data Science), Taiyuan University of Technology(太原科技大学计算机科学(数据科学)学院) Max Planck Institute for Human Cognitive and Brain Sciences(马克斯·普朗克人类认知与脑科学研究所) Department of Psychiatry & Behavioral Sciences, Stanford University(斯坦福大学精神病学与行为科学系) Montreal Neurological Institute, McGill University(蒙特利尔神经科学研究所,麦吉尔大学)

AI总结 本文提出MSalNET框架,通过节点信息组装机制提升fMRI功能连接特征提取,结合对抗学习平衡分类与站点回归任务,实验表明在ABIDE和ADHD-200数据集上性能优于其他算法。

Journal ref Proceedings of the Annual Meeting of the Cognitive Science Society 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13667 2026-04-16 cs.CV cs.ET

From Pixels to Nucleotides: End-to-End Token-Based Video Compression for DNA Storage

从像素到核苷酸:面向DNA存储的端到端令牌化视频压缩

Cihan Ruan, Lebin Zhou, Bingqing Zhao, Rongduo Han, Qiming Yuan, Chenchen Zhu, Linyi Han, Liang Yang, Wei Wang, Wei Jiang, Nam Ling

机构 * Santa Clara University(圣克拉拉大学) Stanford University(斯坦福大学) Nankai University(南开大学) Futurewei Technologies(未来科技)

AI总结 本文提出HELIX,首个端到端神经网络,联合优化视频压缩与DNA编码,通过令牌化表示与DNA四字母碱基自然对齐,实现1.91 bits per nucleotide的高效压缩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13386 2026-04-16 cs.LG

Linear Probe Accuracy Scales with Model Size and Benefits from Multi-Layer Ensembling

线性探针的准确性与模型规模成正比,并受益于多层集成

Erik Nordby, Tasha Pais, Aviel Parrack

机构 * Georgia Institute of Technology, Atlanta, Georgia, USA(佐治亚理工学院) Independent Researcher(独立研究者) Stanford University, Stanford, California, USA(斯坦福大学)

AI总结 研究显示,多层集成的线性探针在多个任务中优于单层探针,提升AUROC性能,并发现模型规模与探针准确性呈正相关。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13325 2026-04-16 cs.RO cs.SY eess.SY

Boundary Sampling to Learn Predictive Safety Filters via Pontryagin's Maximum Principle

边界采样以通过庞特里亚金极大值原理学习预测安全过滤器

James Dallas, Thomas Lew, John Talbot, Jonathan DeCastro, Somil Bansal, John Subosits

机构 * Toyota Research Institute(丰田研究院) Department of Aeronautics and Astronautics, Stanford University(斯坦福大学航空与宇航科学系)

AI总结 本文通过庞特里亚金极大值原理识别接近安全违规的轨迹,用于指导数据收集,提升学习效率,改进安全过滤器性能。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13295 2026-04-16 cs.LG math.PR stat.ML

Some Theoretical Limitations of t-SNE

t-SNE的一些理论限制

Rupert Li, Elchanan Mossel

机构 * Stanford University(斯坦福大学) Massachusetts Institute of Technology(麻省理工学院)

AI总结 本文探讨了t-SNE在降维过程中丢失数据重要特征的理论限制,通过不同场景的结果分析揭示了其信息丢失机制。

Comments 19 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11165 2026-04-16 stat.ML cs.AI cs.LG math.ST stat.TH

Cost-optimal Sequential Testing via Doubly Robust Q-learning

通过双重鲁棒Q学习实现成本最优的顺序测试

Doudou Zhou, Yiran Zhang, Dian Jin, Yingye Zheng, Lu Tian, Tianxi Cai

机构 * Department of Statistics and Data Science, National University of Singapore(国立新加坡大学统计与数据科学系) Public Health Sciences Division, Fred Hutch(Fred Hutch公共卫生科学系) Department of Biomedical Data Science, Stanford University School of Medicine(斯坦福大学医学院生物医学数据科学系) Department of Biostatistics, Harvard T.H. Chan School of Public Health(哈佛大学T.H. Chan公共卫生学院生物统计学系)

AI总结 本文研究从回顾数据中学习成本最优的顺序决策策略,提出双重鲁棒Q学习框架以估计最优策略,通过路径特定逆概率加权和辅助对比模型构建正交伪结果,实现无偏策略学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14234 2026-04-16 cs.CV

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body

ViBES:一个具有行为智能的3D虚拟身体对话代理

Juze Zhang, Changan Chen, Xin Chen, Heng Yu, Tiange Xiang, Ali Sartaz Khan, Shrinidhi K. Lakshmikanth, Ehsan Adeli

机构 * Stanford University(斯坦福大学) ByteDance(字节跳动)

AI总结 ViBES通过联合规划语言和运动,实现对话条件下的身体动作生成,提升了多轮对话中的社会交互能力。

Comments Project page: https://ai.stanford.edu/~juze/ViBES/. Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14755 2026-04-16 cs.RO cs.LG cs.SY eess.SY

Robust Verification of Controllers under State Uncertainty via Hamilton-Jacobi Reachability Analysis

通过汉密尔顿-雅可比可达性分析实现感知控制器的鲁棒验证

Albert Lin, Alessandro Pinto, Somil Bansal

机构 * Stanford University(斯坦福大学) NASA Jet Propulsion Laboratory(美国宇航局喷气推进实验室)

AI总结 本文提出RoVer-CoRe框架,利用汉密尔顿-雅可比可达性分析对感知系统进行鲁棒验证,通过整合控制器、观测函数和状态估计模块,实现对非线性、非凸、学习驱动等复杂系统的安全性和性能验证。

Comments Accepted to the 8th Annual Learning for Dynamics & Control Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05056 2026-04-16 cs.LG

Modeling Student Learning with 3.8 Million Program Traces

用380万条程序轨迹建模学生学习

Alexis Ross, Megha Srivastava, Jeremiah Blanchard, Jacob Andreas

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Stanford University(斯坦福大学) University of Florida(佛罗里达大学)

AI总结 本文通过分析380万条程序轨迹,探讨训练语言模型以理解学生编程行为及学习过程,发现基于真实轨迹的模型能更准确预测学生行为并生成更正确的代码。

Comments Accepted to 27th International Conference on AI in Education (AIED 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17674 2026-04-16 math.OC cs.LG cs.RO cs.SY eess.SY

Convex Hulls of Reachable Sets

可达集的凸包

Thomas Lew, Riccardo Bonalli, Marco Pavone

机构 * Toyota Research Institute(丰田研究院) Paris-Saclay University(巴黎-萨克雷大学) CentraleSupélec(中央超导大学) Stanford University(斯坦福大学)

AI总结 研究非线性系统在有界扰动和不确定初始条件下的可达集的凸包,提出基于常微分方程解的有限维表征,设计高效采样算法并分析误差界,应用于神经反馈回路分析和鲁棒MPC。

Comments 20 pages. IEEE Transactions on Automatic Control 2025. Simplified maximality condition (no minus sign)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13022 2026-04-15 quant-ph cs.LG math.OC stat.ML

Classical and Quantum Speedups for Non-Convex Optimization via Energy Conserving Descent

非凸优化中的经典与量子加速方法:能量守恒下降

Yihang Sun, Huaijin Wang, Patrick Hayden, Jose Blanchet

机构 * Stanford University(斯坦福大学) Stanford University, Google DeepMind(斯坦福大学,谷歌深Mind)

AI总结 本文首次分析了能量守恒下降算法,提出了一种具有能量守恒噪声的随机ECD动态和量子ECD哈密顿量的量子类比,证明了在双谷目标函数中,sECD和qECD在梯度下降基础上实现指数加速,尤其在高壁垒目标中qECD进一步加速。

Comments 33 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13961 2026-04-15 cs.CL cs.LG

Olmo 3

Olmo 3:先进全开源语言模型家族

Team Olmo, :, Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groeneveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, Jacob Morrison, Jake Poznanski, Kyle Lo, Luca Soldaini, Matt Jordan, Mayee Chen, Michael Noukhovitch, Nathan Lambert, Pete Walsh, Pradeep Dasigi, Robert Berry, Saumya Malik, Saurabh Shah, Scott Geng, Shane Arora, Shashank Gupta, Taira Anderson, Teng Xiao, Tyler Murray, Tyler Romero, Victoria Graf, Akari Asai, Akshita Bhagia, Alexander Wettig, Alisa Liu, Aman Rangapur, Chloe Anastasiades, Costa Huang, Dustin Schwenk, Harsh Trivedi, Ian Magnusson, Jaron Lochner, Jiacheng Liu, Lester James V. Miranda, Maarten Sap, Malia Morgan, Michael Schmitz, Michal Guerquin, Michael Wilson, Regan Huff, Ronan Le Bras, Rui Xin, Rulin Shao, Sam Skjonsberg, Shannon Zejiang Shen, Shuyue Stella Li, Tucker Wilde, Valentina Pyatkin, Will Merrill, Yapei Chang, Yuling Gu, Zhiyuan Zeng, Ashish Sabharwal, Luke Zettlemoyer, Pang Wei Koh, Ali Farhadi, Noah A. Smith, Hannaneh Hajishirzi

机构 * Allen Institute for AI(Allen人工智能研究所) University of Washington(华盛顿大学) Carnegie Mellon University(卡内基梅隆大学) Stanford University(斯坦福大学) Princeton University(普林斯顿大学) Massachusetts Institute of Technology(麻省理工学院) University of Maryland(马里兰大学)

AI总结 Olmo 3是一款7B和32B参数规模的先进语言模型,专注于长上下文推理、函数调用、编程、指令遵循、通用聊天和知识回忆。

Comments minor edit updates

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12364 2026-04-15 hep-ex cs.LG hep-ph physics.data-an

Cross-Domain Transfer with Particle Physics Foundation Models: From Jets to Neutrino Interactions

跨领域迁移与粒子物理基础模型:从喷注到中微子相互作用

Gregor Krzmanc, Vinicius Mikuni, Benjamin Nachman, Callum Wilkinson

机构 * Department of Physics, Stanford University(斯坦福大学物理系) Nagoya University, Kobayashi-Maskawa Institute(名古屋大学滨坂-马萨卡瓦研究所) Fundamental Physics Directorate, SLAC National Accelerator Laboratory(SLAC国家加速器实验室基础物理部门) Lawrence Berkeley National Laboratory(伯克利国家实验室)

AI总结 本文探讨了基于粒子物理基础模型的跨领域迁移能力,通过在中微子实验中实现能量回归和分类任务,验证了预训练模型在不同能量尺度和探测器技术下的泛化能力。

Comments 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12223 2026-04-15 cs.CL cs.AI cs.LG

LLM-Guided Semantic Bootstrapping for Interpretable Text Classification with Tsetlin Machines

基于Tsetlin机的语义引导语义引导可解释文本分类

Jiechao Gao, Rohan Kumar Yadav, Yuangang Li, Yuandong Pan, Jie Wang, Ying Liu, Michael Lepech

机构 * Stanford University(斯坦福大学) University of California, Irvine(加州大学伊文斯顿分校) University of the Chinese Academy of Sciences(中国科学院大学)

AI总结 本文提出一种将LLM知识转化为符号形式的框架,通过三阶段课程生成子意图,利用非否定Tsetlin机提取高置信度语义线索,提升文本分类的可解释性和准确性。

Comments Accepted to Findings of the Association for Computational Linguistics (ACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12161 2026-04-15 cs.AI

Development, Evaluation, and Deployment of a Multi-Agent System for Thoracic Tumor Board

胸腔肿瘤板的多智能体系统开发、评估与部署

Tim Ellis-Caleo, Timothy Keyes, Nerissa Ambers, Faraah Bekheet, Wen-wai Yim, Nikesh Kotecha, Nigam H. Shah, Joel Neal

机构 * Division of Oncology, Department of Medicine, Stanford University School of Medicine(斯坦福大学医学院肿瘤学部) Technology and Digital Solutions, Stanford Health Care(斯坦福健康医疗技术与数字解决方案部) Department of Biomedical Data Science, Stanford University School of Medicine(斯坦福大学医学院生物医学数据科学部) Nursing Informatics, Stanford Health Care(斯坦福健康护理信息学部) Department of Medicine, Stanford University School of Medicine(斯坦福大学医学院医学部) Microsoft AI, Redmond, WA(微软人工智能,西雅图) Stanford Cancer Institute, Palo Alto, CA(斯坦福癌症研究所)

AI总结 本文提出了一种多智能体系统,用于生成胸腔肿瘤病例摘要,以提高讨论效率和准确性,并验证了LLM在事实评分中的应用。

Comments 64 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12103 2026-04-15 eess.SY cs.LG cs.SY

Parametric Interpolation of Dynamic Mode Decomposition for Predicting Nonlinear Systems

参数插值动态模态分解用于预测非线性系统

Ananda Chakrabarti, Haitham H. Saleh, Indranil Nayak, Balasubramaniam Shanker, Fernando L. Teixeira, Debdipta Goswami

机构 * Department of Electrical and Computer Engineering, The Ohio State University(俄亥俄州立大学电气与计算机工程系) ElectroScience Laboratory, The Ohio State University(俄亥俄州立大学电科学实验室) Department of Mechanical and Aerospace Engineering, The Ohio State University(俄亥俄州立大学机械与航空航天工程系) SLAC National Accelerator Laboratory, Stanford University(斯坦福大学SLAC国家加速器实验室)

AI总结 本文提出piDMD,一种嵌入已知参数仿射结构的降阶建模框架,通过学习单一参数仿射Koopman近似降阶模型,在多个训练参数样本上预测未见参数值,提升了预测精度和鲁棒性。

Comments 22 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22440 2026-04-15 cs.HC cs.AI cs.CL

AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual Conversations

AI与我的价值观:用户对LLMs从闲聊中提取、体现和解释人类价值观的能力的看法

Bhada Yun, Renn Su, April Yi Wang

机构 * Stanford University(斯坦福大学)

AI总结 研究探讨LLMs从闲聊中提取、体现和解释人类价值观的能力,通过参与者与聊天机器人互动并完成评估访谈,发现13名参与者认为AI能理解人类价值观,警示'武器化共情'风险,并提出VAPT工具用于评估AI价值观对齐。

Comments To appear in CHI '26

Journal ref Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI '26), April 13--17, 2026, Barcelona, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10226 2026-04-15 cs.CV cs.RO

Latent Chain-of-Thought World Modeling for End-to-End Driving

潜在链式思维世界建模用于端到端驾驶

Shuhan Tan, Kashyap Chitta, Yuxiao Chen, Ran Tian, Yurong You, Yan Wang, Wenjie Luo, Yulong Cao, Philipp Krahenbuhl, Marco Pavone, Boris Ivanovic

机构 * UT Austin(得克萨斯大学奥斯汀分校) NVIDIA(英伟达) Stanford University(斯坦福大学)

AI总结 本文提出Latent-CoT-Drive模型,通过潜在语言整合链式思维推理与决策,提升驾驶性能与安全性,实现更快推理和更优轨迹质量。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19695 2026-04-15 cs.CL cs.AI cs.IR

DyBBT: Dynamic Balance via Bandit-inspired Targeting for Dialog Policy with Cognitive Dual-Systems

DyBBT:通过老虎机启发式目标实现对话策略的动态平衡

Shuyu Zhang, Yifan Wei, Jialuo Yuan, Xinru Wang, Yanmin Zhu, Bin Li, Yujie Liu

机构 * Shanghai Jiao Tong University(上海交通大学) Beihang University(北京航空航天大学) Stanford University(斯坦福大学) University of Sydney(悉尼大学) SIAT, CAS(中国科学院上海硅酸盐研究所) Beijing Institute of Graphic Communication(北京印刷学院)

AI总结 本文提出DyBBT框架,通过结构化认知状态空间解决对话探索挑战,结合快速直觉推理与慢速 deliberative 推理,提升对话策略的成功率、效率和泛化能力。

Comments Accepted in ACL2026 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11271 2026-04-15 cs.LG cs.CL cs.CV cs.MA

OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning

OctoTools: 一个具有可扩展工具的代理框架,用于复杂推理

Pan Lu, Bowen Chen, Sheng Liu, Rahul Thapa, Joseph Boen, James Zou

机构 * Stanford University(斯坦福大学)

AI总结 OctoTools通过标准化工具卡、规划器和执行器,提供一种无需训练的多代理框架,实现跨领域复杂推理,其在16种任务中达到9.3%的平均准确率提升。

Comments 88 pages, 18 figures. Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11462 2026-04-14 cs.AI

Escaping the Context Bottleneck: Active Context Curation for LLM Agents via Reinforcement Learning

突破上下文瓶颈:通过强化学习实现LLM代理的主动上下文精修

Xiaozhe Li, Tianyi Lyu, Yizhao Yang, Liang Shan, Siyi Yang, Ligao Zhang, Zhuoyi Huang, Qingwen Liu, Yang Li

机构 * Tongji University(同济大学) Stanford University(斯坦福大学) CurrentsAI Research(CurrentsAI 研究院)

AI总结 本文提出一种解耦上下文管理和任务执行的框架,通过强化学习训练轻量级策略模型ContextCurator,有效减少工作内存中的信息熵,提升LLM在长周期任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19691 2026-04-14 cs.AI stat.AP

Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight

可扩展的LLM辅助临床基准的 stewardship 与医生监督

Junze Ye, Daniel Tawfik, Alex J. Goodell, Nikhil V. Kotha, Mark K. Buyyounouski, Mohsen Bayati

机构 * Department of Operations, Information and Technology, Stanford Graduate School of Business, Stanford, CA, USA(斯坦福大学商学院运营、信息与技术系) Department of Pediatrics, Division of Critical Care Medicine, Stanford University School of Medicine, Stanford, CA, USA(斯坦福大学医学院儿科系重症监护医学部) Department of Anesthesiology, Perioperative and Pain Medicine, Stanford University School of Medicine, Stanford, CA, USA(斯坦福大学医学院麻醉学、围手术期与疼痛医学系) Department of Radiation Oncology, Stanford University School of Medicine, Stanford, CA, USA(斯坦福大学医学院放射肿瘤学系) Department of Electrical Engineering, Stanford University School of Engineering, Stanford, CA, USA(斯坦福大学工程学院电气工程系)

AI总结 研究通过医生监督流程重新评估LLM辅助生成的临床基准标签,发现其可靠性不足,且影响模型性能。

Comments Github codebase: https://github.com/junzeye/validate-medcalc-labels

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17516 2026-04-14 cs.CL cs.AI cs.CY cs.LG

SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors

SimBench:大型语言模型模拟人类行为能力的基准测试

Tiancheng Hu, Joachim Baumann, Lorenzo Lupo, Nigel Collier, Dirk Hovy, Paul Röttger

机构 * University of Cambridge(剑桥大学) Stanford University(斯坦福大学) Bocconi University(博科尼大学) University of Oxford(牛津大学)

AI总结 SimBench通过统一20个多样化的数据集,评估大型语言模型模拟人类行为的 fidelity,发现模型性能与模型大小呈对数线性关系,但与计算资源无关,且存在指令微调与模拟准确性之间的权衡。

Comments Accepted at ICLR 2026. Project Website: http://simbench.tiancheng.hu/ Data: https://huggingface.co/datasets/pitehu/SimBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09389 2026-04-14 cs.LG cs.AI

Design Principles for Sequence Models via Coefficient Dynamics

通过系数动力学设计序列模型的原则

Jerome Sieber, Antonio Orvieto, Melanie N. Zeilinger, Carmen Amo Alonso

机构 * ETH Zurich(苏黎世联邦理工学院) ELLIS Institute Tübingen(图宾根ELLIS研究所) Stanford University(斯坦福大学)

AI总结 本文通过系数动力学框架系统比较了序列模型,揭示了不同架构间的数学共同主题,并提出了设计原则以平衡表达能力、效率和稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏