arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Stanford University(斯坦福大学)

共收录 2246
2602.06176 2026-02-09 cs.AI cs.CL cs.LG

Large Language Model Reasoning Failures

大语言模型推理失败

Peiyang Song, Pengrui Han, Noah Goodman

机构 * California Institute of Technology(加州理工学院) Stanford University(斯坦福大学) Carleton College(卡尔顿学院)

AI总结 本文首次系统调查大语言模型推理失败问题,提出分类框架并分析其根本原因,旨在提升模型的推理能力与鲁棒性。

Comments Repository: https://github.com/Peiyang-Song/Awesome-LLM-Reasoning-Failures. Published at TMLR 2026 with Survey Certification

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06078 2026-02-09 cs.DL cs.AI cs.LG

Allocate Marginal Reviews to Borderline Papers Using LLM Comparative Ranking

使用LLM比较排名分配边缘论文的边际评审

Elliot L. Epstein, Rajat Dwaraknath, John Winnicki, Thanawat Sornwanee

机构 * Stanford University(斯坦福大学)

AI总结 本文提出利用LLM比较排名技术,在人类评审前识别边缘论文并分配额外评审资源,以提高评审效率和决策准确性。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21355 2026-02-09 cs.LG

SMMILE: An Expert-Driven Benchmark for Multimodal Medical In-Context Learning

SMMILE:一个多模态医学上下文学习的专家驱动基准

Melanie Rieff, Maya Varma, Ossian Rabow, Subathra Adithan, Julie Kim, Ken Chang, Hannah Lee, Nidhi Rohatgi, Christian Bluethgen, Mohamed S. Muneer, Jean-Benoit Delbrouck, Michael Moor

机构 * ETH Zurich(苏黎世联邦理工学院) Stanford University(斯坦福大学) Lund University(隆德大学) Jawaharlal Institute of Postgraduate Medical Education and Research(贾瓦哈拉尔·尼赫鲁医学教育与研究学院) UCSF(加州大学旧金山分校) University of Zurich(苏黎世大学) University Hospital Zurich(苏黎世大学医院) HOPPR

AI总结 SMMILE是一个专家驱动的多模态医学上下文学习基准,评估了15个MLLMs在医学任务中的多模态ICL能力,发现其表现有限且易受无关示例和最近性偏见影响。

Comments NeurIPS 2025 (Datasets & Benchmarks Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16175 2026-02-06 cs.LG cs.AI

Learning to Discover at Test Time

在测试时间学习以发现

Mert Yuksekgonul, Daniel Koceja, Xinhao Li, Federico Bianchi, Jed McCaleb, Xiaolong Wang, Jan Kautz, Yejin Choi, James Zou, Carlos Guestrin, Yu Sun

机构 * Stanford University(斯坦福大学) NVIDIA(英伟达) Astera Institute(Astera研究院) UC San Diego(加州大学圣地亚哥分校) Together AI

AI总结 通过在测试时间进行强化学习,TTT-Discover方法在多个科学领域实现了新的前沿状态,利用开放模型和公开代码实现高效解决方案。

Comments Code: https://github.com/test-time-training/discover

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13579 2026-02-06 cs.LG cs.AI

Learning to summarize user information for personalized reinforcement learning from human feedback

学习总结用户信息以实现基于人类反馈的个性化强化学习

Hyunji Nam, Yanming Wan, Mickel Liu, Peter Ahnn, Jianxun Lian, Natasha Jaques

机构 * Stanford University(斯坦福大学) University of Washington(华盛顿大学) DoorDash Microsoft Research Asia(微软亚洲研究院)

AI总结 PLUS通过学习用户信息摘要实现个性化强化学习,提升奖励模型准确性并实现零样本个性化。

Comments 10 pages for main text, 10 pages for appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15120 2026-02-06 stat.ML cs.AI cs.IT cs.LG math.IT math.ST stat.TH

Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit

神经网络在信息论极限附近学习通用多索引模型

Bohan Zhang, Zihao Wang, Hengyu Fu, Jason D. Lee

机构 * Peking University(北京大学) Stanford University(斯坦福大学) UC Berkeley(加州大学伯克利分校)

AI总结 神经网络通过梯度下降在信息论极限内高效学习多索引模型,证明两层网络在特定条件下能以最优样本和时间复杂度实现目标学习。

Comments 85 pages, 2 figures. The order of the first two authors was determined by a coin flip. Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13342 2026-02-06 cs.AI cs.CL cs.LG

Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers

验证验证者:揭示事实验证器中的陷阱与潜力

Wooseok Seo, Seungju Han, Jaehun Jung, Benjamin Newman, Seungwon Lim, Seungbeen Lee, Ximing Lu, Yejin Choi, Youngjae Yu

机构 * Yonsei University(延世大学) Stanford University(斯坦福大学) University of Washington(华盛顿大学) Seoul National University(首尔国立大学)

AI总结 本研究评估了多种大语言模型和事实验证器,揭示了数据标注问题、前沿模型性能及小型验证器的改进潜力。

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10960 2026-02-06 cs.LG cs.AI cs.DB

Relational Graph Transformer

关系图变换器

Vijay Prakash Dwivedi, Sri Jaladi, Yangyi Shen, Federico López, Charilaos I. Kanatsoulis, Rishi Puri, Matthias Fey, Jure Leskovec

机构 * Stanford University(斯坦福大学) NVIDIA

AI总结 RelGT通过新颖的多元素分词策略,解决关系图中异构性、时效性和拓扑结构的编码问题,实现对关系数据的高效建模,优于GNN基线。

Comments ICLR 2026, Code: https://github.com/snap-stanford/relgt

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03312 2026-02-05 q-bio.BM cs.LG

Unlocking hidden biomolecular conformational landscapes in diffusion models at inference time

在推理时间解锁扩散模型中隐藏的生物分子构象景观

Daniel D. Richman, Jessica Karaguesian, Carl-Mikael Suomivuori, Ron O. Dror

机构 * Stanford University(斯坦福大学) Yale School of Medicine(耶鲁医学院)

AI总结 ConforMix通过结合分类引导、过滤和自由能估计,提升扩散模型在推理时对生物分子构象变化的采样能力,实现更高效的构象发现。

Comments Project page: https://github.com/drorlab/conformix

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04063 2026-02-05 cs.CV

iSight: Towards expert-AI co-assessment for improved immunohistochemistry staining interpretation

iSight:迈向专家-人工智能协同评估以提高免疫组化染色解释

Jacob S. Leiby, Jialu Yao, Pan Lu, George Hu, Anna Davidian, Shunsuke Koga, Olivia Leung, Pravin Patel, Isabella Tondi Resta, Rebecca Rojansky, Derek Sung, Eric Yang, Paul J. Zhang, Emma Lundberg, Dokyoon Kim, Serena Yeung-Levy, James Zou, Thomas Montine, Jeffrey Nirschl, Zhi Huang

机构 * Department of Pathology and Laboratory Medicine, Perelman School of Medicine, University of Pennsylvania(病理与实验室医学系,佩尔曼医学院,宾夕法尼亚大学) Department of Electrical and Systems Engineering, University of Pennsylvania(电气与系统工程系,宾夕法尼亚大学) Department of Biomedical Data Science, Stanford University School of Medicine(生物医学数据科学系,斯坦福大学医学院) Department of Pathology, University of Maryland School of Medicine(病理学系,马里兰大学医学院) Department of Pathology, Stanford University School of Medicine(病理学系,斯坦福大学医学院)

AI总结 iSight通过多任务学习框架提升IHC染色评估的准确性,优于基础模型并改善专家一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03975 2026-02-05 cs.AI

Adaptive Test-Time Compute Allocation via Learned Heuristics over Categorical Structure

通过学习的启发式方法在分类结构上实现自适应测试时计算分配

Shuhui Qu

机构 * Stanford University(斯坦福大学)

AI总结 本文提出了一种基于学习启发式的自适应测试时计算分配方法,通过状态级选择性验证框架在MATH基准测试中实现更高的准确率,同时减少验证器调用次数。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03974 2026-02-05 cs.AI

Active Epistemic Control for Query-Efficient Verified Planning

主动认知控制用于查询高效的验证规划

Shuhui Qu

机构 * Stanford University(斯坦福大学)

AI总结 主动认知控制通过结合基于模型的信念管理和范畴可行性检查,实现查询高效的验证规划,在交互环境中减少重新规划轮次。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16381 2026-02-05 cs.CL cs.LG

PaTH Attention: Position Encoding via Accumulating Householder Transformations

PaTH Attention: 通过累积Householder变换实现位置编码

Songlin Yang, Yikang Shen, Kaiyue Wen, Shawn Tan, Mayank Mishra, Liliang Ren, Rameswar Panda, Yoon Kim

机构 * Massachusetts Institute of Technology(麻省理工学院) MIT-IBM Watson AI Lab(MIT-IBM Watson人工智能实验室) Stanford University(斯坦福大学) Microsoft(微软公司)

AI总结 PaTH Attention通过累积Householder变换实现数据依赖的位置编码,提升Transformer模型的表达能力。

Comments NeurIPS 2025 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12083 2026-02-05 cs.LO cs.LG

PICID: Proof-Driven Clause Learning in Neural Network Verification

PICID:基于证明的神经网络验证中的子句学习

Omri Isac, Idan Refaeli, Haoze Wu, Clark Barrett, Guy Katz

机构 * The Hebrew University of Jerusalem(特拉维夫大学) Amherst College(阿默斯特学院) Stanford University(斯坦福大学)

AI总结 PICID是一种生成标准Alethe格式证明的神经网络验证器,通过并行CDCL(T)架构结合先进SAT求解器和Marabou验证器,提升验证的可靠性和证明生成效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04636 2026-02-05 cs.LG stat.ML

Data-driven Error Estimation: Excess Risk Bounds without Class Complexity as Input

数据驱动的误差估计:无需类复杂度作为输入的超额风险界

Sanath Kumar Krishnamurthy, Anna Lyubarskaja, Emma Brunskill, Susan Athey

机构 * Institute of Computational and Mathematical Engineering, Stanford University(计算与数学工程研究所,斯坦福大学) Meta AI, Sunnyvale, CA(Meta AI,硅谷,CA) Department of Computer Science, Stanford University(计算机科学系,斯坦福大学) Graduate School of Business, Stanford University(商学院,斯坦福大学) Stanford Institute for Human-Centered AI(斯坦福人本人工智能研究所)

AI总结 本文提出一种数据驱动的方法,无需类复杂度输入,为估计类推导高概率误差上界,应用于置信区间、超额风险控制和上下文老虎机算法优化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04118 2026-02-05 cs.LG cs.AI cs.CL

Policy Learning with a Language Bottleneck

具有语言瓶颈的策略学习

Megha Srivastava, Cedric Colas, Dorsa Sadigh, Jacob Andreas

机构 * Stanford University(斯坦福大学) Massachusetts Institute of Technology(麻省理工学院) Inria(法国国家信息与自动化技术研究院)

AI总结 通过语言瓶颈框架,AI代理能生成可解释的策略规则,提升与人类的协作效率。

Comments Accepted to TMLR (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03822 2026-02-04 cs.CL

They Said Memes Were Harmless-We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural References

他们说迷因是无害的——我们发现了那些有害的:解码笑话、符号和文化参考

Sahil Tripathi, Gautam Siddharth Kashyap, Mehwish Nasim, Jian Yang, Jiechao Gao, Usman Naseem

机构 * Macquarie University(麦考瑞大学) The University of Western Australia(西澳大学) Stanford University(斯坦福大学) Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-罗quentcourt研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒尔研究实验室)

AI总结 CROSS-ALIGN+通过三阶段框架解决迷因滥用检测中的文化盲区、边界模糊和可解释性问题,实现性能提升和决策可解释性。

Comments Accepted at the The Web Conference 2026 (Research Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03320 2026-02-04 cs.CV cs.AI

MedSAM-Agent: Empowering Interactive Medical Image Segmentation with Multi-turn Agentic Reinforcement Learning

MedSAM-Agent: 通过多轮代理强化学习赋能交互式医学图像分割

Shengyuan Liu, Liuxin Bao, Qi Yang, Wanting Geng, Boyun Zheng, Chenxin Li, Wenting Chen, Houwen Peng, Yixuan Yuan

机构 * Chinese University of Hong Kong, Hong Kong SAR, China(香港中文大学) Hunyuan Group, Tencent(腾讯洪音集团) Institute of Automation, the Chinese Academy of Sciences, Beijing, China(中国科学院自动化研究所) Dalian University of Technology, Dalian, China(大连理工大学) Stanford University, Stanford, USA(斯坦福大学)

AI总结 MedSAM-Agent通过多轮代理强化学习提升医学图像分割的交互效率与准确性

Comments 23 Pages, 4 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03132 2026-02-04 cs.LG cs.AI cs.NE

Contrastive Concept-Tree Search for LLM-Assisted Algorithm Discovery

对比概念树搜索用于大语言模型辅助算法发现

Timothee Leleu, Sudeera Gunathilaka, Federico Ghimenti, Surya Ganguli

机构 * NTT Research(NTT研究院) Stanford University(斯坦福大学)

AI总结 本研究提出对比概念树搜索(CCTS)方法,通过提取层次化概念表示和对比学习模型,提升LLM辅助算法发现的搜索效率和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02961 2026-02-04 cs.AI

Generative Engine Optimization: A VLM and Agent Framework for Pinterest Acquisition Growth

生成引擎优化:一个面向Pinterest获取增长的VLM和代理框架

Faye Zhang, Qianyu Cheng, Jasmine Wan, Vishwakarma Singh, Jinfeng Rao, Kofi Boakye

机构 * Stanford University(斯坦福大学)

AI总结 Pinterest提出GEO框架,通过微调VLM和AI代理预测用户搜索需求,构建语义连贯的集合页面,提升生成搜索时代的视觉平台流量增长。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02864 2026-02-04 cs.RO

Accelerating Structured Chain-of-Thought in Autonomous Vehicles

加速自动驾驶中的结构化链式推理

Yi Gu, Yan Wang, Yuxiao Chen, Yurong You, Wenjie Luo, Yue Wang, Wenhao Ding, Boyi Li, Heng Yang, Boris Ivanovic, Marco Pavone

机构 * NVIDIA University of Southern California(南加州大学) Harvard University(哈佛大学) Stanford University(斯坦福大学)

AI总结 FastDriveCoT通过并行解码方法加速自动驾驶中的结构化链式推理,实现3-4倍的生成速度提升和端到端延迟的显著降低,同时保持推理效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02604 2026-02-04 econ.EM cs.AI

AI Assisted Economics Measurement From Survey: Evidence from Public Employee Pension Choice

基于调查的AI辅助经济学测量:来自公共员工养老金选择的证据

Tiancheng Wang, Krishna Sharma

机构 * Hoover Institution, Stanford University(霍夫曼研究所,斯坦福大学)

AI总结 本文提出一种基于AI的迭代框架,用于从调查中提取经济测量结构,通过软映射和有效性测试优化分类法,揭示养老金选择中的行为信号与经济机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16936 2026-02-04 cs.LG

SPAR: Self-supervised Placement-Aware Representation Learning for Distributed Sensing

SPAR:用于分布式传感的自监督位置感知表示学习

Yizhuo Chen, Tianchen Wang, You Lyu, Yanlan Hu, Jinyang Li, Tomoyoshi Kimura, Hongjue Zhao, Yigong Hu, Denizhan Kara, Tarek Abdelzaher

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Stanford University(斯坦福大学)

AI总结 SPAR通过信号与位置的二元性原则,提出了一种自监督位置感知表示学习框架,提升分布式传感在多种模态和任务中的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13197 2026-02-04 cs.LG physics.bio-ph q-bio.QM

Inferring stochastic dynamics with growth from cross-sectional data

从横截面数据推断具有增长的随机动力学

Stephen Zhang, Suryanarayana Maddu, Xiaojie Qiu, Victor Chardès

机构 * School of Mathematics and Statistics, University of Melbourne(墨尔本大学数学与统计学学院) Center for Computational Biology, Flatiron Institute(Flatiron研究所计算生物学中心) Department of Genetics, Stanford University School of Medicine(斯坦福大学医学院遗传学系)

AI总结 该研究提出了一种新的方法,通过利用福克-计划克方程的拉格朗日公式,从横截面数据推断具有增长的随机动力学,提高了准确性并简化了训练过程。

Comments 10 pages, 5 figures, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01816 2026-02-03 cs.CV

Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies

看见即相信?多模态大语言模型在视觉错觉和异常上的基准测试

Wenjin Hou, Wei Liu, Han Hu, Xiaoxiao Sun, Serena Yeung-Levy, Hehe Fan

机构 * Zhejiang University(浙江大学) Tencent Hunyuan Team(腾讯文言团队) Stanford University(斯坦福大学)

AI总结 本文提出VIA-Bench基准测试,评估多模态大语言模型在视觉错觉和异常上的性能,揭示其在复杂视觉任务中的鲁棒性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01689 2026-02-03 cs.AI cs.LG

What LLMs Think When You Don't Tell Them What to Think About?

当你不告诉LLMs该思考什么时,它们会想什么?

Yongchan Kwon, James Zou

机构 * Together AI Stanford University(斯坦福大学)

AI总结 研究揭示了LLMs在无明确主题输入下生成内容的分布特征,发现不同模型家族存在显著的主题偏好和内容深度差异,并公开了相关数据集和代码。

Comments NA

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22083 2026-02-03 cs.LG cs.AI

Latent Adversarial Regularization for Offline Preference Optimization

潜在对抗正则化用于离线偏好优化

Enyi Jiang, Yibo Jacky Zhang, Yinglun Xu, Andreas Haupt, Nancy Amato, Sanmi Koyejo

机构 * Department of Computer Science, Stanford University(斯坦福大学计算机科学系) Siebel School of Computing(塞贝尔计算学院) Data Science, University of Illinois at Urbana-Champaign(数据科学,伊利诺伊大学厄巴纳-香槟分校)

AI总结 本文提出GANPO,通过潜在空间正则化提升语言模型的离线偏好优化性能,增强鲁棒性并减少计算开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01439 2026-02-03 cs.LG cs.AI

TQL: Scaling Q-Functions with Transformers by Preventing Attention Collapse

TQL: 通过防止注意力崩溃来扩展Q函数

Perry Dong, Kuo-Han Hung, Alexander Swerdlow, Dorsa Sadigh, Chelsea Finn

机构 * Stanford University(斯坦福大学)

AI总结 TQL通过防止注意力崩溃,提升Transformer在强化学习中扩展价值函数的性能,实现43%的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01200 2026-02-03 cs.CV

Med3D-R1: Incentivizing Clinical Reasoning in 3D Medical Vision-Language Models for Abnormality Diagnosis

Med3D-R1: 促进3D医学视觉-语言模型的临床推理

Haoran Lai, Zihang Jiang, Kun Zhang, Qingsong Yao, Rongsheng Wang, Zhiyang He, Xiaodong Tao, Wei Wei, Shaohua Kevin Zhou

机构 * School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China(生物医学工程学院,生命科学与医学系,中国科学技术大学) Suzhou Institute for Advanced Research, University of Science and Technology of China(先进研究所,中国科学技术大学) Stanford University(斯坦福大学) Medical Business Department, iFlytek Co.Ltd(iFlytek公司医学业务部) The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine University of Science and Technology of China(中国科学技术大学第一附属医院,生命科学与医学系)

AI总结 Med3D-R1通过两阶段训练框架提升3D医学视觉-语言模型的临床推理能力,实现更准确的异常诊断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01128 2026-02-03 cs.LG

Tangent Space Fine-Tuning for Directional Preference Alignment in Large Language Models

切线空间微调用于大语言模型中的方向偏好对齐

Mete Erdogan

机构 * Stanford University(斯坦福大学)

AI总结 TS-DPO通过切线空间微调实现多偏好维度的可控对齐,提升模型在帮助性与冗余性之间的平衡能力。

详情

展开后加载摘要…

URL PDF HTML 收藏