arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Stanford University(斯坦福大学)

共收录 220
2403.06025 2026-07-21 cs.CV cs.AI 版本更新

CarbonNet: How Computer Vision Plays a Role in Climate Change? Application: Learning Geomechanics from Subsurface Geometry of CCS to Mitigate Global Warming

CarbonNet:计算机视觉如何在气候变化中发挥作用?应用:从碳捕获与封存的地下几何结构学习地质力学以缓解全球变暖

Wei Chen, Yunan Li, Yuan Tian

机构 * Stanford University(斯坦福大学)

AI总结 研究利用计算机视觉从碳捕获与封存的地下几何图像预测地表位移,以应对相关挑战。实现多种模型用于静态和瞬态力学问题,实验表明ResNetUNet在静态问题中表现出色,LSTM在瞬态问题中与Transformer性能相当,为CCS项目决策提供支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16073 2026-07-20 cs.LG stat.ML 版本更新

Stop the Sampler! Classifier-Based Adaptive Stopping for Sampling Kernels

停止采样器!基于分类器的采样核自适应停止

Kirill Korolev, Nikita Morozov, Stepan Pavlenko, Esmeralda S. Whitammer, Sergey Samsonov

机构 * Stanford University(斯坦福大学)

AI总结 提出将MCMC轨迹终止作为可学习组件,利用非循环生成流网络训练状态依赖分类器,在保证详细平衡条件下自适应停止采样,显著缩短轨迹长度并改善模式覆盖与混合。

Comments ICML 2026 SPIGM Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01514 2026-07-20 cs.CL cs.AI cs.CR 版本更新

Decoupled Alignment for Robust Plug-and-Play Adaptation

用于鲁棒即插即用适应的解耦对齐

Haozheng Luo, Jiahao Yu, Wenxin Zhang, Jialong Li, Chenghao Qiu, Yimin Wang, Eric Hanchen Jiang, Jerry Yao-Chieh Hu, Yan Chen, Binghui Wang, Xinyu Xing, Han Liu

机构 * Northwestern University(西北大学) New York University Abu Dhabi(纽约大学阿布扎克分校) Stanford University(斯坦福大学) Texas A&M University(德克萨斯农工大学) University of California, Los Angeles(加州大学洛杉矶分校) Illinois Institute of Technology(伊利诺伊理工学院)

AI总结 研究提出无需训练的大语言模型对齐方法,利用知识蒸馏提取对齐信号,经模型融合实现即插即用的对齐校正,采用增量调试识别关键知识组件,在有害问题数据集上显著提升防御成功率,且不损性能。

Comments Revised to correct the Acknowledgments section. Previous versions inadvertently included acknowledgments of NSF and NIH awards that did not support this work. Those funding acknowledgments have been removed. The technical content, results, and conclusions are unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.04599 2026-07-20 cs.LG cs.AI 版本更新

Perception-Aligned AI Outputs: End-to-End Visual Prediction for Uncertainty Communication in Clinical Decision-Making

感知对齐的人工智能输出:临床决策中用于不确定性通信的端到端视觉预测

Mohammad Eslami, Solale Tabarestani, Saber Kazeminasab, Ehsan Adeli, Glyn Elwyn, Tobias Elze, Mengyu Wang, Nazlee Zebardast, Lucia Sobrin, Nassir Navab, Daniel Shu Wei Ting, Malek Adjouadi

机构 * Harvard Ophthalmology AI Lab(哈佛眼科人工智能实验室) Schepens Eye Research Institute of Massachusetts Eye and Ear(马萨诸塞眼耳医院施佩恩眼科研究所) Harvard Medical School(哈佛医学院) Center for Advanced Technology and Education(先进教育技术中心) Florida International University(佛罗里达国际大学) Dartmouth Institute for Health Policy and Clinical Practice(达特茅斯健康政策与临床实践研究所) Dartmouth College(达特茅斯学院) Computer Aided Medical Procedures(医学辅助程序) Technical University of Munich(慕尼黑技术大学) Singapore Eye Research Institute(新加坡眼科研究所) Singapore National Eye Centre(新加坡国家眼科中心) Department of Ophthalmology, Byers Eye Institute, Stanford University(眼科部门,比尔斯眼科研究所,斯坦福大学)

AI总结 研究针对医疗保健中可解释人工智能的问题,提出以人为本的机器学习可视化学习框架VL4ML,通过直观视觉表示传达模型预测与不确定性,经多临床任务验证及评估,结果显示其能有效支持临床决策,具有广泛可及性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21317 2026-07-17 cs.HC cs.AI cs.CY 版本更新

Warning labels shift perceptions of sycophantic AI, but not its influence

警告标签改变对谄媚AI的认知,但不改变其影响

Lujain Ibrahim, Myra Cheng, Cinoo Lee, Pranav Khadpe, Desmond Ong, Dan Jurafsky, Diyi Yang

机构 * University of Oxford(牛津大学) Stanford University(斯坦福大学) Microsoft(微软) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 通过实验发现,警告标签能改变用户对谄媚AI的感知,但未能减少其对用户判断和冲突修复意愿的影响,揭示了认知与影响之间的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12963 2026-07-16 cs.CL 版本更新

The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context

稳健性的错觉:聚合准确率掩盖了与任务无关的上下文下的预测翻转

Yanzhe Zhang, Sanmi Koyejo, Diyi Yang

机构 * Georgia Tech(佐治亚理工学院) Stanford University(斯坦福大学)

AI总结 研究大语言模型在含无关上下文环境中的表现,发现聚合准确率掩盖了单个示例预测的不稳定性,如随机伪词会改变部分预测,且此不稳定性受多种因素调节,揭示了尾部风险,推动对模型进行单个示例可靠性评估。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13903 2026-07-16 cs.MA cs.AI cs.LG 版本更新

Benefits and Limitations of Communication in Multi-Agent Reasoning

多智能体推理中通信的益处与局限性

Michael Rizvi-Martel, Satwik Bhattamishra, Neil Rathi, Guillaume Rabusseau, Michael Hahn

机构 * Mila & Université de Montréal(Mila与蒙特利尔大学) University of Oxford(牛津大学) Stanford University(斯坦福大学) Saarland University(萨尔兰大学)

AI总结 研究多智能体推理中通信的益处与局限,提出理论框架分析其表达能力,应用于三个算法族,得出相关界限,通过实验验证关键量权衡,为设计可扩展多智能体推理系统提供指导。

Comments 34 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24563 2026-07-16 cs.CV cs.CL 版本更新

NeMo: Needle in a Montage for Video-Language Understanding

NeMo:视频语言理解中的蒙太奇之针

Zi-Yuan Hu, Shuo Liang, Duo Zheng, Yanyang Li, Yeyao Tao, Shijia Huang, Wei Feng, Jia Qin, Jianguang Yu, Jing Huang, Meng Fang, Yin Li, Liwei Wang

机构 * The Chinese University of Hong Kong(香港中文大学) Phoenix TV(凤凰电视台) Stanford University(斯坦福大学) University of Liverpool(利物浦大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

AI总结 为评估视频大语言模型时间理解能力,提出蒙太奇之针(NeMo)任务,开发自动数据生成管道并构建NeMoBench基准,能生成高质量数据,评估了20个模型,展现其能力与局限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07491 2026-07-15 cs.RO 版本更新

Smooth Operator: A Real-Time Sampling-Based Algorithm for Kinematic Hand Retargeting

平滑算子:一种基于实时采样的运动学手部重定向算法

Robert Jomar Malate, Erik Bauer, Norica Bacuieti, Stefanos Charalambous, Elvis Nava, Robert K. Katzschmann, Benedek Forrai

机构 * ETH Zurich(苏黎世联邦理工学院) Stanford University(斯坦福大学) mimic robotics(模仿机器人公司)

AI总结 针对基于梯度的手部重定向算法易收敛到不同局部最小值影响数据质量的问题,提出基于采样的无梯度重定向方法SBR,经模拟和真实用户研究评估,其总体任务成功率最高且显著降低操作员疲劳,为灵巧操作提供高效重定向器及基准测试方法。

Comments Minor cosmetic updates to figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24267 2026-07-15 cs.CL cs.AI 版本更新

Pigeonholing: how bad prompts hurt models, causing collapse and mistakes

鸽笼效应:不良提示导致模型崩溃和犯错

Hyunji Nam, Keertana Chidambaram, Dorottya Demszky, Natasha Jaques

机构 * Stanford University(斯坦福大学) University of Washington(华盛顿大学)

AI总结 研究不良上下文导致大语言模型性能下降和模式崩溃的“鸽笼效应”,发现重复错误答案、收敛于狭窄答案集等问题,并提出RLVR合成错误缓解方法。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04029 2026-07-15 cs.DB cs.AI cs.LG 版本更新

PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models

PluRel: 合成数据解锁关系基础模型的扩展定律

Vignesh Kothapalli, Rishabh Ranjan, Valter Hudovernik, Vijay Prakash Dwivedi, Johannes Hoffart, Carlos Guestrin, Jure Leskovec

机构 * Stanford University(斯坦福大学) SAP Labs LLC(SAP实验室)

AI总结 PluRel通过合成多表关系数据库,实现了关系基础模型在扩展定律上的突破,展示了合成数据在提升模型泛化能力方面的潜力。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01241 2026-07-15 cs.CY cs.AI 版本更新

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

首先,不伤害:迈向临床安全的大语言模型

David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj, Priyank Jain, Jessica Tran, Matthew Gwiazdon, Arjun Rustagi, Jenelle Jindal, Jacob M. Koshy, Vinay Kadiyala, Anup Agarwal, Bassman Tappuni, Brianna French, Sirus Jesudasen, Christopher V. Cosgriff, Rebanta Chakraborty, Jillian Caldwell, Susan Ziolkowski, David J. Iberri, Robert Diep, Rahul S. Dalal, Kira L. Newman, Kristin Galetta, J. Carl Pallais, Nancy Wei, Kathleen M. Buchheit, David I. Hong, Vartan Pahalyants, Ernest Y. Lee, Allen Shih, Tamara B. Kaplan, Vishnu Ravi, Sarita Khemani, Thomas A. Buckley, April S. Liang, Daniel Shirvani, Advait Patil, Nicholas Marshall, Kanav Chopra, Joel Koh, Adi Badhwar, Anastasia Perez, Austin J. Schoeffler, Mahbuba Tusty, Chase M. Walton, Liam G. McCoy, David J. H. Wu, Yingjie Weng, Sumant Ranji, Kevin Schulman, Nigam H. Shah, Jason Hom, Arnold Milstein, Arjun K. Manrai, Adam Rodman, Jonathan H. Chen, Ethan Goh

机构 * Harvard Combined Dermatology Program(哈佛联合皮肤科项目) Department of Dermatology, Mass General Brigham(麻省总医院皮肤科) Harvard Medical School(哈佛医学院) Stanford Center for Biomedical Informatics Research(斯坦福生物医学信息学研究中心) Stanford University(斯坦福大学) Division of Hospital Medicine, Department of Medicine, Stanford University School of Medicine(斯坦福大学医学院医院医学科) Department of Medicine, Cambridge Health Alliance(剑桥健康联盟医学科) Beth Israel Deaconess Hospital–Plymouth(贝塞斯达德acons医院-普利茅斯) Department of Medicine, University of California, San Francisco(加州大学旧金山分校医学科) Department of Neurology, Stanford University School of Medicine(斯坦福大学医学院神经科) Department of Medicine, Beth Israel Deaconess Medical Center(贝塞斯达德acons医学中心医学科) Division of Cardiology, Department of Medicine, Cambridge Health Alliance(剑桥健康联盟心脏病科) Department of Cardiovascular Medicine, Summa Health System(Summa健康系统心血管医学科) Division of Allergy, Pulmonary, and Critical Care Medicine, Department of Medicine, University of Wisconsin-Madison(威斯康星大学麦迪逊分校医学科过敏、呼吸科和危重医学科) Division of Pulmonary and Critical Care Medicine, Department of Medicine, Massachusetts General Hospital(麻省总医院呼吸科和危重医学科) Center for Immunology and Inflammatory Diseases, Department of Medicine, Massachusetts General Hospital(麻省总医院免疫和炎症疾病中心) Broad Institute of MIT and Harvard(MIT和哈佛Broad研究所) Division of Pulmonary, Critical Care, and Sleep Medicine, Cambridge Health Alliance(剑桥健康联盟呼吸科、危重医学科和睡眠医学科)

AI总结 提出NOHARM基准,包含1100个初级到专科咨询案例,评估28个LLM的医疗建议安全性,发现高达22.6%的案例存在严重危害风险,其中遗漏错误占80%以上。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18089 2026-07-15 cs.CV 版本更新

Together, Then Apart: Balancing Alignment and Distinctiveness for Multimodal Survival Analysis

在一起,然后分开:平衡多模态生存分析中的对齐与独特性

Wenjing Liu, Qin Ren, Wen Zhang, Yuewei Lin, Chenyu You

机构 * Stony Brook University(石溪大学) Stanford University(斯坦福大学) Johns Hopkins University(约翰霍普金斯大学) Brookhaven National Laboratory(布鲁赫林国家实验室)

AI总结 针对多模态生存分析,提出TTA框架,先基于原型对齐捕获跨模态共享结构,再通过锚定引导对比目标鼓励模态特定独特性,用不平衡最优传输处理模态不平衡和噪声对应,在多个癌症队列上评估,提升了生存预测并揭示可解释模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31651 2026-07-14 cs.AI 版本更新

FARS: A Fully Automated Research System Deployed at Scale

FARS:一个大规模部署的全自动研究系统

Qiong Tang, Tianxiang Sun, Xiangkun Hu, Xiangyang Liu, Yiran Chen, Yunfan Shao, Bobo Li, Changze Lv, Cheng Xu, Chengsong Huang, Chunyang Li, Dizhan Xue, Hao Bai, Haodong Duan, Hengquan Guo, Hongyang He, Hongyi Chen, Hui Shen, Jiahao Yuan, Jiankai Sun, Jikang Cheng, Jinfeng Xu, Jingqi Tong, Jingye Chen, Jinxiu Liu, Jixuan Leng, Junchi Yu, Kaixun Jiang, Kun Xiang, Kunpeng Yao, Lang Feng, Liangqi Yuan, Longsen Gao, Meng Li, Qi Jia, Qiushi Sun, Shengyuan Ding, Shizhan Gong, Siru Zhong, Terry Jingchen Zhang, Tianle Gu, Tianyi Liang, Weijie Liu, Weikai Yang, Weizhi Fei, Xin Wang, Xinpeng Liu, Xuanwen Ding, Yihong Tang, Yuanli Wang, Yukun Jiang, Yuming Yang, Zhengbao He, Zhikai Chen, Zhikun Xu, Zhuang Li, Zihao Huang

机构 * Analemma National University of Singapore(新加坡国立大学) Fudan University(复旦大学) University College Dublin(都柏林大学) Washington University in St. Louis(圣路易斯华盛顿大学) The Hong Kong University of Science and Technology(香港科技大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ByteDance(字节跳动) ShanghaiTech University(上海科技大学) University of Warwick(华威大学) Carnegie Mellon University(卡内基梅隆大学) University of Michigan, Ann Arbor(密歇根大学安娜堡分校) East China Normal University(华东师范大学) Stanford University(斯坦福大学) Tencent(腾讯) The University of Hong Kong(香港大学) Shanghai Innovation Institute(上海创新研究院) Nex-AGI Team(Nex-AGI团队)

AI总结 提出FARS系统,通过分阶段智能体协作自动生成研究项目,在67个AI/ML主题上产出166篇论文,经282份评审验证其可产出有价值成果,同时暴露实验范围窄、方法局限和诚信问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08878 2026-07-14 cs.CL cs.MA 版本更新

PerspectiveGap: A Benchmark for Multi-Agent Orchestration Prompting

PerspectiveGap: 多智能体编排提示的基准测试

Youran Sun, Xingyu Ren, Kejia Zhang, Xinpeng Liu, Jiaxuan Guo

机构 * University of Maryland(马里兰大学) The Chinese University of Hong Kong(香港中文大学) Stanford University(斯坦福大学)

AI总结 提出PerspectiveGap基准,评估LLM为多智能体系统编写编排提示的能力,实验显示模型平均通过率仅14.9%,表明该能力独特且未被充分评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15753 2026-07-14 cs.RO cs.CV 版本更新

Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces

层次化和整体化的开放词汇功能3D场景图用于室内空间

Xinggang Hu, Chenyangguang Zhang, Alexandros Delitzas, Xiangkui Zhang, Marc Pollefeys, Francis Engelmann, Xiangyang Ji

机构 * Tsinghua University(清华大学) ETH Zürich(苏黎世联邦理工学院) MPI for Informatics(信息研究所) Dalian University of Technology(大连理工大学) Microsoft(微软) Stanford University(斯坦福大学) University of Lugano(卢加诺大学)

AI总结 本文提出一种开放词汇管道,结合2D视觉定位和3D图优化,解决小规模密集相似实例的场景图推理问题,通过时间图优化和全局层次塑造提升室内空间的功能3D场景图生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25917 2026-07-14 cs.AI cs.CL cs.LG 版本更新

Recursive Multi-Agent Systems

递归多智能体系统

Jiaru Zou, Rui Pan, Ruizhong Qiu, Pan Lu, Shizhe Diao, Jindong Jiang, Hanghang Tong, Tong Zhang, Markus J. Buehler, Jingrui He, James Zou

机构 * UIUC(伊利诺伊大学香槟分校) Stanford University(斯坦福大学) NVIDIA(英伟达) MIT(麻省理工学院)

AI总结 本文提出递归多智能体框架RecursiveMAS,通过递归计算提升多智能体协作效率,实验证明其在多个基准测试中准确率提升8.3%,推理速度提升1.2-2.4倍,token使用减少34.6%-75.6%。

Comments Project Website: https://recursivemas.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06205 2026-07-14 cs.CL cs.AI 版本更新

Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation

Tool-MCoT:用于内容安全审核的工具增强多模态链式思维

Shutong Zhang, Dylan Zhou, Yinxiao Liu, Yang Yang, Huiwen Luo, Wenfei Zou

机构 * Stanford University(斯坦福大学) Google(谷歌) Google DeepMind(谷歌DeepMind)

AI总结 本文提出Tool-MCoT,一种基于工具增强的多模态链式思维模型,用于提升内容安全审核的效率与准确性,通过训练小语言模型以有效利用外部工具进行推理和决策。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11503 2026-07-14 q-bio.NC cs.AI cs.GT 版本更新

People use fast and flat simulation to reason about new games

人们使用快速且扁平的模拟来对新游戏进行推理

Katherine M. Collins, Cedegao E. Zhang, Lionel Wong, Mauricio Barba da Costa, Graham Todd, Adrian Weller, Samuel J. Cheyette, Thomas L. Griffiths, Joshua B. Tenenbaum

机构 * Massachusetts Institute of Technology(麻省理工学院) Princeton University(普林斯顿大学) University of Cambridge(剑桥大学) Stanford University(斯坦福大学) New York University(纽约大学) The Alan Turing Institute(艾伦·图灵研究所)

AI总结 研究人们对新游戏的推理,通过超千名参与者和121种新棋盘游戏的研究,发现人们玩新游戏或评估时具系统性和适应性理性,用“直观玩家”模型解释,为新问题应对及类人AI系统设计提供见解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09497 2026-07-14 cs.RO cs.AI 版本更新

Toward Autonomous Soft Robotic Endovascular Navigation via Imitation Learning

通过模仿学习实现自主软机器人血管内导航

Noah Barnes, Ji Woong Kim, Lingyun Di, Hannah Qu, Anuruddha Bhattacharjee, Miroslaw Janowski, Dheeraj Gandhi, Bailey Felix, Shaopeng Jiang, Olivia Young, Mark Fuge, Ryan D. Sochol, Jeremy D. Brown, Axel Krieger

机构 * Johns Hopkins University(约翰霍普金斯大学) Stanford University(斯坦福大学) McGill University(麦吉尔大学) University of Maryland(马里兰大学) Swiss Federal Institute of Technology in Lausanne (EPFL)(日内瓦联邦理工学院(EPFL)) ETH Zurich(苏黎世联邦理工学院)

AI总结 研究旨在实现自主软机器人血管内导航,开发基于变压器的模仿学习框架,通过目标条件等实现通用导航,在多种几何结构上训练并评估,策略成功率高,还进行了相关研究及扩展,提升了在未见几何结构上的成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01042 2026-07-14 cs.IR cs.AI cs.CL 版本更新

Can Argus Judge Them All? Comparing VLMs Across Domains

阿格斯能评判一切吗?跨领域比较视觉语言模型

Harsh Joshi, Gautam Siddharth Kashyap, Rafiq Ali, Ebad Shabbir, Niharika Jain, Sarthak Jain, Jiechao Gao, Usman Naseem

机构 * Bharati Vidyapeeth(巴哈提大学) Macquarie University(麦考瑞大学) DSEU-Okhla Vivekananda Institute of Professional Studies(维维kananda专业研究学院) IIIT-Delhi(德里印度理工学院) Center for SDGC(SDGC中心) Stanford University(斯坦福大学)

AI总结 研究跨领域视觉语言模型,提出ARGUS-EVAL评估框架,通过多种指标刻画模型行为,评估多个模型在下游任务表现,发现能力与可靠性导向排名有显著差异,如Qwen-2.5VL-3B-Instruct能力强,CLIP延迟和内存占用低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03578 2026-07-14 cs.LG stat.ML 版本更新

Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal Algorithms

具有交互式数据收集的分布鲁棒强化学习:基本难度与近最优算法

Miao Lu, Han Zhong, Tong Zhang, Jose Blanchet

机构 * Department of Management Science and Engineering, Stanford University(斯坦福大学管理科学与工程系) Center for Data Science, Peking University(北京大学数据科学中心) Department of Computer Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系)

AI总结 针对强化学习中模拟与现实的差距,提出分布鲁棒强化学习,通过交互式数据收集应对挑战。因支持转移难题,引入消失最小值假设,给出近最优算法,扩展到新公式和博弈,应用于库存控制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22294 2026-07-13 cs.CL cs.AI 版本更新

SLIDERS: Systematic Reviews via Automated Evidence Synthesis and Reconciliation

上下文从不够长:用于长文档集可扩展问答的结构化推理

Harshit Joshi, Priyank Shethia, Jadelynn Dao, Monica S. Lam

机构 * Computer Science Department, Stanford University(斯坦福大学计算机科学系)

AI总结 本文提出SLIDERS框架,通过结构化推理解决长文档集问答问题,利用关系数据库和SQL实现可扩展推理,优于现有基准。

Comments 53 pages (10 main), preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13356 2026-07-10 cs.CL cs.AI cs.GT 版本更新

Peer-Predictive Self-Training for Language Model Reasoning

同伴预测自训练用于语言模型推理

Shi Feng, Hanlin Zhang, Fan Nie, Sham Kakade, Yiling Chen

机构 * Harvard University(哈佛大学) Stanford University(斯坦福大学)

AI总结 本文提出PST框架,通过多模型协作利用交叉模型聚合响应作为内部训练信号,提升数学推理任务的准确率并减少生成-验证差距。

Comments 22 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14584 2026-07-10 cs.CV 版本更新

Anatomically Guided Latent Diffusion for Brain MRI Progression Modeling

用于脑 MRI 进展建模的解剖学引导潜在扩散

Cheng Wan, Bahram Jafrasteh, Ehsan Adeli, Miaomiao Zhang, Qingyu Zhao

机构 * Cornell University(康奈尔大学) Weill Cornell Medicine(韦尔医学院) Stanford University(斯坦福大学) University of Virginia(弗吉尼亚大学)

AI总结 研究旨在准确建模脑 MRI 进展,提出解剖学引导潜在扩散模型 AG-LDM,通过融合多因素简化训练流程,经实验验证其在图像质量、误差降低及特征捕捉等方面表现优异,是可靠的脑 MRI 进展建模框架。

Comments 24 pages, 7 figures, 7 tables. Code available at https://github.com/JornyWan/AG-LDM

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21414 2026-07-10 cs.CV cs.LG 版本更新

A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding

用于临床信息丰富且可解释的医学图像理解的工具瓶颈框架

Christina Liu, Alan Q. Wang, Joy Hsu, Jiajun Wu, Ehsan Adeli

机构 * California Institute of Technology(加州理工学院) Stanford University(斯坦福大学)

AI总结 针对医学图像理解中工具组合难的问题,提出工具瓶颈框架(TBF),利用工具瓶颈模型(TBM)组合VLM选择的工具,通过神经网络计算融合工具输出,在组织病理学和皮肤病学任务中表现出色,提升医学图像理解且使预测更具可解释性。

Journal ref Proceedings of the 9th International Conference on Medical Imaging with Deep Learning, Proceedings of Machine Learning Research 315 (2026) 2958-2986

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15245 2026-07-09 cs.CL cs.HC 版本更新

Practicing with Language Models Cultivates Human Empathic Communication

与语言模型练习培养人类共情沟通

Aakriti Kumar, Nalin Poungpeth, Diyi Yang, Bruce Lambert, Matthew Groh

机构 * Kellogg School of Management, Northwestern University(西北大学凯洛格管理学院) Northwestern Institute for Complex Systems, Northwestern University(西北大学复杂系统研究所) Ryan Institute on Complexity, Northwestern University(西北大学复杂性研究院) Department of Computer Science, Stanford University(斯坦福大学计算机科学系) Department of Communication Studies, Northwestern University(西北大学传播学系) Department of Computer Science, Northwestern University(西北大学计算机科学系)

AI总结 研究通过实验平台探讨共情沟通技能的提升,发现简短的LLM指导干预能有效提升参与者与规范共情沟通模式的匹配度,并发现沉默共情效应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16529 2026-07-09 cs.AI cs.HC 版本更新

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

SycoEval-EM:急诊护理模拟临床场景中大语言模型的谄媚评估

Dongshen Peng, Yi Wang, Austin Schoeffler, Sun-ha Hong, Brian Suffoletto, David Kim, Carl Preiksaitis, Christian Rose

机构 * UNC Chapel Hill(UNC夏洛特山分校) University of Waterloo(滑铁卢大学) Stanford University(斯坦福大学)

AI总结 提出SycoEval-EM多智能体模拟框架,评估19个LLM在急诊医学中对患者说服的鲁棒性,发现谄媚率呈双峰分布,静态医学基准无法预测安全表现。

Comments 19 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18164 2026-07-09 cs.RO 版本更新

GrandTour: A Legged Robotics Dataset in the Wild for Multi-Modal Perception and State Estimation

GrandTour: 一种用于多模态感知与状态估计的野外腿式机器人数据集

Turcan Tuna, Jonas Frey, Frank Fu, Katharine Patterson, Tianao Xu, Maurice Fallon, Cesar Cadena, Marco Hutter

机构 * ETH Zurich(苏黎世联邦理工学院) Stanford University(斯坦福大学) UC Berkeley(伯克利大学) Oxford University(牛津大学)

AI总结 GrandTour数据集为多模态感知与状态估计提供了大规模野外腿式机器人数据,支持SLAM和高精度状态估计的研究与开发。

Comments Turcan Tuna, and Jonas Frey contributed equally. Submitted to Sage The International Journal of Robotics Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21556 2026-07-09 cs.AI cs.GT 版本更新

Power and Limitations of Aggregation in Compound AI Systems

复合人工智能系统中聚合的力量与局限性

Nivasini Ananthakrishnan, Meena Jagadeesan

机构 * UC Berkeley(伯克利大学) Stanford University(斯坦福大学)

AI总结 研究复合人工智能系统中聚合的力量与局限性,在委托 - 代理框架内揭示可行性扩展等三种机制,证明其对引出能力扩展的作用,通过玩具任务实证说明结果,朝表征系统克服相关局限迈进。

详情

展开后加载摘要…

URL PDF HTML 收藏