arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Stanford University(斯坦福大学)

共收录 2237
2606.00477 2026-06-02 cs.CL cs.CV

Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs

文本编辑能否泛化到视觉生成?统一多模态模型中的跨模态知识编辑基准

Xin Gao, Cheng Yang, Chufan Shi, Taylor Berg-Kirkpatrick

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学) University of Toronto(多伦多大学) University of Washington(华盛顿大学)

AI总结 提出跨模态知识编辑基准UniKE,发现文本编辑在图像生成中效果显著下降(VQA准确率仅18.5%),并提出推理增强参数编辑方法提升跨模态迁移效果。

Comments Published at ICML 2026; Code and data available at https://github.com/gxx27/UniKE

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00440 2026-06-02 cs.AI

SDR: Set-Distance Rewards for Radiology Report Generation

SDR:用于放射学报告生成的集合距离奖励

Halil Ibrahim Gulluk, Max Van Puyvelde, Wim Van Criekinge, Olivier Gevaert

机构 * Stanford University(斯坦福大学) Stanford University School of Medicine(斯坦福大学医学院) Ghent University(根特大学)

AI总结 针对胸部X光报告生成中标准奖励不兼容的问题,提出基于集合距离的连续置换不变奖励,通过GRPO后训练和测试时缩放显著提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00439 2026-06-02 cs.CV

Physical Object Understanding with a Physically Controllable World Model

基于物理可控世界模型的物理对象理解

Rahul Venkatesh, Klemen Kotar, Lilian Naing Chen, Wanhee Lee, Gia Ancone, Seungwoo Kim, Luca Thomas Wheeler, Jared Watrous, Honglin Chen, Daniel Bear, Stefan Stojanov, Daniel LK Yamins

机构 * Stanford University(斯坦福大学) OpenAI(开放人工智能公司) Noetik Inc.(Noetik公司) Google(谷歌)

AI总结 提出一类概率世界模型,通过自回归序列建模高效训练,从视频中推断对象及其物理交互,实现对象发现、3D操控和物理关系计算。

Comments CVPR 2026 Highlight. Project page at: https://neuroailab.github.io/psi-website/blog.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00374 2026-06-02 cs.RO

Constrained Whole-Body Tracking for Humanoid Robots

人形机器人的约束全身跟踪

Daniel Morton, Pranit Mohnot, Marco Pavone

机构 * Stanford University(斯坦福大学) NVIDIA Research(NVIDIA研究)

AI总结 提出 ConstrainedMimic 框架,结合操作空间控制与控制障碍函数,在强化学习跟踪策略中实现实时约束满足,用于人形机器人全身运动跟踪与遥操作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00340 2026-06-02 cs.LG

Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks

跨层学习率平衡:线性神经网络中的精确两步动力学与最优缩放

Tianyu Pang, Vignesh Kothapalli, Shenyang Deng, Haohui Wang, Dawei Zhou, Yaoqing Yang

机构 * Dartmouth College(达特茅斯学院) Stanford University(斯坦福大学)

AI总结 本文通过精确推导线性神经网络在梯度下降两步后的梯度和测试损失闭式表达式,研究了层间学习率的最优选择,揭示了初始步骤不等学习率可最小化测试损失而后续步骤等学习率最优的早期训练机制。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00320 2026-06-02 cs.LG

Adversarially Robust Control of Conditional Value-at-Risk via Rockafellar-Uryasev Conformal Inference

通过Rockafellar-Uryasev共形推断的条件风险价值对抗鲁棒控制

Catherine Chen, Jingyan Shen, Zhun Deng, Lihua Lei

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学)

AI总结 提出一种在线无分布框架,通过结合共形尾风险控制、在线学习和CVaR的变分表示,在非平稳和对抗环境下实现对条件风险价值(CVaR)的鲁棒控制,并提供渐近保证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00267 2026-06-02 cs.CV cs.AI cs.LG cs.RO

StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement

StressDream: 引导视频世界模型实现鲁棒的策略评估与改进

Junwon Seo, Sushant Veer, Ran Tian, Wenhao Ding, Apoorva Sharma, Karen Leung, Edward Schmerling, Marco Pavone, Andrea Bajcsy

机构 * Carnegie Mellon University(卡内基梅隆大学) NVIDIA Research(NVIDIA研究) University of Washington(华盛顿大学) Stanford University(斯坦福大学)

AI总结 提出StressDream方法,通过优化扩散视频世界模型的初始噪声,在推理时引导生成高影响且合理的未来场景,以支持鲁棒的策略评估与改进。

Comments Project page: https://junwon.me/StressDream/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00202 2026-06-02 cs.LG cs.AI

From Rashomon Theory to PRAXIS: Efficient Decision Tree Rashomon Sets

从Rashomon理论到PRAXIS:高效决策树Rashomon集

Zakk Heile, Hayden McTavish, Varun Babbar, Margo Seltzer, Cynthia Rudin

机构 * Stanford University(斯坦福大学)

AI总结 针对决策树Rashomon集计算开销大的问题,提出PRAXIS算法,在运行时和内存使用上实现数量级改进,并能恢复几乎完整的Rashomon集。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00156 2026-06-02 eess.IV cs.AI

A physics-informed foundation model for quantitative diffusion MRI

一种用于定量扩散MRI的物理信息基础模型

Zihan Li, Jialan Zheng, Ziyu Li, Xun Yuan, Kasidit Anmahapong, Ziang Wang, Mingxuan Liu, Hongjia Yang, Yifei Chen, Zhuhao Wang, Yuhang He, Fang Chen, Rui Li, Huaiqiang Sun, Yi Liao, Congyu Liao, Yang Yang, Haibo Qu, Xue Zhang, Hongen Liao, Qiyuan Tian

机构 * School of Biomedical Engineering, Tsinghua University(清华大学生物医学工程系) Oxford Centre for Integrative Neuroimaging, FMRIB, Nuffield Department of Clinical Neurosciences, University of Oxford(牛津大学整合神经影像中心、FMRIB、临床神经科学系) Department of Radiology, West China Second University Hospital, Sichuan University(四川大学华西第二医院放射科) School of Biomedical Engineering and the Institute of Medical Robotics, Shanghai Jiaotong University(上海交通大学生物医学工程学院和医学机器人研究院) Department of Radiology, Institution of Radiology and Medical Imaging, West China Hospital, Sichuan University(四川大学华西医院放射科、放射医学与影像研究所) Department of Radiology and Biomedical Imaging, University of California San Francisco(加州大学旧金山分校放射科和生物医学影像系) Department of Psychiatry and Behavioral Sciences, Stanford University School of Medicine(斯坦福大学医学院精神病学与行为科学系)

AI总结 提出物理信息生成微结构网络(PIGMENT),通过零样本适应实现从稀疏数据中恢复可靠的定量扩散MRI参数映射。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00047 2026-06-02 cs.CY cs.AI

Comprehensive AI governance requires addressing non-model gains

全面的人工智能治理需要解决非模型增益问题

Arthur Goemans, Dan Altman, Noemi Dreksler, Jonas Freund, Milan Gandhi, Zhengdong Wang, Sarah Cogan, Sebastien Krier, Demetra Brady, Lewis Ho, Allan Dafoe

机构 * Stanford University(斯坦福大学) UC Berkeley(加州大学伯克利分校) Open Philanthropy(开放哲学基金会)

AI总结 本文提出非模型增益的概念,包括推理增益、系统增益和资产增益,并论证这些增益会削弱以模型为中心的治理有效性,进而提出超越模型层面的治理方法。

Comments This paper has been accepted to ICML 2026 (Position paper track): https://openreview.net/forum?id=V3O1sHpKxX

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29548 2026-06-02 cs.LG

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

为什么更大的模型学习得更多:容量、干扰和稀有任务保留的影响

Jing Huang, Daniel Wurgaft, Rachit Bansal, Laura Ruis, Naomi Saphra, David Alvarez-Melis, Andrew Kyle Lampinen, Christopher Potts, Ekdeep Singh Lubana

机构 * Stanford University(斯坦福大学) Kempner Institute at Harvard University(哈佛大学凯普纳研究所) MIT(麻省理工学院) Anthropic

AI总结 通过理论分析和合成实验,研究模型规模对学习能力的影响,发现更大模型通过减少梯度干扰来学习稀有和复杂任务,并在OLMo模型上验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29341 2026-06-02 cs.CV cs.CL

WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction

WorldMemArena: 通过动作-世界交互评估多模态智能体记忆

Chengzhi Liu, Yuzhe Yang, Sophia Xiao Pu, Yepeng Liu, Lin Long, Yichen Guo, Nuo Chen, Zhaotian Weng, Elena Kochkina, Simerjot Kaur, Charese Smiley, Xiaomo Liu, James Zou, Sheng Liu, Yuheng Bu, Songyou Peng, Xin Eric Wang

机构 * University of California, Santa Barbara(加州大学圣芭芭拉分校) J.P. Morgan Chase(摩根大通) ETH Zurich(苏黎世联邦理工学院) Stanford University(斯坦福大学) Johns Hopkins University(约翰霍普金斯大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 提出WorldMemArena基准,通过动作-世界交互循环的四阶段生命周期评估多模态智能体记忆,揭示现有方法在写入、维护、检索和使用中的失败点。

Comments 25 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22285 2026-06-02 cs.LG

Demystifying Mergeability: Interpretable Properties to Predict Model Merging Success

揭秘可合并性:预测模型合并成功的可解释属性

Luca Zhou, Bo Zhao, Rose Yu, Emanuele Rodolà

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学)

AI总结 通过架构无关框架和L1正则化线性优化,发现合并成功取决于合并方法和伙伴任务,梯度对齐是最基本的兼容性信号。

Comments 9 pages of main paper, 3 figures in the main paper, 4 tables in the main paper, many more figures and tables in the appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09883 2026-06-02 cs.CV cs.AI

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space

笛卡尔捷径:在极坐标空间中重新评估视觉推理

Xia Hu, Zhenrui Yue, Brian Potetz, Howard Zhou, Leonidas Guibas, Chun-Ta Lu, Zhicheng Wang

机构 * Stanford University(斯坦福大学) Google Research(谷歌研究院)

AI总结 针对多模态大语言模型在视觉推理中利用笛卡尔坐标捷径的问题,提出Polaris-Bench基准,将任务转换至极坐标空间,揭示模型缺乏拓扑不变性视觉推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15231 2026-06-02 cs.AI

RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography

RadAgent:一种用于胸部CT逐步解读的工具型AI智能体

Mélanie Roschewitz, Kenneth Styppa, Yitian Tao, Jiwoong Sohn, Jean-Benoit Delbrouck, Benjamin Gundersen, Nicolas Deperrois, Christian Bluethgen, Julia E. Vogt, Bjoern Menze, Farhad Nooralahzadeh, Michael Krauthammer, Michael Moor

机构 * Department of Biosystems Science and Engineering, ETH Zurich(生物系统科学与工程系,苏黎世联邦理工学院) ETH AI Center, Zurich(ETH人工智能中心,苏黎世) Department of Computer Science, ETH Zurich(计算机科学系,苏黎世联邦理工学院) Faculty of Computer Science and Mathematics, Heidelberg University(计算机科学与数学学院,海德堡大学) Stanford Center for Artificial Intelligence in Medicine and Imaging, Stanford University(斯坦福大学人工智能在医学和影像中的中心) Department of Radiology, Stanford University(放射科,斯坦福大学) Department of Quantitative Biomedicine, University of Zurich(定量生物医学系,苏黎世大学) Institute of Computer Science, Zurich University of Applied Sciences(应用科学大学计算机科学研究所)

AI总结 提出RadAgent,一种通过逐步、可解释过程生成CT报告的工具型AI智能体,在临床准确性、鲁棒性和忠实度上优于3D VLM方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03789 2026-06-02 cs.LG cs.AI

Automated Conjecture Resolution with Formal Verification

自动猜想解决与形式化验证

Haocheng Ju, Guoxiong Gao, Jiedong Jiang, Bin Wu, Zeming Sun, Shurui Liu, Leheng Chen, Yutong Wang, Yuefeng Wang, Zichen Wang, Wanyi He, Peihao Wu, Liang Xiao, Ruochuan Liu, Bryan Dai, Bin Dong

机构 * School of Mathematical Sciences, Peking University(北京大学数学科学学院) Westlake Institute for Advanced Study, Westlake University(西拉雅大学先进研究所) School of Mathematics, Tianjin University(天津大学数学学院) Research Institute for Mathematical Sciences, Kyoto University(京都大学数学研究所) Department of Mathematics, Stanford University(斯坦福大学数学系) IQuest Research(IQuest研究) New Cornerstone Science Laboratory, School of Mathematical Sciences, Peking University(北京大学数学科学学院新基石科学实验室) Beijing International Center for Mathematical Research and the New Cornerstone Science Laboratory, Peking University(北京大学国际数学研究所以及新基石科学实验室) Center for Machine Learning Research, Peking University(北京大学机器学习研究中心) Center for Intelligent Computing, Great Bay Institute for Advanced Study, Great Bay University(大湾大学先进研究所智能计算中心) Zhongguancun Academy(中关村学院)

AI总结 提出一个集成非形式化推理与形式化验证的自动框架,通过两个组件Rethlas和Archon解决研究级数学问题,并成功解决交换代数中的开放问题并在Lean 4中形式化验证。

Comments Code and resources are available at: Rethlas (https://github.com/frenzymath/Rethlas), Rethlas Results (https://github.com/frenzymath/Rethlas_results), Archon (https://github.com/frenzymath/Archon), and the formalization results (https://github.com/frenzymath/Anderson-Conjecture)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25640 2026-06-02 cs.DL cs.CL

RenoBench: A Citation Parsing Benchmark

RenoBench: 引文解析基准

Parth Sarin, Juan Pablo Alperin, Adam Buttrick, Dione Mentis

机构 * Graduate School of Education, Stanford University(斯坦福大学教育研究生院) Public Knowledge Project, Simon Fraser University(公共知识项目,西蒙弗雷泽大学) DataCite, Hannover, Germany(DataCite,德国汉诺威) California Digital Library, University of California Office of the President(加州数字图书馆,加州大学校长办公室)

AI总结 针对现有引文解析评估方法不可泛化、基于合成数据或不可公开获取的问题,提出从四个出版生态系统收集的公开基准RenoBench,通过自动验证和特征采样构建多语言、多类型数据集,并评估多种解析系统,结果表明语言模型(尤其微调后)表现优异,为可重复标准化评估奠定基础。

Comments Presented as a conference paper at CiteX 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11537 2026-06-02 cs.RO

MiNI-Q: A Miniature, Wire-Free Quadruped with Unbounded, Independently Actuated Leg Joints

MiNI-Q:一种微型、无线四足机器人,具有无界、独立驱动的腿关节

Daniel Koh, Suraj Shah, Yufeng Wu, Dennis Hong

机构 * Stanford University(斯坦福大学)

AI总结 本文提出MiNI-Q^2微型无线四足机器人,通过无界独立驱动腿关节设计实现多种运动模式,速度达0.46 m/s,可折叠至2.5 cm高度,并开源所有设计文件。

Comments 7 pages, 11 figures. Submitted to the IEEE RAS Conference on Ubiquitous Robots (UR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03741 2026-06-02 cs.RO cs.AI

HALO: Learning Human-Robot Collaboration via Heterogeneous-Agent Lyapunov Policy Optimization

HALO:通过异质智能体李雅普诺夫策略优化学习人机协作

Hao Zhang, Yaru Niu, Yikai Wang, Ding Zhao, H. Eric Tseng

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学)

AI总结 针对人机协作中人类行为多样性和环境变化导致的泛化与鲁棒性问题,提出异质智能体李雅普诺夫策略优化(HALO)框架,通过李雅普诺夫收缩稳定去中心化多智能体强化学习,并利用最优二次投影修正梯度,实现理性差距的单调收缩,提升协作性能。

Comments https://HaoZhang-THU.github.io/HALO/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03291 2026-06-02 cs.CL cs.AI

One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models

一个接一个的偏差:语言奖励模型中的机械奖励塑造与持续偏差

Daniel Fein, Max Lamparth, Violet Xiang, Mykel J. Kochenderfer, Nick Haber

机构 * Stanford University(斯坦福大学) University of California, Berkeley(加州大学伯克利分校)

AI总结 本文通过系统测量五个高质量奖励模型中的偏差,发现长度、谄媚、过度自信等持续问题,并提出一种简单的后处理干预方法(机械奖励塑造)来减轻低复杂度偏差。

Comments ICML 2026 Camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07650 2026-06-02 cs.LG cs.AI

Value Flows

Value Flows

Perry Dong, Chongyi Zheng, Chelsea Finn, Dorsa Sadigh, Benjamin Eysenbach

机构 * Stanford University(斯坦福大学) Princeton University(普林斯顿大学)

AI总结 本文利用基于流的生成模型估计完整未来回报分布,通过新的流匹配目标满足分布贝尔曼方程,并利用流导数ODE估计回报不确定性以优先学习,在离线与在线设置中平均成功率提升1.3倍。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18195 2026-06-02 cs.LG cs.AI

LERD: Latent Event-Relational Dynamics for Neurodegenerative Classification

LERD: 用于神经退行性疾病分类的潜在事件-关系动力学

Yicheng Feng, Hairong Chen, Ziyu Jia, Samir Bhatt, Hengguan Huang

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学) University of Washington(华盛顿大学) University of California, San Diego(加州大学圣地亚哥分校) University of California, Los Angeles(加州大学洛杉矶分校)

AI总结 提出LERD,一种端到端贝叶斯潜在事件-关系动力系统,直接从多通道脑电图推断潜在神经事件及其关系结构,无需事件或交互标注,在阿尔茨海默病分类中优于基线方法并提供生理对齐的动力学摘要。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16745 2026-06-02 cs.LG cs.AI

PETS: A Principled Framework Towards Optimal Trajectory Allocation for Efficient Test-Time Self-Consistency

PETS:一种面向高效测试时自一致性的最优轨迹分配原则性框架

Zhangyi Liu, Huaizhi Qu, Xiaowei Yin, He Sun, Yanjun Han, Tianlong Chen, Zhun Deng

机构 * Stanford University(斯坦福大学) UNC at Chapel Hill(Chapel Hill 大学) Yale University(耶鲁大学) New York University(纽约大学)

AI总结 提出PETS框架,通过将轨迹分配建模为优化问题并引入自一致性率度量,在离线(连接众包理论)和在线流式场景下实现样本高效的测试时自一致性,显著降低采样预算。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05139 2026-06-02 cs.LG

Adaptive Exploration for Latent-State Bandits

潜在状态赌博机的自适应探索

Jikai Jin, Kenneth Hung, Sanath Kumar Krishnamurthy, Baoyi Shi, Congshan Zhang

机构 * The Institute for Computational and Mathematical Engineering(计算与数学工程研究所) Stanford University(斯坦福大学) Meta Platforms, Inc.(Meta平台公司) Ads Online Experimentation(广告在线实验部) Central Applied Science(应用科学中央研究所)

AI总结 针对奖励依赖于未观测马尔可夫状态的赌博机问题,提出基于LinUCB的自适应算法,通过滞后动作-奖励对和探针指纹两种摘要来区分状态,并采用残差、边际和过时测试动态更新指纹,在合成压力测试中相比标准、对抗和非平稳基线降低了动态遗憾。

Comments 12 pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24069 2026-06-02 cs.LG cs.AI

Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures

LLM 能否进行结构性推理?通过数据结构视角进行基准测试

Yu He, Yingxi Li, Colin White, Ellen Vitercik

机构 * Stanford University(斯坦福大学)

AI总结 本文提出 DSR-Bench 基准,通过 20 种数据结构、35 种操作和 4140 个问题实例评估 LLM 的结构性推理能力,发现顶级模型在挑战性实例上仅得 0.46/1,且在空间数据、上下文丰富场景及自身代码推理上表现不佳。

Comments Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07218 2026-06-02 cs.LG cs.AI stat.ML

Collaborative and Efficient Fine-tuning: Leveraging Task Similarity

协作高效微调:利用任务相似性

Gagik Magakyan, Amirhossein Reisizadeh, Chanwoo Park, Pablo A. Parrilo, Asuman Ozdaglar

机构 * Massachusetts Institute of Technology(麻省理工学院) Stanford University(斯坦福大学)

AI总结 提出CoLoRA方法,通过共享适配器和个性化适配器利用任务相似性进行协作微调,提升数据稀缺下的模型性能,并在理论和实验上验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04861 2026-06-02 cs.LG cond-mat.mtrl-sci cs.AI physics.chem-ph

From Evaluation to Design: Using Potential Energy Surface Smoothness Metrics to Guide Machine Learning Interatomic Potential Architectures

从评估到设计:利用势能面平滑度指标指导机器学习原子间势架构

Ryan Liu, Eric Qu, Tobias Kreiman, Samuel M. Blau, Aditi S. Krishnapriyan

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学)

AI总结 提出键平滑度表征测试(BSCT)作为高效评估机器学习原子间势(MLIP)势能面平滑度的指标,并与分子动力学稳定性强相关,同时指导模型设计以减少伪影。

Comments Accepted at the International Conference on Machine Learning (ICML) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04094 2026-06-02 cs.CV

VideoBrain: Learning Adaptive Frame Sampling for Long Video Understanding

VideoBrain: 学习自适应帧采样以理解长视频

Junbo Zou, Ziheng Huang, Shengjie Zhang, Liwen Zhang, Weining Shen

机构 * Stanford University(斯坦福大学)

AI总结 提出VideoBrain框架,通过CLIP和均匀采样双智能体策略,使视觉语言模型自适应获取关键帧,在减少30-40%帧数的同时提升长视频理解准确率3.5%-9.0%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04791 2026-06-02 cs.LG

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing

DuetServe: 通过自适应GPU多路复用协调LLM服务的预填充与解码

Lei Gao, Chaoyi Jiang, Hossein Entezari Zarch, Daniel Wong, Mark Hill, Murali Annavaram

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学) University of Toronto(多伦多大学) University of Washington(华盛顿大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 针对LLM服务中预填充与解码阶段的干扰问题,提出DuetServe框架,通过自适应SM级GPU空间多路复用实现单GPU内的阶段隔离,在保证低延迟的同时提升吞吐量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02086 2026-06-02 cs.CV

Markerless Augmented Reality Registration for Surgical Guidance: A Multi-Anatomy Clinical Accuracy Study

用于手术引导的无标记增强现实配准:一项多解剖结构临床精度研究

Yue Yang, Fabian Necker, Christoph Leuze, Michelle Chen, Andrey Finegersh, Jake Lee, Vasu Divi, Bruce Daniel, Brian Hargreaves, Jie Ying Wu, Fred M Baik

机构 * School of Medicine, Stanford University(斯坦福大学医学院) Vanderbilt Institute of Surgery and Engineering(范德比尔特手术与工程研究院)

AI总结 本文开发并临床评估了一种基于深度相机的无标记增强现实配准方法,在头戴式显示器上实现多解剖结构(足、耳、小腿)的手术引导,中位误差约3-4 mm,接近临床可接受阈值。

详情

展开后加载摘要…

URL PDF HTML 收藏