arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Stanford University(斯坦福大学)

共收录 2251
2509.07506 2025-12-04 cs.DC cs.AI cs.CL cs.LG cs.SE

Astra: A Multi-Agent System for GPU Kernel Performance Optimization

Astra:一种用于GPU内核性能优化的多智能体系统

Anjiang Wei, Tianran Sun, Yogesh Seenichamy, Hang Song, Anne Ouyang, Azalia Mirhoseini, Ke Wang, Alex Aiken

机构 * Stanford University(斯坦福大学) Shanghai Jiao Tong University(上海交通大学) Nanjing University(南京大学)

AI总结 Astra是首个基于LLM的多智能体系统,用于优化GPU内核性能,通过协作生成高效内核并实现显著加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05745 2025-12-04 cs.AI cs.LG

SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning Models

SPRINT: 使推理模型能够实现交错规划与并行执行

Emil Biju, Shayan Talaei, Zhemin Huang, Mohammadreza Pourreza, Azalia Mirhoseini, Amin Saberi

机构 * Stanford University(斯坦福大学) Microsoft(微软) Google(谷歌)

AI总结 SPRINT通过动态识别并利用并行化机会,使推理模型在复杂任务中提升效率,减少序列token生成量。

Comments Published at NeurIPS 2025. Emil Biju, Shayan Talaei, and Zhemin Huang contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11056 2025-12-04 cs.CV

Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image Tokenization

流到模式:用于最新图像标记化的模式寻求扩散自编码器

Kyle Sargent, Kyle Hsu, Justin Johnson, Li Fei-Fei, Jiajun Wu

机构 * Stanford University(斯坦福大学) University of Michigan(密歇根大学)

AI总结 FlowMo是一种基于Transformer的扩散自编码器,通过模式匹配和模式寻求阶段实现图像标记化的新SOTA,无需卷积、对抗损失等。

Comments ICCV 2025, 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14332 2025-12-04 cs.LG stat.ML

Accelerating data-driven algorithm selection for combinatorial partitioning problems

加速组合划分问题的数据驱动算法选择

Vaggos Chatziafratis, Ishani Karmarkar, Yingxi Li, Ellen Vitercik

机构 * UC Santa Cruz(加州大学圣克鲁兹分校) Stanford University(斯坦福大学)

AI总结 本文提出了一种理论基础,用于数据驱动算法选择中的大小泛化,通过在较小样本上评估算法性能来预测大规模实例的表现,并验证了三种聚类算法和两种max-cut算法的泛化能力。

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03036 2025-12-03 cs.CV cs.AI

ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation

ViSAudio:端到端视频驱动的双耳空间音频生成

Mengchen Zhang, Qi Chen, Tong Wu, Zihan Liu, Dahua Lin

机构 * Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院) Stanford University(斯坦福大学) Beihang University(北航) The Chinese University of Hong Kong(香港中文大学)

AI总结 ViSAudio通过端到端双耳空间音频生成框架,从静音视频直接生成高质量空间音频,提升空间沉浸感和适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02450 2025-12-03 cs.CV cs.AI

HouseLayout3D: A Benchmark and Training-Free Baseline for 3D Layout Estimation in the Wild

HouseLayout3D: 一个用于野外3D布局估计的基准和无需训练的基线

Valentin Bieri, Marie-Julie Rakotosaona, Keisuke Tateno, Francis Engelmann, Leonidas Guibas

机构 * ETH Zurich(苏黎世联邦理工学院) Google(谷歌) Stanford University(斯坦福大学)

AI总结 HouseLayout3D提出一个无需训练的基线,通过现实世界数据推动多楼层建筑的3D布局估计研究。

Comments NeurIPS 2025 (Datasets and Benchmarks Track) Project Page: https://houselayout3d.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02268 2025-12-03 cs.CV cs.AI cs.LG eess.IV stat.ML

Spatiotemporal Pyramid Flow Matching for Climate Emulation

时空金字塔流匹配用于气候模拟

Jeremy Andrew Irvin, Jiaqi Han, Zikui Wang, Abdulaziz Alharbi, Yufei Zhao, Nomin-Erdene Bayarsaikhan, Daniele Visioni, Andrew Y. Ng, Duncan Watson-Parris

机构 * Stanford University(斯坦福大学) Cornell University(康奈尔大学) University of California, San Diego(加州大学圣地亚哥分校)

AI总结 本文提出时空金字塔流匹配方法,用于高效、准确的多时间尺度气候模拟,并通过ClimateSuite数据集验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17117 2025-12-03 cs.CL cs.AI cs.IT math.IT

From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning

从标记到思考:LLMs和人类如何在压缩与意义之间进行权衡

Chen Shani, Liron Soffer, Dan Jurafsky, Yann LeCun, Ravid Shwartz-Ziv

机构 * Stanford University(斯坦福大学) Tel Aviv University(特拉维夫大学) New York University(纽约大学) Meta - FAIR

AI总结 本文通过信息瓶颈框架比较人类与LLMs的概念结构,发现LLMs在压缩效率上优于人类,但牺牲了语义丰富性,揭示了人工与自然智能的本质差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15216 2025-12-03 cs.CR cs.AI cs.CL cs.LG

BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems

BountyBench: AI代理攻击者和防御者对现实世界网络安全系统的影响

Andy K. Zhang, Joey Ji, Celeste Menders, Riya Dulepet, Thomas Qin, Ron Y. Wang, Junrong Wu, Kyleen Liao, Jiliang Li, Jinghan Hu, Sara Hong, Nardos Demilew, Shivatmica Murgai, Jason Tran, Nishka Kacheria, Ethan Ho, Denis Liu, Lauren McLane, Olivia Bruvik, Dai-Rong Han, Seungwoo Kim, Akhil Vyas, Cuiyuanxiu Chen, Ryan Li, Weiran Xu, Jonathan Z. Ye, Prerit Choudhary, Siddharth M. Bhatia, Vikram Sivashankar, Yuxuan Bao, Dawn Song, Dan Boneh, Daniel E. Ho, Percy Liang

机构 * Stanford University(斯坦福大学) UC Berkeley(加州大学伯克利分校)

AI总结 BountyBench通过评估AI代理在漏洞检测、利用和修补中的表现,揭示AI在网络安全中的影响。

Comments 113 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07958 2025-12-03 cs.CV

Detect Anything 3D in the Wild

在野外检测任意3D对象

Hanxue Zhang, Haoran Jiang, Qingsong Yao, Yanan Sun, Renrui Zhang, Hao Zhao, Hongyang Li, Hongzi Zhu, Zetong Yang

机构 * OpenDriveLab at Shanghai AI Laboratory(上海人工智能实验室开放驾驶实验室) Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学) Stanford University(斯坦福大学) CUHK MMLab(香港大学多模态实验室) Tsinghua University(清华大学) GAC R&D Center(广汽研发中心)

AI总结 DetAny3D通过结合2D基础模型知识和3D解释器,实现了在任意相机配置下检测任意新物体的3D检测基础模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01993 2025-12-02 cs.RO cs.AI cs.CV cs.LG

RoaD: Rollouts as Demonstrations for Closed-Loop Supervised Fine-Tuning of Autonomous Driving Policies

RoaD: 通过回滚作为演示实现自动驾驶策略的闭环监督微调

Guillermo Garcia-Cobo, Maximilian Igl, Peter Karkus, Zhejun Zhang, Michael Watson, Yuxiao Chen, Boris Ivanovic, Marco Pavone

机构 * NVIDIA Research(NVIDIA研究部) Huawei VN Research Center(华为越南研究中心) Stanford University(斯坦福大学)

AI总结 RoaD通过利用自动驾驶策略自身的闭环回滚作为额外训练数据,有效缓解协变量偏移问题,提升闭环监督微调性能。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01992 2025-12-02 cs.AI cs.CL

LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess

LLM CHESS: 通过国际象棋评估大语言模型的推理与指令遵循能力

Sai Kolasani, Maxim Saplin, Nicholas Crispino, Kyle Montgomery, Jared Quincy Davis, Matei Zaharia, Chi Wang, Chenguang Wang

机构 * UC Berkeley(加州大学伯克利分校) Independent Researcher(独立研究者) UC Santa Cruz(加州大学圣克鲁兹分校) Stanford University(斯坦福大学) Google DeepMind(谷歌DeepMind)

AI总结 LLM CHESS通过国际象棋评估大语言模型的推理与指令遵循能力,揭示了推理模型与非推理模型的显著差异,并提供了评估框架和公开排行榜。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01888 2025-12-02 cs.LG cs.NA math-ph math.MP math.NA physics.comp-ph

Domain-Decomposed Graph Neural Network Surrogate Modeling for Ice Sheets

域分解图神经网络代理建模用于冰盖

Adrienne M. Propp, Mauro Perego, Eric C. Cyr, Anthony Gruber, Amanda A. Howard, Alexander Heinlein, Panos Stinis, Daniel M. Tartakovsky

机构 * Institute for Computational and Mathematical Engineering, Stanford University(计算与数学工程研究所,斯坦福大学) Department of Scientific Machine Learning, Sandia National Laboratories(科学机器学习系,桑迪亚国家实验室) Advanced Computing, Mathematics and Data Division, Pacific Northwest National Laboratory(先进计算、数学和数据 division,太平洋西北国家实验室) Delft Institute of Applied Mathematics, Delft University of Technology(应用数学研究所,代尔夫特理工大学) Department of Energy Science Engineering, Stanford University(能源科学与工程系,斯坦福大学)

AI总结 本文提出一种基于域分解和迁移学习的图神经网络代理模型,用于高效模拟冰盖动力学并提升大规模PDE系统的代理建模能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22451 2025-12-02 cs.CV cond-mat.mes-hall cs.LG

Benchmarking machine learning models for multi-class state recognition in double quantum dot data

在双量子点数据中对多类状态识别的机器学习模型基准测试

Valeria Díaz Moreno, Ryan P Khalili, Daniel Schug, Patrick J. Walsh, Justyna P. Zwolak

机构 * Department of Physics, University of Wisconsin-Madison(物理系,威斯康星大学麦迪逊分校) Department of Computer Science, University of Maryland(计算机科学系,马里兰大学) Department of Applied Physics, Stanford University(应用物理系,斯坦福大学) National Institute of Standards and Technology(国家标准与技术研究院) Joint Center for Quantum Information and Computer Science, University of Maryland(量子信息与计算机科学联合中心,马里兰大学) Department of Physics, University of Maryland(物理系,马里兰大学)

AI总结 本研究比较了四种机器学习模型在双量子点数据中的多类状态识别性能,发现CNNs在实验数据中表现最佳,具有较高的准确性和效率。

Comments 12 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17989 2025-12-02 cs.LG cs.AI

Outcome-based Reinforcement Learning to Predict the Future

基于结果的强化学习用于预测未来

Benjamin Turtel, Danny Franklin, Kris Skotheim, Luke Hewitt, Philipp Schoenegger

机构 * Lightning Rod Labs Stanford University(斯坦福大学) London School of Economics and Political Science(伦敦政治经济学院)

AI总结 本文提出基于结果的强化学习方法,通过训练紧凑模型提升预测准确性并增强概率校准,实验证明其在预测市场中的实际应用价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07867 2025-12-02 cs.CV

Continuous Perception Matters: Diagnosing Temporal Integration Failures in Multimodal Models

持续感知至关重要:多模态模型中时间整合失败的诊断

Zeyu Wang, Zhenzhen Weng, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

AI总结 本文提出CP-Bench,通过简单任务揭示多模态模型在持续感知中的时间整合缺陷,指出现有模型无法有效跨时间积累证据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01324 2025-12-02 hep-ex cs.CV

Panda: Self-distillation of Reusable Sensor-level Representations for High Energy Physics

Panda:用于高能物理的可重用传感器级表示的自蒸馏

Samuel Young, Kazuhiro Terao

机构 * Stanford University(斯坦福大学) SLAC National Accelerator Laboratory(SLAC国家加速器实验室)

AI总结 Panda通过自蒸馏方法从原始未标记数据中学习可重用的传感器级表示,显著提升了高能物理重建的效率和质量。

Comments 23 pages, 15 figures, preprint. Project page at https://youngsm.com/panda/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01048 2025-12-02 cs.CV

TRoVe: Discovering Error-Inducing Static Feature Biases in Temporal Vision-Language Models

TRoVe: 发现时间视觉-语言模型中的错误引发静态特征偏差

Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier, Akshay Chaudhari, Curtis Langlotz

机构 * Stanford University(斯坦福大学) HOPPR

AI总结 TRoVe通过识别时间视觉-语言模型中的错误引发静态特征偏差,提升模型在下游任务中的表现。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00762 2025-12-02 cs.CV

Seeing the Wind from a Falling Leaf

从飘落的树叶中看到风

Zhiyuan Gao, Jiageng Mao, Hong-Xing Yu, Haozhe Lou, Emily Yue-Ting Jia, Jernej Barbic, Jiajun Wu, Yue Wang

机构 * University of Southern California(南加州大学) Stanford University(斯坦福大学)

AI总结 本文提出一种端到端可微逆图形框架,通过视频恢复不可见的物理力,应用于风场估计和基于物理的视频生成与编辑。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00434 2025-12-02 cs.LG cs.CR stat.ML

Privacy-Preserving Generative Modeling and Clinical Validation of Longitudinal Health Records for Chronic Disease

隐私保护的生成建模与慢性病纵向健康记录的临床验证

Benjamin D. Ballyk, Ankit Gupta, Sujay Konda, Kavitha Subramanian, Chris Landon, Ahmed Ammar Naseer, Georg Maierhofer, Sumanth Swaminathan, Vasudevan Venkateshwaran

机构 * Vironix Health Inc(Vironix健康公司) University of Oxford(牛津大学) University of Cambridge(剑桥大学) Stanford University(斯坦福大学) University of Southern California(南加州大学)

AI总结 本文提出DP-TimeGAN模型,通过隐私保护生成模型处理纵向健康记录,提升慢性病诊断的隐私与效用平衡。

Comments To appear in Proceedings of Machine Learning Research Volume 297 - Proceedings of ML4H 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19476 2025-12-02 cs.RO

Gentle Object Retraction in Dense Clutter Using Multimodal Force Sensing and Imitation Learning

在密集障碍物中使用多模态力感知和模仿学习实现温和的对象回退

Dane Brouwer, Joshua Citron, Heather Nolte, Jeannette Bohg, Mark Cutkosky

机构 * Department of Mechanical Engineering, Stanford University, USA(机械工程系,斯坦福大学) Department of Computer Science, Stanford University, USA(计算机科学系,斯坦福大学)

AI总结 本研究通过多模态力感知和模仿学习,实现机器人在密集障碍物中温和地提取物体,显著提升成功率和效率。

Comments Accepted in IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04490 2025-12-02 cs.LG q-bio.BM

Multiscale guidance of protein structure prediction with heterogeneous cryo-EM data

多尺度指导基于异质冷冻电镜数据的蛋白质结构预测

Rishwanth Raghu, Axel Levy, Gordon Wetzstein, Ellen D. Zhong

机构 * Princeton University(普林斯顿大学) Stanford University(斯坦福大学)

AI总结 CryoBoltz通过结合冷冻电镜密度图与蛋白质结构预测模型,实现对动态生物分子复合物构象多样性的多尺度指导预测。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18884 2025-12-02 cs.LG cs.AI cs.CV math.OC

LORE: Lagrangian-Optimized Robust Embeddings for Visual Encoders

LORE: 基于拉格朗日优化的视觉编码器鲁棒嵌入

Borna Khodabandeh, Amirabbas Afzali, Amirhossein Afsharrad, Seyed Shahabeddin Mousavi, Sanjay Lall, Sajjad Amini, Seyed-Mohsen Moosavi-Dezfooli

机构 * Stanford University(斯坦福大学) Aktus AI University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Apple(苹果公司)

AI总结 LORE通过约束优化方法提升视觉编码器的对抗鲁棒性,同时保持清洁数据性能,有效平衡鲁棒性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00207 2025-12-02 cs.LG cs.AI

Constructing Efficient Fact-Storing MLPs for Transformers

构建高效的事实存储MLP用于Transformer

Owen Dugan, Roberto Garcia, Ronny Junkins, Jerry Liu, Dylan Zinsley, Sabri Eyuboglu, Atri Rudra, Chris Ré

机构 * Computer Science Department, Stanford University(斯坦福大学计算机科学系) Institute for Computational & Mathematical Engineering, Stanford University(斯坦福大学计算与数学工程研究所) Computer Science Department, University of Wisconsin–Madison(威斯康星大学麦迪逊分校计算机科学系) Computer Science and Engineering Department, University at Buffalo(布法罗大学计算机科学与工程系)

AI总结 本文提出了一种改进的MLP构造框架,提升了事实存储效率和实用性,并揭示了MLP事实存储能力与Transformer实用性之间的权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02139 2025-12-02 cs.AI

The Unified Cognitive Consciousness Theory for Language Models: Anchoring Semantics, Thresholds of Activation, and Emergent Reasoning

语言模型的统一认知意识理论:锚定语义、激活阈值与涌现推理

Edward Y. Chang, Zeyneb N. Kaya, Ethan Chang

机构 * Stanford University(斯坦福大学) UIUC(伊利诺伊大学香槟分校)

AI总结 该研究提出统一认知意识理论,通过语义锚定解释语言模型如何将预训练能力转化为目标导向行为,并通过实验验证了锚定强度对模型性能的影响。

Comments 21 pages, 7 figure, 4 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18151 2025-12-02 cs.GR cs.AI cs.CV

WonderPlay: Dynamic 3D Scene Generation from a Single Image and Actions

WonderPlay:从单张图像和动作生成动态3D场景

Zizhang Li, Hong-Xing Yu, Wei Liu, Yin Yang, Charles Herrmann, Gordon Wetzstein, Jiajun Wu

机构 * Stanford University(斯坦福大学) University of Utah(犹他大学)

AI总结 WonderPlay通过结合物理模拟与视频生成,实现从单张图像和动作生成多样化动态3D场景。

Comments ICCV 2025 (Highlight). The first two authors contributed equally. Project website: https://kyleleey.github.io/WonderPlay/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00983 2025-12-02 cs.GR cs.AI cs.CV

WorldScore: A Unified Evaluation Benchmark for World Generation

WorldScore: 一个用于世界生成的统一评估基准

Haoyi Duan, Hong-Xing Yu, Sirui Chen, Li Fei-Fei, Jiajun Wu

机构 * Stanford University(斯坦福大学)

AI总结 WorldScore提出一个统一评估基准,用于评估不同世界生成方法,涵盖3D、4D场景生成及视频生成,通过可控性、质量和动态性三个维度评估19种模型。

Comments ICCV 2025. Project website: https://haoyi-duan.github.io/WorldScore/ The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22990 2025-12-01 cs.CV cs.AI

MIMM-X: Disentangling Spurious Correlations for Medical Image Analysis

MIMM-X:解构医学图像分析中的虚假相关性

Louisa Fay, Hajer Reguigui, Bin Yang, Sergios Gatidis, Thomas Küstner

机构 * Medical Image and Data Analysis, University Hospital of Tübingen, Germany(医学图像与数据分析,图宾根大学医院,德国) Institute for Signal Processing and System Theory, University of Stuttgart, Germany(信号处理与系统理论研究所,斯图加特大学,德国) Stanford University, Department of Radiology, Stanford, USA(斯坦福大学放射科,斯坦福,美国)

AI总结 MIMM-X通过最小化多重虚假相关性的互信息,解构医学图像分析中的因果特征,提升模型在新环境下的泛化能力。

Journal ref FAIMI 2025. Lecture Notes in Computer Science, vol 15976. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22819 2025-12-01 math.AP cs.LG physics.flu-dyn

Resolving Sharp Gradients of Unstable Singularities to Machine Precision via Neural Networks

通过神经网络解析不稳定奇异性尖锐梯度以达到机器精度

Yongji Wang, Tristan Léger, Ching-Yao Lai, Tristan Buckmaster

机构 * New York University(纽约大学) Yale University(耶鲁大学) Stanford University(斯坦福大学)

AI总结 通过神经网络和多阶段架构,解决高梯度不稳定奇异性问题,实现高精度验证并发现新解。

Comments 27 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09082 2025-12-01 cs.CV

Taming generative video models for zero-shot optical flow extraction

驯服生成视频模型以实现零样本光流提取

Seungwoo Kim, Khai Loong Aw, Klemen Kotar, Cristobal Eyzaguirre, Wanhee Lee, Yunong Liu, Jared Watrous, Stefan Stojanov, Juan Carlos Niebles, Jiajun Wu, Daniel L. K. Yamins

机构 * Stanford University(斯坦福大学)

AI总结 本文提出KL-tracing方法,通过反事实提示实现生成视频模型的零样本光流提取,无需微调即可在现实和合成数据集上竞争现有最佳模型。

Comments Project webpage: https://neuroailab.github.io/projects/kl_tracing

详情

展开后加载摘要…

URL PDF HTML 收藏