arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Cornell University(康奈尔大学)

共收录 82
2607.03525 2026-07-16 cs.SE cs.CL 版本更新

GameEngineBench: Evaluating Coding Agents on Real C++ Runtime Environments

GameEngineBench:在真实C++运行时环境中评估编码智能体

Brian La, Sejoon Chang, Ben Kim, Junyoung Bae, Aamish Ahmad Beg, Sei Chang, Gonzalo Gonzalez-Pumariega, Kanav Goyal

机构 * Nitrode Nexon Intelligence Labs Dartmouth University(达特茅斯大学) Columbia University(哥伦比亚大学) Cornell University(康奈尔大学)

AI总结 研究利用GameEngineBench在虚幻引擎5项目中评估编码智能体,核心方法是构建来自九个真实游戏仓库的基准测试集,主要贡献是揭示智能体在实时交互软件的C++开发中面临挑战,凸显游戏引擎基准测试的价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03529 2026-07-15 cs.RO 版本更新

LapSurgie: Humanoid Robots Performing Surgery via Teleoperated Handheld Laparoscopy

LapSurgie: 人形机器人通过远程操控手持腹腔镜进行手术

Zekai Liang, Xiao Liang, Soofiyan Atar, Sreyan Das, Zoe Chiu, Peihan Zhang, Calvin Joyce, Florian Richter, Shanglei Liu, Michael C. Yip

机构 * Department of Electrical and Computer Engineering, University of California San Diego(加州大学圣地亚哥分校电气与计算机工程系) School of Electrical and Computer Engineering, Cornell University(康奈尔大学电气与计算机工程学院) UC San Diego Health(加州大学圣地亚哥医学中心)

AI总结 LapSurgie是一种基于人形机器人的腹腔镜远程操控框架,通过逆映射策略实现精准手到工具控制,并通过用户研究验证了其在微创手术中的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18315 2026-07-15 cs.LG cs.AI 版本更新

Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task

更高的嵌入维度为简单排序任务创建更强的世界模型

Brady Bhalla, Honglu Fan, Nancy Chen, Tony Yue YU

机构 * California Institute of Technology(加利福尼亚理工学院) Google DeepMind(谷歌DeepMind) Cornell University(康奈尔大学)

AI总结 研究在强化学习训练的变压器中嵌入维度对内部“世界模型”的影响,发现更高维度能产生更好的内部表示,经数百实验观察到两种机制,结果证明变压器构建结构化内部世界模型且模型大小可提升表示质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17670 2026-07-15 cs.CV cs.LG q-bio.QM q-bio.TO 版本更新

The TopCoW Challenge -- Topology-Aware Circle of Willis Segmentation for CT and MR Angiography

TopCoW挑战——用于CT和MR血管造影的拓扑感知Willis环分割

Kaiyuan Yang, Fabio Musio, Yihui Ma, Norman Juchler, Johannes C. Paetzold, Rami Al-Maskari, Luciano Höher, Hongwei Bran Li, Ibrahim Ethem Hamamci, Anjany Sekuboyina, Suprosanna Shit, Houjing Huang, Chinmay Prabhakar, Ezequiel de la Rosa, Bastian Wittmann, Diana Waldmannstetter, Florian Kofler, Fernando Navarro, Martin J. Menten, Ivan Ezhov, Daniel Rueckert, Iris N. Vos, Ynte M. Ruigrok, Birgitta K. Velthuis, Hugo J. Kuijf, Pengcheng Shi, Wei Liu, Ting Ma, Maximilian R. Rokuss, Yannick Kirchhoff, Fabian Isensee, Klaus Maier-Hein, Chengcheng Zhu, Huilin Zhao, Philippe Bijlenga, Julien Hämmerli, Catherine Wurster, Laura Westphal, Jeroen Bisschop, Elisa Colombo, Hakim Baazaoui, Hannah-Lea Handelsmann, Andrew Makmur, James Hallinan, Amrish Soundararajan, Benedikt Wiestler, Jan S. Kirschke, Evamaria O. Riedel, Roland Wiest, Emmanuel Montagnon, Laurent Letourneau-Guillon, Kwanseok Oh, Dahye Lee, Orhun Utku Aydin, Adam Hilbert, Jana Rieger, Dimitrios Rallios, Satoru Tanioka, Alexander Koch, Dietmar Frey, Abdul Qayyum, Moona Mazher, Steven Niederer, Nico Disch, Julius C. Holzschuh, Dominic LaBella, Francesco Galati, Daniele Falcetta, Maria A. Zuluaga, Chaolong Lin, Haoran Zhao, Zehan Zhang, Minghui Zhang, Xin You, Hanxiao Zhang, Guang-Zhong Yang, Yun Gu, Sinyoung Ra, Jongyun Hwang, Hyunjin Park, Junqiang Chen, Marek Wodzinski, Henning Müller, Nesrin Mansouri, Florent Autrusseau, Cansu Yalcin, Rachika E. Hamadache, Clara Lisazo, Joaquim Salvi, Adrià Casamitjana, Xavier Lladó, Uma Maria Lal-Trehan Estrada, Valeriia Abramova, Luca Giancardo, Arnau Oliver, Paula Casademunt, Adrian Galdran, Matteo Delucchi, Oscar Camara, Jialu Liu, Haibin Huang, Yue Cui, Zehang Lin, Yusheng Liu, Shunzhi Zhu, Tatsat R. Patel, Adnan H. Siddiqui, Vincent M. Tutino, Maysam Orouskhani, Huayu Wang, Mahmud Mossa-Basha, Yuki Sato, Sven Hirsch, Susanne Wegener, Bjoern Menze

机构 * Department of Quantitative Biomedicine, University of Zurich, Zurich, Switzerland Institute of Computational Life Sciences, Zurich University of Applied Sciences (ZHAW), Waedenswil, Switzerland Department of Neuroradiology, University Hospital of Zurich, Zurich, Switzerland Department of Neurosurgery, Zhongnan Hospital of Wuhan University, Wuhan, China Department of Radiology at Weill Cornell Medicine, Cornell University, New York, USA Institute for Tissue Engineering School of Computation, Information Technology, Technical University of Munich, Germany Athinoula A. Martinos Center for Biomedical Imaging, Harvard Medical School, Boston, USA School of Medicine Health, TUM Klinikum, Technical University of Munich, Germany Munich Center for Machine Learning, Munich, Germany Department of Computing, Imperial College London, London, UK Image Sciences Institute, UMC Utrecht, Utrecht, The Netherlands Department of Neurology Neurosurgery, University Medical Center Utrecht, Utrecht, The Netherlands Department of Radiology, University Medical Center Utrecht, Utrecht, The Netherlands Electronic \& Information Engineering School, Harbin Institute of Technology (Shenzhen), China Peng Cheng Laboratory, Shenzhen, China Division of Medical Image Computing, German Cancer Research Center (DKFZ), Heidelberg, Germany Faculty of Mathematics Computer Science, Heidelberg University, Germany Helmholtz Imaging, German Cancer Research Center, Heidelberg, Germany Data Science School for Health, Karlsruhe/Heidelberg, Germany Learning Group, Department of Radiation Oncology, Heidelberg University Hospital Department of Radiology, University of Washington, Seattle, WA, USA Department of Radiology, Ren Ji Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China Department of Clinical Neurosciences, Division of Neurosurgery, Geneva University Hospitals, Geneva, Switzerland Department of Neurology, University Hospital of Zurich, Zurich, Switzerland Department of Physiology, University of Toronto, Canada Department of Neurosurgery, University Hospital of Zurich, Zurich, Switzerland Department of Diagnostic Imaging, National University Hospital, Singapore University of Chicago, USA Department of Diagnostic Interventional Neuroradiology, University Hospital Berne University of Berne, Berne, Switzerland Centre de Recherche du Centre Hospitalier de l’Université de Montréal (CRCHUM), Montréal, Québec, Canada DEEPNOID Inc., Seoul, South Korea Department of Artificial Intelligence, Korea University, Seoul, South Korea Charité Lab for AI in Medicine (CLAIM), Charité Universitätsmedizin Berlin, Berlin, Germany Lung Institute, Faculty of Medicine, Imperial College London, London, UK Centre for Medical Image Computing, Department of Computer Science, University College London, London, UK Department of Radiation Oncology, Duke University Medical Center, Durham, NC, USA Institute of Medical Technology, Peking University Health Science Center, Beijing, China Hangzhou Genlight MedTech Co., Ltd., China Institute of Medical Robotics, Shanghai Jiao Tong University, Shanghai, China Department of Automation, Shanghai Jiao Tong University, Shanghai, China Department of Artificial Intelligence, Sungkyunkwan University, Seoul, South Korea Department of Electrical Computer Engineering, Sungkyunkwan University, Seoul, South Korea Shanghai MediWorks Precision Instruments Co., Ltd., China Institute of Informatics, HES-SO Valais-Wallis, Switzerland Department of Measurement Electronics, AGH University of Krakow, Poland Laboratoire de Thermique et Energie de Nantes (LTeN), Université Nantes, Polytech’Nantes, Nantes, France Research Institute of Computer Vision Center for Precision Health, McWilliams School of Biomedical Informatics, University of Texas Health Science Center at Houston, USA Physense, BCN-Medtech, Department of Communication Information Technologies, Universitat Pompeu Fabra, Barcelona, Spain Department of Mathematical Modeling Machine Learning, University of Zurich, Zurich, Switzerland Laboratory of Brain Atlas Brain-inspired Intelligence, Institute of Automation, Chinese Academy of Sciences, Beijing, China School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China School of Computer Information Engineering, Xiamen University of Technology, Xiamen, China Vascular Research Center, University at Buffalo, NY, USA Department of Pathology Anatomical Sciences, University at Buffalo, NY, USA Department of Neurosurgery, University at Buffalo, NY, USA LPIXEL Inc., Tokyo, Japan

AI总结 组织TopCoW基准挑战,发布含125对MRA和CTA扫描的注释数据集,参与者提交CoW分割和变体分类算法,经评估,最佳算法在多任务中表现出色,证明CoW分割算法对下游临床应用有可解释性效用。

Comments Summary paper for the TopCoW Challenge: 4 figures, 1 table, and supplementary material in appendix. Accepted for publication in NEJM AI. Datasets and best-performing algorithm Dockers are available at https://zenodo.org/records/15692630 and https://zenodo.org/records/15665435

Journal ref NEJM AI 2026;3(8)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12339 2026-07-15 cs.HC cs.AI 版本更新

SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation

SheetMind:一个由端到端大语言模型驱动的用于电子表格自动化的多智能体框架

Xi Cheng, Ruiyan Zhu, Ke Liu, Rakesh Chowdary Machineni, Lyuhao Chen, Brian Zhu, Daniel Jin, Zheng Qi, Neeraj Parihar, Zhoutian Xu, Oliver Gao

机构 * Cornell University(康奈尔大学) University of California, Berkeley(加州大学伯克利分校) University of Michigan(密歇根大学) Carnegie Mellon University(卡内基梅隆大学) Hong Kong University of Science and Technology (GZ)(香港科学与技术大学)

AI总结 介绍由大语言模型驱动的多智能体框架SheetMind用于电子表格自动化,其分层系统含三个智能体,经评估在SheetCopilot基准测试中执行成功率达100%、功能正确性54.8%超对手,消融研究验证配置优势,还集成到谷歌表格。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09218 2026-07-14 cs.RO cs.AI 版本更新

TACTIC: Tactile and Vision Conditioned Contact-Centric Control for Whole-Arm Manipulation

用于全臂操作的触觉和视觉条件下的以接触为中心的控制

Rishabh Madan, Angchen Xie, Samantha Saak, Andres Blanco, Dohyeok Lee, Sarah Grace Brown, Yunting Yan, Mark Zolotas, Jose Barreiros, Tapomayukh Bhattacharjee

机构 * Cornell University(康奈尔大学) Carnegie Mellon University(卡内基梅隆大学) Toyota Research Institute(丰田研究机构)

AI总结 研究全臂操作问题,提出TACTIC控制器,它采用以接触为中心的混合预测模型,结合多种传感,通过接触雅可比矩阵耦合动力学与运动学,集成到MPC规划器,在模拟和实际任务中表现出色,优于其他方法。

Comments RSS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26795 2026-07-14 cs.CV cs.AI cs.MM 版本更新

NaviCache: Test-Time Self-Calibration Caching for Video Generation

NaviCache: 视频生成的测试时自校准缓存

Zheqi Lv, Zhibo Zhu, Jinke Wang, Qi Tian, Shengyu Zhang, Zhengyu Chen, Chengxi Zang, Zhou Zhao, Fei Wu

机构 * Zhejiang University(浙江大学) Cornell University(康奈尔大学) Tencent Hunyuan(腾讯文生视频)

AI总结 针对视频扩散模型计算成本高的问题,提出NaviCache方法,将特征演化重构思为惯性导航系统问题,通过双状态估计架构自适应跟踪特征变化比和潜在漂移,实现有界误差的计算跳过,在多个模型上取得优异性能。

Comments Published at ICML 2026: Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13213 2026-07-14 stat.ML cs.LG math.OC physics.chem-ph 版本更新

Rare Event Analysis via Stochastic Optimal Control

基于随机最优控制的稀有事件分析

Yuanqi Du, Jiajun He, Dinghuai Zhang, Eric Vanden-Eijnden, Carles Domingo-Enrich

机构 * Microsoft Research New England(微软研究院新英格兰分部) Cornell University(康奈尔大学) University of Cambridge(剑桥大学) Courant Institute of Mathematical Sciences, NYU(纽约大学Courant数学科学研究所)

AI总结 提出将稀有事件分析中的committor函数估计转化为随机最优控制问题,通过反馈控制引导轨迹采样,并开发两种损失函数及处理亚稳态的方法,在基准系统上获得更准确的结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02905 2026-07-14 cs.AI 版本更新

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

FIRE-Bench:评估人工智能代理在科学见解重新发现方面的表现

Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, Zhiting Hu, Eric P. Xing

机构 * Johns Hopkins University(约翰霍普金斯大学) Cornell University(康奈尔大学)

AI总结 研究旨在评估大语言模型驱动的人工智能代理在科学见解重新发现上的能力。核心方法是引入FIRE-Bench基准,让代理基于高层次研究问题进行全周期探索。主要贡献是为衡量代理驱动的科学发现进展提供了严格诊断框架,揭示当前代理系统在全周期科研上的挑战。

Comments 34 pages, 3 figures, 16 tables; ICML 2026 Camera-ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03979 2026-07-13 cs.LG cs.AI 版本更新

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories

语言模型需要睡眠:学习自我修改和巩固记忆

Ali Behrouz, Farnoosh Hashemi, Adel Javanmard, Vahab Mirrokni

机构 * Google(谷歌) Cornell University(康奈尔大学)

AI总结 受人类学习过程启发,提出“睡眠”范式,通过记忆巩固(知识播种)和梦境(自我改进)两阶段,使模型持续学习、将短期记忆转化为长期知识并自我提升。

Comments A version of this work has been publicly available from September 2025 on OpenReview

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00473 2026-07-10 cs.LG cs.AI 版本更新

Deep Neural Networks as Discrete Dynamical Systems: Implications for Physics-Informed Learning

深度神经网络作为离散动力系统:对物理信息学习的启示

Abhisek Ganguly, Santosh Ansumali, Sauro Succi

机构 * Engineering Mechanics Unit, Jawaharlal Nehru Centre for Advanced Scientific Research(纳拉扬·德赛高级科学研究中心工程力学单元) Italian Institute of Technology(意大利理工学院) University of Roma Tre(罗马三大学) Physics Department, Harvard University(哈佛大学物理系) Cornell University(康奈尔大学)

AI总结 本文探讨了深度神经网络与离散动力系统之间的类比,通过比较Burgers方程和Eikonal方程的数值/精确解与PINNs获得的解,展示了PINN学习在近似相同系统动力学时提供了一种不同的计算路径,同时指出PINNs的密集参数表示在高维情况下可能具有优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14584 2026-07-10 cs.CV 版本更新

Anatomically Guided Latent Diffusion for Brain MRI Progression Modeling

用于脑 MRI 进展建模的解剖学引导潜在扩散

Cheng Wan, Bahram Jafrasteh, Ehsan Adeli, Miaomiao Zhang, Qingyu Zhao

机构 * Cornell University(康奈尔大学) Weill Cornell Medicine(韦尔医学院) Stanford University(斯坦福大学) University of Virginia(弗吉尼亚大学)

AI总结 研究旨在准确建模脑 MRI 进展,提出解剖学引导潜在扩散模型 AG-LDM,通过融合多因素简化训练流程,经实验验证其在图像质量、误差降低及特征捕捉等方面表现优异,是可靠的脑 MRI 进展建模框架。

Comments 24 pages, 7 figures, 7 tables. Code available at https://github.com/JornyWan/AG-LDM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18077 2026-07-10 stat.ML cs.LG econ.EM stat.AP 版本更新

Bayesian Deep Learning for Discrete Choice

用于离散选择的贝叶斯深度学习

Daniel F. Villarraga, Ricardo A. Daziano

机构 * School of Civil and Environmental Engineering, Cornell University(土木与环境工程学院,康奈尔大学) Industrial Engineering Department, Universidad de Los Andes(工业工程系,洛斯安德斯大学)

AI总结 研究针对离散选择模型,提出将深度学习模型与近似贝叶斯推理方法集成的架构,数据有限时可减轻过拟合等问题,通过模拟研究和实证案例展示该方法,评估相关指标,提升离散选择模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22136 2026-07-08 cs.LG cs.AI cs.CR cs.SE 版本更新

StepShield: When, Not Whether to Intervene on Rogue Agents

StepShield:何时干预流氓代理,而非是否干预

Gloria Felicia, Zitha Sasindran, Jinfeng He, Michael Eniolade, Hemant Kumar, Milan Hussain Angati

机构 * University of Virginia(弗吉尼亚大学) Indian Institute of Science, Bangalore(印度科学研究院,班加罗尔) Cornell University(康奈尔大学) University of the Cumberlands(库姆伯兰兹大学) University of Arizona(亚利桑那大学) California State University, Northridge(加州州立大学,北岭分校)

AI总结 研究代理安全检测时机问题,引入StepShield基准及早期干预率(EIR),揭示取证陷阱,指出基于规则的护栏虽召回率高但时机不佳,现有方法无法兼顾高召回、低误报和及时干预,凸显逐步骤流氓检测待解决。

Comments 20 pages, 4 figures, 10 ablation studies. Code and data: https://github.com/glo26/stepshield

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20677 2026-07-07 cs.CR cs.CL 版本更新

Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models

基于学习的自动化对抗红队测试用于大型语言模型的鲁棒性评估

Zhang Wei, Hanxuan Chen, Peilu Hu, Zhenyuan Wei, Chenwei Liang, Jiayi Gu, Wenqian Weng, Jacqueline Pang, Hao Yan, Li Mei, Shengning Lang, Kuan Lu, Xi Xiao, Zhimo Han, Yijin Wang, Yichao Zhang, Chen Yang, Zhenyu Yu, Riyang Bao, Xinyuan Song, Junfeng Hao, Mu-Jiang-Shan Wang

机构 * Independent Researcher(独立研究者) Hunan University(湖南大学) Shenzhen Kaihong Digital Industry Development Co., Ltd.(深圳凯鸿数字产业开发有限公司) Central University of Finance and Economics(中央财经大学) Chongqing University(重庆大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Stevens Institute of Technology(斯蒂文斯理工学院) Cornell University(康奈尔大学) Oak Ridge National Laboratory(橡树岭国家实验室) Zhengzhou University of Light Industry(郑州轻工业大学) Xidian University(西安电子科技大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校) AI Safety Research Lab, Institute of Advanced Computing(高级计算研究所人工智能安全研究实验室) University of Malaya(马来亚大学) Emory University(埃默里大学) Department of Nephrology, Affiliated Hospital of Guangdong Medical University(广东医学院附属医院肾内科)

AI总结 本文提出一种学习驱动的自动化红队测试框架,用于高效发现大型语言模型的漏洞,通过元提示引导的对抗性提示生成和分层执行检测流程,在六个威胁类别中实现标准化评估,实验发现47个漏洞,包括21个高严重性失败和12种新攻击模式。

Comments ACL ARR minor revision

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00819 2026-07-07 math.NA cs.LG cs.NA math.ST stat.TH 版本更新

A short tour of operator learning theory: Convergence rates, statistical limits, and open questions

算子学习理论简览:收敛速率、统计极限及开放问题

Simone Brugiapaglia, Nicola Rares Franco, Nicholas H. Nelsen

机构 * Concordia University(康科迪亚大学) Politecnico di Milano(米兰理工学院) Cornell University and The University of Texas at Austin(康奈尔大学和德克萨斯大学奥斯汀分校)

AI总结 综述算子学习、统计学习理论和逼近理论交叉领域进展,先回顾经验风险最小化的误差界,再从极小极大角度说明样本量的基本性能极限,最后讨论两者相互作用及相关开放问题。

Comments To appear in Numerical Mathematics and Advanced Applications (Proceedings of ENUMATH 2025); 13 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13465 2026-07-07 cs.AI cs.LG 版本更新

Graph Neural Networks are Heuristics

图神经网络是启发式算法

Yimeng Min, Carla P. Gomes

机构 * Department of Computer Science(计算机科学系) Cornell University(康奈尔大学) Ithaca, NY, USA(纽约州伊萨卡市)

AI总结 研究表明图神经网络在组合优化中辅助角色非固有,可自身作启发式算法。以欧几里得旅行商问题为例,训练无标签等的非自回归图神经网络,仅用可微哈密顿循环目标监督,模型能快速生成完整路线且多样,优于贪心基线。

Comments 12 pages, 3 tables with 2 figures, code repo included in the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23845 2026-07-07 cs.LG cs.AI cs.CL cs.CY 版本更新

Position: Use Sparse Autoencoders to Discover Unknowns

定位:使用稀疏自编码器发现未知因素

Kenny Peng, Rajiv Movva, Jon Kleinberg, Emma Pierson, Nikhil Garg

机构 * Cornell University(康奈尔大学) UC Berkeley(伯克利大学)

AI总结 研究围绕稀疏自编码器(SAEs)的不同观点,指出其虽在作用于已知概念时效果欠佳,但在发现未知概念上很强大,还给出了在机器学习及社会健康科学等方面的应用案例。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18546 2026-07-01 cs.LG eess.SP math.OC 版本更新

Wasserstein Distributionally Robust Risk-Sensitive Estimation via Conditional Value-at-Risk

基于条件风险值的Wasserstein分布鲁棒风险敏感估计

Feras Al Taha, Eilyan Bitar

机构 * Cornell University(康奈尔大学)

AI总结 本文提出了一种分布鲁棒的风险敏感估计方法,通过求解可解的半正定规划计算最小化最坏情况下的条件风险值,应用于批发电力价格预测,降低样本外风险值。

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11001 2026-07-01 cs.CY cs.AI 版本更新

RCTs for Frontier AI Governance: Methodological Challenges and Solutions for Human Uplift Studies

随机对照试验与人类提升研究:前沿AI评估的方法论挑战与实践解决方案

Patricia Paskov, Kevin Wei, Shen Zhou Hong, Dan Bateyko, Xavier Roberts-Gaal, Carson Ezell, Gailius Praninskas, Valerie Chen, Umang Bhatt, Ella Guest

机构 * RAND Johns Hopkins University(约翰霍普金斯大学) Cornell University(康奈尔大学) Harvard University(哈佛大学) University of Cambridge(剑桥大学) London School of Economics(伦敦经济学院)

AI总结 本文通过访谈16位专家,系统梳理了人类提升研究(测量AI对人类绩效影响)在随机对照试验中面临的方法论挑战,包括内部效度、外部效度和构念效度问题,并提出了相应的解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17581 2026-07-01 cs.LG cs.CV 版本更新

EgoCogNav: Cognition-aware Human Egocentric Navigation

EgoCogNav: 认知感知的人类自我中心导航

Zhiwen Qiu, Ziang Liu, Wenqian Niu, Tapomayukh Bhattacharjee, Saleh Kalantari

机构 * Cornell University(康奈尔大学) Georgia Institute of Technology(佐治亚理工学院)

AI总结 提出EgoCogNav多模态自我中心导航框架,联合预测路径不确定性、轨迹和头部运动,并引入CEN数据集,实验表明其能学习与人类扫描、犹豫等行为相关的感知不确定性。

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02379 2026-06-29 cs.CV 版本更新

Honey, I Shrunk the Arc de Triomphe!

亲爱的,我把凯旋门缩小了!

Yuanbo Xiangli, Hanyu Chen, Xueqing Tsang, Noah Snavely

机构 * Cornell University(康奈尔大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 针对单目度量几何估计中的“尺度坍缩”现象,通过构建新数据集MetricScenes并采用两阶段泊松补全方法提升深度图质量,微调MoGe-2模型显著缓解了尺度低估问题。

Comments Project page: https://metricscenes.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08648 2026-06-29 astro-ph.HE astro-ph.IM cs.LG hep-ph 版本更新

High-dimensional inference for the $γ$-ray sky with differentiable programming

高维推断用于γ射线天空的不同可微编程

Siddharth Mishra-Sharma, Tracy R. Slatyer, Yitian Sun, Yuqing Wu

机构 * Faculty of Computing Data Sciences, Boston University, Boston, MA 02215, USA The NSF AI Institute for Artificial Intelligence Center for Theoretical Physics -- a Leinweber Institute, Massachusetts Institute of Technology, Cambridge, MA 02139, USA Department of Physics, Harvard University, Cambridge, MA 02138, USA Center for Theoretical Physics, Massachusetts Institute of Technology, Cambridge, MA 02139, USA Trottier Space Institute \& Department of Physics, McGill University, Montreal, QC H3A 2T8, Canada Department of Physics, Cornell University, Ithaca, NY 14853, USA

AI总结 本文利用可微概率编程技术解决银河系中心γ射线异常问题,通过GPU加速和向量化方法构建可微前向模型和似然函数,实现对多种空间形态的灵活推断。

Comments 20 pages, 16 figures. Code available at https://github.com/smsharma/fermi-prob-prog. V2: Updated with a more complete set of coverge tests

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12390 2026-06-25 cs.LG cs.AI cs.NA math.NA 版本更新

Rational Neural Networks have Expressivity Advantages

有理神经网络具有表达优势

Maosen Tang, Alex Townsend

机构 * Center for Applied Mathematics, Cornell University, United States(应用数学中心,康奈尔大学,美国) Department of Mathematics, Cornell University, United States(数学系,康奈尔大学,美国)

AI总结 本文证明可训练低次有理激活函数网络在表达性和参数效率上优于常见固定激活函数,理论上有指数级参数节省,实验验证其与标准架构兼容。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21321 2026-06-24 cs.LG cs.AR math.OC 版本更新

Dynamic Symmetric Point Tracking: Tackling Non-ideal Reference in Analog In-memory Training

动态对称点跟踪:解决模拟内存训练中的非理想参考问题

Quan Xiao, Jindan Li, Zhaoxian Wu, Tayfun Gokmen, Tianyi Chen

机构 * Department of Electrical and Computer Engineering, Cornell University, New York, NY(康奈尔大学电气与计算机工程系) IBM T. J. Watson Research Center, Yorktown Heights, NY(IBM 沃森研究中心) Rensselaer Polytechnic Institute, Troy, NY(伦塞拉尔理工学院)

AI总结 针对模拟内存计算中非理想器件导致的权重更新偏向对称点问题,提出动态对称点估计方法,在训练中跟踪对称点并保证收敛,结合数字信号处理技术增强性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21056 2026-06-24 cs.LG cs.CL math.OC 版本更新

Bilevel Data Curation for LLM Fine-tuning: Offline Selection and Online Self-Refining Generation

大语言模型微调的双层数据策展:离线选择与在线自优化生成

Quan Xiao, Yutong Xuan, Gaowen Liu, Ramana Rao Kompella, Tianyi Chen

机构 * Cornell University(康奈尔大学) Cisco Research(思科研究)

AI总结 提出双层框架结合离线数据选择与在线自优化生成,提升微调数据质量,理论证明优于直接混合,实验验证效果。

Comments updated the theories and experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04374 2026-06-23 cs.RO cs.AI cs.HC 版本更新

Towards Considerate Human-Robot Coexistence: A Dual-Space Framework of Robot Design and Human Perception in Healthcare

迈向体贴的人机共存:医疗保健中机器人设计与人类感知的双空间框架

Yuanchen Bai, Zijian Ding, Ruixiang Han, Niti Parikh, Wendy Ju, Angelique Taylor

机构 * Cornell Tech(康奈尔科技) Cornell University(康奈尔大学) University of Maryland(马里兰大学)

AI总结 通过14周联合设计研究的后续访谈,识别人类感知空间的四个解释维度,提出人机共存作为设计空间与感知空间共同演化的循环,强调人类作为解释者和中介者的角色。

Comments This paper has been accepted for publication at the 35th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20819 2026-06-23 cs.LG cs.SY eess.SY stat.ML 版本更新

Achieving $\widetilde{O}(1/ε)$ Sample Complexity for Bilinear Systems Identification under Bounded Noises

在有限噪声下实现双线性系统辨识的 $\widetilde{O}(1/ε)$ 样本复杂度

Hongyu Yi, Chenbei Lu, Jing Yu

机构 * Department of Electrical and Computer Engineering, University of Washington(华盛顿大学电气与计算机工程系) Cornell University AI for Science Institute, Cornell University(康奈尔大学AI for Science研究所)

AI总结 针对有界对称对数凹扰动下的离散时间双线性系统,提出有限样本集成员辨识方法,证明可行参数集直径以 $\widetilde{\mathcal O}(1/\epsilon)$ 的样本复杂度收缩,并通过仿真验证了不确定性量化的优势。

Comments 14 pages, 2 figures. Accepted by IEEE Control Systems Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24160 2026-06-23 eess.IV cs.CV 版本更新

Beyond the LUMIR challenge: The pathway to foundational registration models

超越LUMIR挑战:走向基础配准模型

Junyu Chen, Shuwen Wei, Joel Honkamaa, Pekka Marttinen, Hang Zhang, Min Liu, Yichao Zhou, Zuopeng Tan, Zhuoyuan Wang, Yi Wang, Hongchao Zhou, Shunbo Hu, Yi Zhang, Qian Tao, Lukas Förner, Thomas Wendler, Bailiang Jian, Benedikt Wiestler, Tim Hable, Jin Kim, Dan Ruan, Frederic Madesta, Thilo Sentker, Wiebke Heyer, Lianrui Zuo, Yuwei Dai, Jing Wu, Jerry L. Prince, Harrison Bai, Yong Du, Yihao Liu, Alessa Hering, Reuben Dorent, Lasse Hansen, Mattias P. Heinrich, Aaron Carass

机构 * The Russell H. Morgan Department of Radiology(Russell H. Morgan放射科) Radiological Science, Johns Hopkins Medical School(约翰霍普金斯医学院放射科学) Department of Computer Science, Aalto University(阿尔托大学计算机科学系) Cornell University(康奈尔大学) Canon Medical Systems (China) Co. Ltd.(佳能医疗系统(中国)有限公司) School of Biomedical Engineering, Shenzhen University Medical School(深圳大学医学院生物医学工程学院) Department of Imaging Physics, Delft University of Technology(代尔夫特理工大学成像物理系) Technical University of Munich(慕尼黑技术大学) Radboud University Medical Center(拉德伯德大学医学中心) Inria, Paris, France(法国巴黎Inria)

AI总结 提出LUMIR挑战,通过大规模无监督脑MRI配准任务,验证深度学习方法在生成解剖合理变形场和跨域鲁棒性上的优势,推动通用医学图像配准基础模型的发展。

Comments Accepted to Medical Image Analysis ((c) MedIA). Code available at https://github.com/JHU-MedImage-Reg/LUMIR_L2R

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19836 2026-06-23 cs.CL physics.ed-ph 版本更新

A Framework for Deductive Semantic Content Analysis at Scale in Science Education Using Text Embeddings

使用文本嵌入进行科学教育中大规模演绎语义内容分析的框架

Jonas Timmann Mjaaland, Markus Fleten Kreutzer, Halvor Tyseng, Rebeckah K. Fussell, Gina Passante, N. G. Holmes, Anders Malthe-Sørenssen, Tor Ole B. Odden

机构 * Center for Interdisciplinary Education, University of Oslo(interdisciplinary Education 中心,奥斯陆大学) Laboratory of Atomic and Solid State Physics, Cornell University(原子与固体物理实验室,康奈尔大学) Department of Physics, California State University Fullerton(物理系,弗里蒙特加州州立大学) Center for Computing in Science Education, University of Oslo(科学教育计算中心,奥斯陆大学)

AI总结 提出基于文本嵌入的分类框架DeSCA,仅需少量示例即可实现大规模开放题编码,与人类编码者高度一致,并支持编码一致性审计。

Comments 47 pages plus supplementary information, 5 figures. Version 2 has been lightly edited and formatted to fit better with the field of science education research, including updating the title and adding a brief literature review of NLP methods applied to textual datasets in science education. Results are unchanged since original version

详情

展开后加载摘要…

URL PDF HTML 收藏