arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

New York University(纽约大学)

2026-06-10 至 2026-06-10 共收录 13
2606.11166 2026-06-10 stat.OT cs.AI 新提交

Flaws in the LLM Automation Narrative

LLM自动化叙事中的缺陷

George Perrett, Javae Elliott, Jennifer Hill, Marc Scott

机构 * New York University(纽约大学)

AI总结 通过编写代码完成数据分析任务的新基准测试,发现前沿LLM在平均性能、方差和错误幅度上均不如人类专家,挑战了LLM达到人类专家水平的说法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10765 2026-06-10 cs.CL 新提交

ArabiGEE: A Hierarchical Taxonomy for Arabic Grammatical Error Explanation

ArabiGEE:阿拉伯语语法错误解释的层次分类体系

Khaled Elhady, Omar Kallas, Nizar Habash, Bashar Alhafni

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) New York University Abu Dhabi(纽约大学阿布扎克分校)

AI总结 提出首个基于显式错误类型的阿拉伯语语法错误解释层次分类体系,涵盖正字法、形态、句法和词汇四个维度,包含27种错误类型、140种修正类型和324种解释,并用于人工标注现有语料库以支持大语言模型的自动评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10460 2026-06-10 cs.CL cs.AI 新提交

LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake

LakeQA:百万级数据湖上的探索性问答基准

Haonan Wang, Jiaxiang Liu, Yurong Liu, Austin Senna Wijaya, Tianle Zhou, Eden Wu, Yijia Chen, Wanting You, Reya Vir, Daniela Pinto, Grace Fan, Yusen Zhang, Juliana Freire, Eugene Wu

机构 * Columbia University(哥伦比亚大学) New York University(纽约大学) Barnard College(巴纳德学院)

AI总结 提出LakeQA基准,要求LLM在9.5TB异构数据湖中搜索并多跳推理,GPT-5.2仅达18.37%精确匹配,挑战性强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10254 2026-06-10 cs.AI cs.CL 新提交

RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning

RealMath-Eval:为何SOTA裁判难以应对真实人类推理

Yiteng Mao, Kenan Xu, Yijia Lyu, Wenhao Li, Jianlong Chen, Xiangfeng Wang

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) East China Normal University(华东师范大学) New York University(纽约大学) Tongji University(同济大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

AI总结 提出RealMath-Eval基准,评估LLM裁判对真实学生数学解答的评分能力,发现与人类评分存在高均方误差,而合成数据上表现更好,揭示评估差距源于人类错误空间的多样性和高信息熵。

Comments Code available at https://github.com/RicharMd/RealMath-Eval , Data available at https://huggingface.co/datasets/RicharMd/RealMath-Eval

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10216 2026-06-10 cs.LG cs.AI 新提交

A Source Domain is All You Need: Source-Only Cross-OS Transfer Learning for APT Anomaly Detection via Semantic Alignment and Optimal Transport

一个源域足矣:基于语义对齐和最优传输的仅源域跨操作系统APT异常检测迁移学习

Sidahmed Benabderrahmanea, Petko Valtchev, James Cheney, Talal Rahwan

机构 * New York University, NYUAD, Division of Science, Computer Science Department(纽约大学,NYUAD,科学学院,计算机科学系) University of Quebec in Montreal, Computer Science Department, Montreal(魁北克大学蒙特利尔分校,计算机科学系,蒙特利尔) University of Edinburgh, School of Informatics, Edinburgh(爱丁堡大学,信息学院,爱丁堡)

AI总结 针对跨操作系统APT检测中目标域无标签的挑战,提出基于最优传输的仅源域异常评分框架,通过语义抽象和三种偏差通道实现零目标监督下的异常排序。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10094 2026-06-10 cs.AI 新提交

Predictive Assistance and the Temporal Dynamics of Exploratory Compression

预测性辅助与探索性压缩的时间动态

Balaraju Battu

机构 * European University Institute(欧洲大学学院) New York University Abu Dhabi(纽约大学阿布扎比分校)

AI总结 提出几何动力学框架,研究预测性AI如何通过外源探索性压缩改变认知探索的时间动态,发现持续稳定会降低探索响应性、曲率不对称积累导致滞后效应、早期干预限制后续探索多样性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10086 2026-06-10 cs.AI 新提交

Exploratory Responsiveness and Adaptive Rigidity under AI-Assisted Optimization

AI辅助优化下的探索响应性与适应性刚性

Balaraju Battu

机构 * European University Institute(欧洲大学研究所) New York University Abu Dhabi(纽约大学阿布扎比分校)

AI总结 本文提出AI辅助优化下的探索适应理论,通过动态框架分析预测辅助如何影响系统探索响应性,揭示收敛预测机制导致适应性降低、刚性增强,而探索增强机制则促进适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09936 2026-06-10 cs.LG cs.AI 新提交

One Lens, Many Worlds : A Capability-Typed Interface for World-Model Interpretability

一个镜头,多个世界:面向世界模型可解释性的能力类型接口

Bhavith Chandra Challagundla, Sanskar Pandey, Param Thakkar, Rishikesh Mallagundla, Yugandhar Reddy Gogireddy, Wenhao Lu, Hindol Roy Choudhury, Shravani Challagundla, Mohamed Deraz Nasr, Spursh Deshpande

机构 * New York University(纽约大学) Independent Researcher(独立研究者) Veermata Jijabai Technological Institute(韦尔玛塔·吉贾拜技术学院) Mercity University of Southern California(南加州大学) Independent Researcher, MIT(麻省理工学院独立研究者) Independent Researcher, GITAM(GITAM独立研究者) Georgia Institute of Technology(佐治亚理工学院)

AI总结 提出WorldModelLens,通过能力类型适配器统一不同世界模型(如PlaNet、IRIS、I-JEPA)的可解释性分析,避免重复实现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11138 2026-06-10 cs.LG cs.NA math.NA 新提交

First-Order Trajectory Matching: Fast Ensemble Predictions of Chaotic, Turbulent, Stochastic Systems

一阶轨迹匹配:混沌、湍流、随机系统的快速集成预测

Shreya Jha, Timo Schorlepp, Nicholas Geissler, Jules Berman, Benjamin Peherstorfer

机构 * Courant Institute of Mathematical Sciences, New York University(纽约大学库朗数学科学研究所)

AI总结 提出一阶轨迹匹配(FTM)方法,通过学习随机系统轨迹的一阶局部概率质量输运,实现低成本的集成预测,并捕捉通量、环流等轨迹量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09639 2026-06-10 cs.LG stat.ML 版本更新

Blind denoising diffusion models and the blessings of dimensionality

盲去噪扩散模型与维度的祝福

Zahra Kadkhodaie, Aram-Alexandre Pooladian, Sinho Chewi, Eero Simoncelli

机构 * Flatiron Institute, Simons Foundation(Flatiron研究院,Simons基金会) Foundations of Data Science, Yale University(数据科学基础,耶鲁大学) Department of Statistics and Data Science, Yale University(统计与数据科学系,耶鲁大学) Ctr. for Neural Science & Courant Institute, New York University(神经科学中心及Courant学院,纽约大学)

AI总结 提出盲去噪扩散模型(BDDM),通过不向神经网络传递噪声幅度来简化设计,并在数据内在维度低于环境维度的假设下证明其正确性,实验显示自适应方案的优势。

Comments 39 pages, 13 figures; Accepted to ICML 2025 FoGen workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09620 2026-06-10 cs.HC cs.AI cs.CY

Full Disclosure, Less Trust? How the Level of Detail about AI Use in News Writing Affects Readers' Trust

全面披露,更少信任?新闻写作中AI使用细节程度如何影响读者信任

Pooja Prajod, Hannes Cools, Thomas Röggla, Karthikeya Puttur Venkatraj, Amber Kusters, Alia ElKattan, Pablo Cesar, Abdallah El Ali

机构 * Centrum Wiskunde & Informatica(数学与信息学中心) University of Amsterdam(阿姆斯特丹大学) New York University(纽约大学) TU Delft(代尔夫特理工大学) Utrecht University(乌得勒支大学)

AI总结 研究探讨新闻写作中AI使用细节披露程度对读者信任的影响,发现详细披露会降低信任,但促使更多读者核查信息源,揭示透明度与信任之间的权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22017 2026-06-10 eess.IV cs.CV 版本更新

Cyst-X: A Multi-Center MRI Benchmark and Federated Learning Framework for Malignancy-Risk Stratification of Pancreatic Cystic Neoplasm

Cyst-X:用于胰腺囊性肿瘤恶性风险分层的多中心MRI基准与联邦学习框架

Hongyi Pan, Gorkem Durak, Elif Keles, Ziliang Hong, Deniz Seyithanoglu, Zheyuan Zhang, Alpay Medetalibeyoglu, Halil Ertugrul Aktas, Andrea Mia Bejar, Yavuz Taktak, Gulbiz Dagoglu Kartal, Mehmet Sukru Erturk, Timurhan Cebeci, Yury Velichko, Lili Zhao, Emil Agarunov, Federica Proietto Salanitri, Concetto Spampinato, Pallavi Tiwari, Ziyue Xu, Sachin Jambawalikar, Ivo G. Schoots, Marco J. Bruno, Chenchan Huang, Candice W. Bolan, Tamas Gonda, Frank H. Miller, Rajesh N. Keswani, Michael B. Wallace, Ulas Bagci

机构 * Machine & Hybrid Intelligence Lab, Department of Radiology, Northwestern University(机器与混合智能实验室,放射科,西北大学) Istanbul Faculty of Medicine, Istanbul University(伊斯坦布尔大学医学学院) Department of Biomedical Engineering and Radiology, University of Wisconsin-Madison(生物医学工程与放射科,威斯康星大学麦迪逊分校) Department of Preventive Medicine, Northwestern University(预防医学系,西北大学) Division of Gastroenterology and Hepatology, New York University(消化内科与肝病科,纽约大学) Department of Electrical, Electronic and Computer Engineering, University of Catania(电气、电子和计算机工程系,卡塔尼亚大学) NVIDIA Department of Radiology, Columbia University(放射科,哥伦比亚大学) Department of Radiology and Nuclear Medicine, Erasmus Medical Center(放射科与核医学科,埃因霍温医学院) Department of Gastroenterology and Hepatology, Erasmus Medical Center(消化内科与肝病科,埃因霍温医学院) Department of Radiology, New York University(放射科,纽约大学) Division of Gastroenterology and Hepatology, Mayo Clinic Florida(消化内科与肝病科,迈阿密诊所佛罗里达分部) Department of Gastroenterology and Hepatology, Northwestern University(消化内科与肝病科,西北大学)

AI总结 提出Cyst-X,一个多中心MRI基准和联邦学习框架,用于IPMN恶性风险分层,结合PanSegNet分割器和3D DenseNet-121分类器,在内部交叉验证中达到0.85的AUC,性能与放射科医生相当。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14753 2026-06-10 cs.CV cs.LG 版本更新

Cost-Aware Routing for Efficient Text-To-Image Generation

面向文本到图像生成的高效路由:成本感知方法

Qinchan Li, Kenneth Chen, Changyue Su, Wittawat Jitkrittum, Qi Sun, Patsorn Sangkloy

机构 * Tandon School of Engineering, New York University(纽约大学Tandon工程学院) Google(谷歌) Eigen 4D Inc.(Eigen 4D公司)

AI总结 提出成本感知路由框架,根据提示复杂度自动选择不同去噪步数或模型,在保证高质量的同时降低计算成本,优于单一模型。

Comments Accepted by TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏