arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

NeurIPS

Conference on Neural Information Processing Systems · 会议 · Machine Learning

共收录 17321
2512.05119 2025-12-08 cs.IR cs.AI cs.CL

RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering

RAG-IGBench: 用于开放领域问答中基于检索增强生成的交错生成的创新评估

Rongyang Zhang, Yuqing Huang, Chengqiang Lu, Qimeng Wang, Yan Gao, Yi Wu, Yao Hu, Yin Xu, Wei Wang, Hao Wang, Enhong Chen

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学) Xiaohongshu Inc.(小红书公司) Xi’an Jiaotong University(西安交通大学)

AI总结 RAG-IGBench通过创新的评估指标和多模态数据,评估基于检索增强生成的交错生成任务,验证了模型在开放领域问答中的性能提升。

Comments 26 pages, 6 figures, NeurIPS 2025 D&B Track poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18526 2025-12-08 cs.AI cs.LG

Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models

反事实推理用于可操控的多元价值观对齐大语言模型

Hanze Guo, Jing Yao, Xiao Zhou, Xiaoyuan Yi, Xing Xie

机构 * Renmin University of China(中国人民大学) Microsoft Research Asia(微软亚洲研究院) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程研究中心,教育部)

AI总结 COUPLE通过反事实推理框架实现多元价值观对齐,解决现有方法在处理细粒度价值目标时的依赖性和优先级控制问题。

Comments NeurIPS 2025. 41 pages, 7 figures

Journal ref The Thirty-Ninth Annual Conference on Neural Information Processing Systems. (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19552 2025-12-08 cs.CV

iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning

iFinder: 结构化零样本视觉基于LLM的地面定位用于行车记录仪视频推理

Manyi Yao, Bingbing Zhuang, Sparsh Garg, Amit Roy-Chowdhury, Christian Shelton, Manmohan Chandraker, Abhishek Aich

机构 * NEC Laboratories, America(NEC美国实验室) University of California, Riverside(加州大学河滨分校) University of California, San Diego(加州大学圣地亚哥分校)

AI总结 iFinder通过结构化语义接地框架,利用行车记录仪视频中的关键线索提升LLM在驾驶视频推理中的性能。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16567 2025-12-08 cs.CV cs.AI

V-CECE: Visual Counterfactual Explanations via Conceptual Edits

V-CECE:通过概念编辑生成视觉反事实解释

Nikolaos Spanos, Maria Lymperaiou, Giorgos Filandrianos, Konstantinos Thomas, Athanasios Voulodimos, Giorgos Stamou

机构 * National Technical University of Athens(国家技术大学雅典)

AI总结 V-CECE通过概念编辑生成反事实解释,无需训练即可产生人类水平的可解释性结果。

Comments Accepted in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07339 2025-12-08 cs.RO cs.AI cs.LG

Real-Time Execution of Action Chunking Flow Policies

实时执行动作分块流策略

Kevin Black, Manuel Y. Galliker, Sergey Levine

机构 * Physical Intelligence(物理智能) UC Berkeley(伯克利大学)

AI总结 本文提出实时分块(RTC)方法,通过在执行当前动作分块的同时生成下一个分块,实现动作分块策略的实时异步执行,提升任务吞吐量和精确任务成功率。

Comments published in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01790 2025-12-08 cs.LG cs.CR

IF-GUIDE: Influence Function-Guided Detoxification of LLMs

IF-GUIDE:影响函数引导的LLM去毒化

Zachary Coalson, Juhan Bae, Nicholas Carlini, Sanghyun Hong

机构 * Oregon State University(俄勒冈州立大学) University of Toronto(多伦多大学) Anthropic

AI总结 IF-GUIDE通过影响函数主动识别并抑制训练数据中的有害标记,有效减少大语言模型的显性和隐性毒性。

Comments Accepted at NeurIPS 2025 [Poster]

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23471 2025-12-08 cs.SE

Synthesizing Performance Constraints for Evaluating and Improving Code Efficiency

为评估和提升代码效率合成性能约束

Jun Yang, Cheng-Chi Wang, Bogdan Alexandru Stoica, Kexin Pei

AI总结 WEDGE通过生成性能压力输入,提升代码优化效果,释放PERFFORGE测试以评估未来高效代码生成方法。

Comments Accepted by Neurips 2025 (main poster)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20583 2025-12-08 stat.ML cs.LG

Balancing Performance and Costs in Best Arm Identification

在最佳臂识别中平衡性能与成本

Michael O. Harding, Kirthevasan Kandasamy

机构 * Department of Statistics University of Wisconsin-Madison(统计学系威斯康星大学麦迪逊分校)

AI总结 本文提出了一种新的最佳臂识别方法,通过平衡性能和成本来优化学习过程,提出了 DBCARE 算法并展示了其在模拟模型中的优越性能。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18473 2025-12-08 math.OC

PDPO: Parametric Density Path Optimization

PDPO:参数密度路径优化

Sebastian Gutierrez Hernandez, Peng Chen, Haomin Zhou

AI总结 PDPO通过参数映射将无限维密度优化转化为有限维问题,有效解决多模态和高维路径优化问题,优于现有方法。

Comments 28 pages, 16 figures

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01592 2025-12-08 cs.CL cs.AI

AURA: A Diagnostic Framework for Tracking User Satisfaction of Interactive Planning Agents

AURA:一个用于跟踪交互式规划代理用户满意度的诊断框架

Takyoung Kim, Janvijay Singh, Shuhaib Mehri, Emre Can Acikgoz, Sagnik Mukherjee, Nimet Beyza Bozdag, Sumuk Shashidhar, Gokhan Tur, Dilek Hakkani-Tür

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 AURA提出一个诊断框架,用于跟踪交互式规划代理的用户满意度,通过评估代理行为阶段和中间行为来提升用户体验。

Comments NeurIPS 2025 MTI-LLM Workshop. Full version is under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04664 2025-12-08 cs.LG math.OC

Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data

隐式偏置与谱下降法和缪恩在多类可分数据中的表现

Chen Fan, Mark Schmidt, Christos Thrampoulidis

机构 * Department of Computer Science University of British Columbia(计算机科学系不列颠哥伦比亚大学) University of British Columbia Canada(不列颠哥伦比亚大学加拿大) CIFAR AI Chair (Amii) &(CIFAR人工智能主席(Amii)) Department of Electrical and Computer Engineering University of British Columbia(电气与计算机工程系不列颠哥伦比亚大学)

AI总结 本文研究了谱下降法和缪恩在多类可分数据中的隐式偏置,证明其收敛到最大化分类矩阵p-范数的边际解,并展示了Adam在预处理下的收敛特性。

Comments NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05078 2025-12-05 astro-ph.IM astro-ph.GA

Improving Posterior Inference of Galaxy Properties with Image-Based Conditional Flow Matching

通过基于图像的条件流匹配改进星系属性的后验推断

Mikaeel Yunus, John F. Wu, Benne W. Holwerda

AI总结 本文提出条件流匹配框架,结合图像与光度学数据提升星系属性推断精度,缓解尘埃-年龄退化问题。

Comments Accepted at NeurIPS 2025 ML4PS workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01926 2025-12-05 cs.AI cs.CL cs.LG

Large language models can learn and generalize steganographic chain-of-thought under process supervision

大语言模型可以在过程监督下学习和泛化隐写链式思维

Joey Skaf, Luis Ibanez-Lissen, Robert McCarthy, Connor Watts, Vasil Georgiv, Hannes Whittingham, Lorena Gonzalez-Manzano, David Lindner, Cameron Tice, Edward James Young, Puria Radmard

机构 * Mentorship for Alignment Research Students (MARS)(对齐研究 mentorship 项目) University College London(伦敦大学学院) Queen Mary University of London(伦敦女王学院) ML Alignment & Theory Scholars (MATS)(对齐与理论学者) Meridian Impact, Cambridge(剑桥 Meridian Impact) Universidad Carlos III de Madrid(马德里卡洛斯三世大学) Geodesic Research and University of Cambridge(Geodesic Research 和剑桥大学)

AI总结 大语言模型在过程监督下能够学习并泛化隐写链式思维,通过替换特定字符串实现推理编码,提升监控可靠性。

Comments 10 pages main text, 3 figures main text, 17 pages supplementary material, 1 figure supplementary material, accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09321 2025-12-05 cs.CV cs.AI cs.LG

DAVE: Diagnostic benchmark for Audio Visual Evaluation

DAVE: 音频视觉评估诊断基准

Gorjan Radevski, Teodora Popordanoska, Matthew B. Blaschko, Tinne Tuytelaars

机构 * KU Leuven(卢森堡大学)

AI总结 DAVE提出一个诊断基准,通过确保两种模态必要性及分解评估子类,解决多模态模型评估中的视觉偏见问题,提供更精准的模型诊断与改进指导。

Comments First two authors contributed equally

Journal ref NeurIPS 2025 Datasets & Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04550 2025-12-05 cs.CL cs.AI

AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees

AdmTree: 通过自适应语义树压缩长上下文

Yangning Li, Shaoshen Chen, Yinghui Li, Yankai Chen, Hai-Tao Zheng, Hui Wang, Wenhao Jiang, Philip S. Yu

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Peng Cheng Laboratory(鹏城实验室) University of Illinois Chicago(伊利诺伊大学芝加哥分校) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳))

AI总结 AdmTree通过自适应语义树实现高效上下文压缩,保留高语义保真度并减少计算开销。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04307 2025-12-05 cs.LG cs.AI

Evaluating Long-Context Reasoning in LLM-Based WebAgents

评估基于大语言模型的WebAgent的长上下文推理能力

Andy Chung, Yichi Zhang, Kaixiang Lin, Aditya Rawal, Qiaozi Gao, Joyce Chai

机构 * University of Michigan(密歇根大学) Amazon(亚马逊)

AI总结 本文评估了基于大语言模型的WebAgent在长上下文场景中的推理能力,发现随着上下文长度增加,性能显著下降,提出隐式RAG方法以改进任务执行。

Comments Accepted NeurIPS 25 LAW Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04241 2025-12-05 math.CO q-bio.NC

Covering Relations in the Poset of Combinatorial Neural Codes

组合神经码的覆盖关系

R. Amzi Jeffs, Trong-Thuc Trang

AI总结 本文研究了组合神经码的覆盖关系,探讨了凸神经码的实现问题,并提出了关于凸神经码与多面体凸神经码等价性的猜想。

Comments To appear in Proceedings of the 4th NeurIPS Workshop on Symmetry and Geometry in Neural Representations, Proceedings of Machine Learning Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04125 2025-12-05 cs.LG

ASCIIBench: Evaluating Language-Model-Based Understanding of Visually-Oriented Text

ASCIIBench: 评估基于语言模型的视觉文本理解

Kerry Luo, Michael Fu, Joshua Peguero, Husnain Malik, Anvay Patil, Joyce Lin, Megan Van Overborg, Ryan Sarmiento, Kevin Zhu

机构 * Algoverse AI Research(Algoverse AI研究院)

AI总结 ASCIIBench通过评估LLM生成ASCII艺术的性能,揭示了多模态表示的局限性,并推动了针对符号视觉模态的新方法发展。

Comments Accepted to The Thirty-Ninth Annual Conference on Neural Information Processing Systems (NeurIPS 2025): LLM Evaluation Workshop & Multimodal Algorithmic Reasoning Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04107 2025-12-05 cs.CY cs.AI cs.HC cs.LG

Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants

重新思考教育中的AI评估:TEACH-AI框架与生成AI助手的基准测试

Shi Ding, Brian Magerko

机构 * Expressive Machinery Lab(表达性机械实验室) Georgia Institute of Technology(佐治亚理工学院)

AI总结 本文提出TEACH-AI框架,旨在通过多视角重新定义教育中AI的有效性评估,促进包容性和长期影响。

Comments 6 pages, NeurIPS 2025 Responsible Foundation Models Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22037 2025-12-05 cs.CY

What AI Speaks for Your Community: Polling AI Agents for Public Opinion on Data Center Projects

人工智能为你的社区发声:通过AI代理收集数据中心项目公众意见

Zhifeng Wu, Yuelin Han, Shaolei Ren

AI总结 本文提出AI代理调查框架,利用大型语言模型评估社区对数据中心项目的意见,以指导负责任的AI发展。

Comments 35 Pages. Accepted to NeurIPS 2025 Workshop on Socially Responsible and Trustworthy Foundation Models (ResponsibleFM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07594 2025-12-05 hep-ex cs.LG

Locality-Sensitive Hashing-Based Efficient Point Transformer for Charged Particle Reconstruction

基于局部敏感哈希的高效点变换器用于带电粒子重建

Shitij Govil, Jack P. Rodgers, Yuan-Tang Chou, Siqi Miao, Amit Saha, Advaith Anand, Kilian Lieret, Gage DeZoort, Mia Liu, Javier Duarte, Pan Li, Shih-Chieh Hsu

机构 * Georgia Institute of Technology(佐治亚理工学院) Purdue University(普渡大学) University of Washington(华盛顿大学) Princeton University(普林斯顿大学) University of California San Diego(加州大学圣地亚哥分校)

AI总结 HEPTv2通过轻量级解码器消除聚类步骤,实现高效端到端推理,提升带电粒子轨迹重建的性能和效率。

Comments Accepted to NeurIPS 2025 Machine Learning and the Physical Sciences Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03163 2025-12-05 cs.CV cs.GR

ROGR: Relightable 3D Objects using Generative Relighting

ROGR:基于生成光照的可重照明3D物体

Jiapeng Tang, Matthew Levine, Dor Verbin, Stephan J. Garbin, Matthias Nießner, Ricardo Martin Brualla, Pratul P. Srinivasan, Philipp Henzler

机构 * Google Research(谷歌研究) Google Deepmind(谷歌DeepMind) Technical University of Munich(慕尼黑技术大学)

AI总结 ROGR通过生成光照模型实现可重照明的3D物体重建,采用双分支架构的光照条件NeRF,高效生成任意环境光照下的物体外观。

Comments NeurIPS 2025 Spotlight. Project page: https://tangjiapeng.github.io/ROGR

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01721 2025-12-05 cs.LG

Convolutional Monge Mapping between EEG Datasets to Support Independent Component Labeling

卷积蒙日映射在EEG数据集之间进行转换以支持独立成分标注

Austin Meek, Carlos H. Mendoza-Cardenas, Austin J. Brockmeier

机构 * Department of Computer and Information Sciences University of Delaware(计算机与信息科学系德克萨斯大学) Twitch Interactive Inc.(Twitch互动公司) Department of Electrical and Computer Engineering Department of Computer and Information Sciences University of Delaware(电气与计算机工程系计算机与信息科学系德克萨斯大学)

AI总结 本文提出了一种改进的卷积蒙日映射方法,通过两种新方法实现EEG数据集之间的映射,以提高独立成分分类的准确性。

Comments Code available at: https://github.com/cniel-ud/ICWaves; Accepted to NeurIPS 2025 Workshop on Learning from Time Series for Health

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16732 2025-12-05 cs.LG cs.AI stat.ML

Sequential Monte Carlo for Policy Optimization in Continuous POMDPs

连续部分可观测马尔可夫决策过程中的策略优化的序列蒙特卡洛方法

Hany Abdulsamad, Sahel Iqbal, Simo Särkkä

AI总结 本文提出了一种基于序列蒙特卡洛的策略优化方法,用于解决连续部分可观测马尔可夫决策过程中的探索与利用平衡问题。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15555 2025-12-05 cs.LG cs.AI stat.ML

Bayesian Concept Bottleneck Models with LLM Priors

具有LLM先验的贝叶斯概念瓶颈模型

Jean Feng, Avni Kothari, Luke Zier, Chandan Singh, Yan Shuo Tan

机构 * University of California, San Francisco(加州大学旧金山分校) Microsoft Research(微软研究院) National University of Singapore(新加坡国立大学)

AI总结 本文提出BC-LLM模型,利用贝叶斯框架和LLM作为先验,实现高效的概念提取和可解释性提升。

Comments 2025 Conference on Neural Information Processing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03750 2025-12-04 cs.LG cond-mat.mtrl-sci

Universally Converging Representations of Matter Across Scientific Foundation Models

跨科学基础模型中物质的普遍收敛表示

Sathya Edamadaka, Soojung Yang, Ju Li, Rafael Gómez-Bombarelli

AI总结 研究揭示科学基础模型在不同模态和数据集上对物质的普遍表示收敛性,表明模型学习了共同的物理现实表示,但受限于训练数据和归纳偏置。

Comments Oral spotlight at NeurIPS 2025 UniReps Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03678 2025-12-04 cs.LG

Feature-aware Modulation for Learning from Temporal Tabular Data

面向特征的调制:学习时序表格数据

Hao-Run Cai, Han-Jia Ye

机构 * School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学) National Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家重点实验室,南京大学)

AI总结 本文提出了一种面向特征的时间调制机制,通过调节特征表示的统计属性来平衡泛化性和适应性,有效应对时序表格数据中的时间偏移问题。

Comments 17 pages, 6 figures, 8 tables. NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03601 2025-12-04 cs.CV

Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene Understanding

Motion4D: 学习3D一致的运动和语义以实现4D场景理解

Haoran Zhou, Gim Hee Lee

AI总结 Motion4D通过整合2D先验和4D高斯点撒表示,提升3D一致性和语义一致性,实现更准确的4D场景理解。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12178 2025-12-04 cs.LG cond-mat.mtrl-sci

All that structure matches does not glitter

所有结构匹配都不发光

Maya M. Martirossyan, Thomas Egg, Philipp Hoellmer, George Karypis, Mark Transtrum, Adrian Roitberg, Mingjie Liu, Richard G. Hennig, Ellad B. Tadmor, Stefano Martiniani

机构 * Center for Soft Matter Research, Department of Physics, New York University(纽约大学软物质研究中心) Simons Center for Computational Physical Chemistry, Department of Chemistry, New York University(纽约大学计算物理化学simons中心) Department of Computer Science & Engineering, University of Minnesota(明尼苏达大学计算机科学与工程系) Department of Physics & Astronomy, Brigham Young University(BYU物理与天文学系) Department of Chemistry, University of Florida(佛罗里达大学化学系) Quantum Theory Project, University of Florida(佛罗里达大学量子理论项目) Department of Materials Science & Engineering, University of Florida(佛罗里达大学材料科学与工程系) Department of Aerospace Engineering & Mechanics, University of Minnesota(明尼苏达大学航空航天工程与力学系) Center for Neural Science, New York University(纽约大学神经科学中心) Courant Institute of Mathematical Sciences, New York University(纽约大学数学科学学院)

AI总结 本文针对晶体结构预测任务中数据集和评估指标的问题,提出改进数据集的修复方法和新的评估指标,以提高模型评估的准确性。

Comments Accepted at Thirty-Ninth Annual Conference on Neural Information Processing Systems (NeurIPS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07150 2025-12-04 stat.ML cs.AI cs.LG math.ST stat.ME stat.TH

Class conditional conformal prediction for multiple inputs by p-value aggregation

通过p值聚合实现多输入的条件分类置信预测

Jean-Baptiste Fermanian, Mohamed Hebiri, Joseph Salmon

AI总结 通过p值聚合改进多输入条件分类置信预测,提升预测集规模并保持覆盖概率保证。

Journal ref NeurIPS 2025, The Thirty-Ninth Annual Conference on Neural Information Processing Systems, Dec 2025, San Diego (CA), United States

详情

展开后加载摘要…

URL PDF HTML 收藏