arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高校专区

Massachusetts Institute of Technology(麻省理工学院)

2026-04-03 至 2026-04-03 共收录 9
2604.02230 2026-04-03 cs.AI

Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs

回答错误的问题:用于LLMs中退避的推理轨迹逆向

Abinitha Gourabathina, Inkit Padhi, Manish Nagireddy, Subhajit Chaudhury, Prasanna Sattigeri

机构 * MIT(麻省理工学院) IBM Research(IBM研究院) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)

AI总结 本文提出轨迹逆向方法,通过分析模型推理轨迹来改进LLMs的退避能力,实验表明其在多个数据集上显著提升退避性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15387 2026-04-03 math.OC cs.LG

Multi-Timescale Primal Dual Hybrid Gradient with Application to Distributed Optimization

多时间尺度对偶双极梯度方法及其在分布式优化中的应用

Junhui Zhang, Patrick Jaillet

机构 * Operations Research Center, MIT(麻省理工学院运筹学研究中心) Department of Electrical Engineering and Computer Science, MIT(麻省理工学院电气工程与计算机科学系)

AI总结 本文提出多时间尺度PDHG方法及其加速变体,通过混合Bregman散度和多时间尺度外推实现不同对偶块的任意更新率收敛,提升异构优化和通信成本下的效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01978 2026-04-03 math.PR cs.LG stat.ML

Homogenized Transformers

均质化变换器

Hugo Koubbi, Borjan Geshkovski, Philippe Rigollet

机构 * Department of Mathematics, Massachusetts Institute of Technology(麻省理工学院数学系)

AI总结 研究深度多头自注意力机制的随机模型,证明在适当标度下,残差流的动力学具有非平凡均质极限,揭示了在均场 regime 中的随机非线性福克-计划克方程,以及在高斯设定下的表示崩溃现象。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01279 2026-04-03 cs.LG cs.AI hep-th math.OC

Sven: Singular Value Descent as a Computationally Efficient Natural Gradient Method

Sven:奇异值下降作为计算高效的自然梯度方法

Samuel Bright-Thonney, Thomas R. Harvey, Andre Lukas, Jesse Thaler

机构 * Massachusetts Institute of Technology(麻省理工学院) The NSF Institute for Artificial Intelligence and Fundamental Interactions(美国国家科学基金会人工智能与基本相互作用研究所) University of Oxford(牛津大学) Institut des Hautes Études Scientifiques(高等科学研究所) CEA Paris-Saclay(巴黎-萨克雷CEA)

AI总结 Sven是一种基于损失函数自然分解的神经网络优化算法,通过截断奇异值分解实现高效参数更新,优于传统自然梯度方法,在回归任务中优于Adam,且比LBFGS更高效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22828 2026-04-03 cs.AI q-bio.NC

Fast dynamical similarity analysis

快速动态相似性分析

Arman Behrad, Mitchell Ostrow, Mohammad Taha Fakharian, Ila Fiete, Christian Beste, Shervin Safavi

机构 * K. Lisa Yang Integrative Computational Neuroscience (ICoN), Massachusetts Institute of Technology(K. Lisa Yang 综合计算神经科学中心,麻省理工学院)

AI总结 本文提出fastDSA,一种高效且准确的非线性动态系统相似性度量方法,结合几何方法的效率与动态方法的准确性,适用于大规模比较。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18123 2026-04-03 cs.CV cs.AI cs.CL cs.LG

Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models

偏差是子空间,而非坐标:视觉-语言模型中事后去偏的几何重思

Dachuan Zhao, Weiyue Li, Zhenda Shen, Yushu Qiu, Bowen Xu, Haoyu Chen, Yongchao Chen

机构 * Harvard University(哈佛大学) MIT(麻省理工学院)

AI总结 本文提出SPD框架,通过几何方法识别并去除线性可解码的偏子空间,提升视觉-语言模型的公平性与任务性能。

Comments Accepted at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03432 2026-04-03 math.OC cs.GT cs.LG

A Polynomial-Time Algorithm for Variational Inequalities under the Minty Condition

变分不等式在 Minty 条件下的多项式时间算法

Ioannis Anagnostides, Gabriele Farina, Tuomas Sandholm, Brian Hu Zhang

机构 * Carnegie Mellon University(卡内基梅隆大学) Massachusetts Institute of Technology(麻省理工学院) Strategy Robot, Inc.(Strategy Robot公司) Strategic Machine, Inc.(Strategic Machine公司) Optimized Markets, Inc.(Optimized Markets公司)

AI总结 本文提出在 Minty 条件下求解 ε-变分不等式的新算法,具有多项式时间复杂度,解决了传统方法中依赖于 1/ε 的指数问题,并展示了其在博弈论中的应用。

Comments V3 polishes the writing and makes a correction to Theorem 5.2

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12215 2026-04-03 cs.LG

Automatic selection of the best neural architecture for time series forecasting

时间序列预测中最佳神经架构的自动选择

Qianying Cao, Shanqing Liu, Alan John Varghese, Jerome Darbon, Michael Triantafyllou, George Em Karniadakis

机构 * Division of Applied Mathematics, Brown University(布朗大学应用数学系) School of Engineering, Brown University(布朗大学工程学院) Department of Mechanical Engineering, Massachusetts Institute of Technology(麻省理工学院机械工程系)

AI总结 本文提出一个灵活的自动框架,通过整合LSTM、GRU、多头注意力和SSM模块,系统设计和评估多样化网络架构,通过多目标优化确定最佳模型配置,验证了在不同应用场景中,最佳模型是根据用户定义的偏好函数和评估目标定制的复合架构。

Comments 35 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10095 2026-04-03 cs.CV

Human Insights Driven Latent Space for Different Driving Perspectives: A Unified Encoder for Efficient Multi-Task Inference

由人类洞察驱动的潜在空间:一种统一编码器用于高效的多任务推理

Huy-Dung Nguyen, Anass Bairouk, Mirjana Maras, Wei Xiao, Tsun-Hsuan Wang, Patrick Chareyre, Ramin Hasani, Marc Blanchon, Daniela Rus

机构 * Hybrid Intelligence, part of Capgemini Engineering(Hybrid Intelligence,Capgemini Engineering 旗下) Computer Science and Artificial Intelligence Laboratory (CSAIL), MIT(麻省理工学院计算机科学与人工智能实验室(CSAIL))

AI总结 本文提出一种统一编码器,通过多任务学习提升自动驾驶中的视觉感知任务性能,并在转向估计中展示出优于其他方法的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏