arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Pennsylvania(宾夕法尼亚大学)

共收录 961
2602.01428 2026-02-24 cs.LG cs.CR

Improving the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language Models

提升语言模型中水印强度与推测采样效率之间的权衡

Weiqing He, Xiang Li, Li Shen, Weijie Su, Qi Long

机构 * University of Pennsylvania(宾夕法尼亚大学)

AI总结 本文提出了一种提升语言模型中水印强度与推测采样效率之间权衡的方法,通过定量指标和系统机制实现两者的平衡。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13904 2026-02-24 cs.CL

Calibrating Large Language Models with Sample Consistency

通过样本一致性校准大型语言模型

Qing Lyu, Kumar Shridhar, Chaitanya Malaviya, Li Zhang, Yanai Elazar, Niket Tandon, Marianna Apidianaki, Mrinmaya Sachan, Chris Callison-Burch

机构 * University of Pennsylvania(宾夕法尼亚大学) ETH Zurich(苏黎世联邦理工学院) University of Washington(华盛顿大学) Allen Institute for AI(人工智能研究院)

AI总结 通过样本一致性方法提升大型语言模型的校准效果,改进其预测置信度评估和性能表现。

Comments AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25124 2026-02-23 cs.RO

Safe Planning in Unknown Environments Using Conformalized Semantic Maps

在未知环境中使用符合化语义地图进行安全规划

David Smith Sundarsingh, Yifei Li, Tianji Tang, George J. Pappas, Nikolay Atanasov, Yiannis Kantaros

机构 * Department of Electrical and Systems Engineering, Washington University in St. Louis(华盛顿大学圣路易斯分校电气与系统工程系) Department of Electrical and Computer Engineering, University of California, San Diego(加州大学圣地亚哥分校电气与计算机工程系) Department of Electrical and Systems Engineering, University of Pennsylvania(宾夕法尼亚大学电气与系统工程系)

AI总结 本文提出了一种在未知环境中使用符合化语义地图进行安全规划的方法,能够实现用户指定的任务完成率,无需依赖传感器模型或噪声知识。

Comments 8 pages, 5 figures, 2 algorithms, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17592 2026-02-23 astro-ph.IM cs.LG

AstroMLab 4: Benchmark-Topping Performance in Astronomy Q&A with a 70B-Parameter Domain-Specialized Reasoning Model

AstroMLab 4: 在天文学问答中通过700亿参数领域专用模型实现基准顶级性能

Tijmen de Haan, Yuan-Sen Ting, Tirthankar Ghosal, Tuan Dung Nguyen, Alberto Accomazzi, Emily Herron, Vanessa Lama, Rui Pan, Azton Wells, Nesar Ramachandra

机构 * Institute of Particle Nuclear Studies (IPNS), High Energy Accelerator Research Organization (KEK), Tsukuba, Ibaraki 305-0801, Japan International Center for Quantum-field Measurement Systems for Studies of the Universe Particles (QUP-WPI), High Energy Accelerator Research Organization (KEK), Tsukuba, Ibaraki 305-0801, Japan Department of Astronomy, The Ohio State University, Columbus, OH, USA Center for Cosmology AstroParticle Physics (CCAPP), The Ohio State University, Columbus, OH, USA National Center for Computational Sciences, Oak Ridge National Laboratory, Oak Ridge, TN, USA Department of Computer Information Science, University of Pennsylvania, Philadelphia, PA, USA Center for Astrophysics, Harvard \& Smithsonian, Cambridge, MA, USA Siebel School of Computing Data Science, University of Illinois at Urbana-Champaign, Urbana-Champaign, IL, USA Computational Science Division, Argonne National Laboratory, Lemont, IL, USA

AI总结 AstroSage-Llama-3.1-70B通过700亿参数领域专用模型在天文学问答中实现顶级性能,优于GPT-5.2等通用模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17633 2026-02-20 cs.LG cs.AI stat.ML

When to Trust the Cheap Check: Weak and Strong Verification for Reasoning

何时相信廉价检查:弱验证与强验证用于推理

Shayan Kiyani, Sima Noorani, George Pappas, Hamed Hassani

机构 * University of Pennsylvania(宾夕法尼亚大学)

AI总结 本文提出弱-强验证策略,通过两个阈值结构控制接受和拒绝错误,以提高LLM推理的可靠性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16690 2026-02-19 stat.ME cs.LG stat.ML

Synthetic-Powered Multiple Testing with FDR Control

由合成数据驱动的FDR控制多重检验

Yonghoon Lee, Meshi Bashari, Edgar Dobriban, Yaniv Romano

机构 * Department of Statistics and Data Science, The Wharton School, University of Pennsylvania, USA(统计与数据科学系,沃顿商学院,宾夕法尼亚大学) Department of Electrical and Computer Engineering, Technion IIT, Israel(电气工程与计算机科学系,技术学院(IIT),以色列) Department of Computer Science, Technion IIT, Israel(计算机科学系,技术学院(IIT),以色列)

AI总结 SynthBH通过利用合成数据提高多重检验的样本效率和功效,同时在不依赖真实数据有效性的情况下控制FDR。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20063 2026-02-19 cs.GT cs.CY cs.LG

Strategic Hiring under Algorithmic Monoculture

算法单一性下的战略性招聘

Jackie Baek, Hamsa Bastani, Shihan Chen

机构 * Stern School of Business, New York University(纽约大学斯特恩商学院) Wharton School, University of Pennsylvania(宾夕法尼亚大学沃顿商学院) Graduate Group in Applied Mathematics and Computational Science, University of Pennsylvania(宾夕法尼亚大学应用数学与计算科学研究生小组)

AI总结 本文研究了在算法单一性劳动力市场中,企业通过共同算法评估竞争导致的拥堵问题,提出均衡策略能显著提升社会福利,需平台揭示拥堵信息以实现最优结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16206 2026-02-19 cs.RO cs.SY eess.SY

Nonplanar Model Predictive Control for Autonomous Vehicles with Recursive Sparse Gaussian Process Dynamics

非平面自主车辆模型预测控制

Ahmad Amine, Kabir Puri, Viet-Anh Le, Rahul Mangharam

机构 * Department of Electrical & Systems Engineering, University of Pennsylvania(电气与系统工程系,宾夕法尼亚大学)

AI总结 本文提出了一种基于递归稀疏高斯过程的非平面模型预测控制框架,用于提升自主车辆在复杂地形中的路径跟踪性能。

Comments 6 pages, 5 figures. Accepted to IEEE Intelligent Vehicles Symposium (IV), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16187 2026-02-19 cs.RO cs.AI cs.SY eess.SY

SIT-LMPC: Safe Information-Theoretic Learning Model Predictive Control for Iterative Tasks

SIT-LMPC:用于迭代任务的安全信息论学习模型预测控制

Zirui Zang, Ahmad Amine, Nick-Marios T. Kokolakis, Truong X. Nghiem, Ugo Rosolia, Rahul Mangharam

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Central Florida(佛罗里达大学) Lyric

AI总结 SIT-LMPC通过信息论方法和归一化流学习,实现对迭代任务中安全性和性能的平衡优化。

Comments 8 pages, 5 figures. Published in IEEE RA-L, vol. 11, no. 1, Jan. 2026. Presented at ICRA 2026

Journal ref IEEE Robotics and Automation Letters, vol. 11, no. 1, pp. 986-993, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04072 2026-02-19 cs.CL cs.HC

Toward Beginner-Friendly LLMs for Language Learning: Controlling Difficulty in Conversation

迈向语言学习的初学者友好型LLM:对话中的难度控制

Meiqing Jin, Liam Dugan, Chris Callison-Burch

机构 * University of Pennsylvania(宾夕法尼亚大学)

AI总结 本文提出通过可控生成技术提升初学者与LLM对话的可理解性,实验表明该方法能显著提高输出清晰度,并引入新的评估指标支持语言学习研究。

Comments EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13949 2026-02-17 cs.LG cs.AI

Experiential Reinforcement Learning

经验强化学习

Taiwei Shi, Sihao Chen, Bowen Jiang, Linxin Song, Longqi Yang, Jieyu Zhao

机构 * University of Southern California(南加州大学) Microsoft(微软公司) University of Pennsylvania(宾夕法尼亚大学)

AI总结 经验强化学习通过显式经验-反思-巩固循环提升语言模型在稀疏奖励环境中的学习效率和性能。

Comments 26 pages, 9 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18784 2026-02-17 cs.LG cs.NA eess.SP math.NA math.ST stat.ML stat.TH

Denoising diffusion probabilistic models are optimally adaptive to unknown low dimensionality

去噪扩散概率模型在未知低维性下最优适应

Zhihan Huang, Yuting Wei, Yuxin Chen

机构 * Department of Statistics and Data Science, the Wharton School, University of Pennsylvania(统计与数据科学系,沃顿商学院,宾夕法尼亚大学)

AI总结 该研究证明了DDPM在未知低维数据下具有最优适应性,其迭代复杂度与内在维度线性相关,且与KL散度度量一致。

Comments Accepted to Mathematics of Operations Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11907 2026-02-17 cs.LG cs.SI

GraphFM: A generalist graph transformer that learns transferable representations across diverse domains

GraphFM: 一种通用的图变换器,能够在不同领域中学习可迁移的表示

Divyansha Lachi, Mehdi Azabou, Vinam Arora, Eva Dyer

机构 * University of Pennsylvania(宾夕法尼亚大学) Columbia University(哥伦比亚大学)

AI总结 GraphFM是一种通用图变换器,通过多图预训练学习可迁移的表示,提升跨不同图结构和任务的性能。

Journal ref Transactions on Machine Learning Research, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04166 2026-02-16 cs.LG stat.CO stat.ML

N$^2$: A Unified Python Package and Test Bench for Nearest Neighbor-Based Matrix Completion

N²:一种统一的Python包和测试平台,用于基于最近邻的矩阵补全

Caleb Chin, Aashish Khubchandani, Harshvardhan Maskara, Kyuseong Choi, Jacob Feitelberg, Albert Gong, Manit Paul, Tathagata Sadhukhan, Anish Agarwal, Raaz Dwivedi

机构 * Cornell University(康奈尔大学) Columbia University(哥伦比亚大学) University of Pennsylvania(宾夕法尼亚大学)

AI总结 本文提出N²,一种统一的Python包和测试平台,用于基于最近邻的矩阵补全,展示了新的NN变体和现实世界数据集的基准测试,证明NN方法在现实应用中的优越性。

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12499 2026-02-16 cs.LG

A Theoretical Analysis of Mamba's Training Dynamics: Filtering Relevant Features for Generalization in State Space Models

Mamba训练动态的理论分析:在状态空间模型中筛选相关特征以实现泛化

Mugunthan Shandirasegaran, Hongkang Li, Songyang Zhang, Meng Wang, Shuai Zhang

机构 * New Jersey Institute of Technology(新泽西理工学院) University of Pennsylvania(宾夕法尼亚大学) University of Louisiana at Lafayette(路易斯安那州立大学拉法叶分校)

AI总结 本文分析了Mamba模型的训练动态,证明其通过选择性递归实现特征筛选,提升泛化能力,为Transformer解释提供了理论依据。

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02083 2026-02-16 cs.CR cs.AI cs.CY

Watermarking Discrete Diffusion Language Models

对离散扩散语言模型进行水印标记

Avi Bagchi, Akhil Bhimaraju, Moulik Choraria, Daniel Alabi, Lav R. Varshney

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Illinois Urbana–Champaign(伊利诺伊大学厄巴纳-香槟分校) Stony Brook University(石溪大学)

AI总结 本文提出了一种针对离散扩散语言模型的水印标记方法,通过保持分布的Gumbel-max采样技巧和序列位置播种随机性,实现了可靠检测和无失真水印。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13458 2026-02-13 stat.ML cs.AI cs.LG math.ST stat.TH

Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs

标签还是偏好?在AI生成输出上的人类判断下的预算受限学习

Zihan Dong, Xiaotian Hou, Ruijia Wu, Linjun Zhang

机构 * Rutgers University(罗切斯特大学) University of Pennsylvania(宾夕法尼亚大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 本文提出PCAL方法,通过半参数推断框架解决AI生成输出中预算受限学习问题,提供统计高效且鲁棒的估计器。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13811 2026-02-13 cs.CR cs.LG

Can LLMs Handle WebShell Detection? Overcoming Detection Challenges with Behavioral Function-Aware Framework

LLMs能否处理WebShell检测?通过行为功能感知框架克服检测挑战

Feijiang Han, Jiaming Zhang, Chuyi Deng, Jianheng Tang, Yunhuai Liu

机构 * University of Pennsylvania(宾夕法尼亚大学) Central South University(中南大学) Peking University(北京大学)

AI总结 基于行为功能感知框架,研究评估了七种LLM在WebShell检测中的表现,发现大模型精度高但召回率低,提出BFAD框架提升检测性能,F1值平均提升13.82%。

Comments Published as a conference paper at COLM 2025 (The new version has been polished and expanded with more detailed future work ideas)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10233 2026-02-12 cs.RO

Proactive Local-Minima-Free Robot Navigation: Blending Motion Prediction with Safe Control

主动避障的机器人导航:融合运动预测与安全控制

Yifan Xue, Ze Zhang, Knut Åkesson, Nadia Figueroa

机构 * Department of Mechanical Engineering and Applied Mechanics, University of Pennsylvania(机械工程与应用力学系,宾夕法尼亚大学) Department of Electrical Engineering, Chalmers University of Technology(电气工程系,查尔姆斯理工大学)

AI总结 本文提出一种结合运动预测与安全控制的主动避障导航方法,通过在线学习屏障函数和自适应参数调节,提升机器人在复杂动态环境中的安全性和效率。

Comments Co-first authors: Yifan Xue and Ze Zhang; Accepted by IEEE RA-L 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11739 2026-02-12 cs.CL cs.AI

ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without Training

ZeroTuning: 解锁初始标记的潜力以无需训练的方式提升大语言模型

Feijiang Han, Xiaodong Yu, Jianheng Tang, Delip Rao, Weihua Du, Lyle Ungar

机构 * University of Pennsylvania(宾夕法尼亚大学) AMD(AMD公司) Peking University(北京大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 ZeroTuning通过仅调整初始标记的注意力实现无需训练的大语言模型性能提升,适用于多种任务和场景。

Comments ICLR 2026 Accepted Version: proofread, introduction rewritten, additional experiments and appendix material added

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10017 2026-02-11 cs.CL

SCORE: Specificity, Context Utilization, Robustness, and Relevance for Reference-Free LLM Evaluation

SCORE:特定性、上下文利用、鲁棒性与相关性用于无参考LLM评估

Homaira Huda Shomee, Rochana Chaturvedi, Yangxinyu Xie, Tanwi Mallick

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Argonne National Laboratory(阿贡国家实验室) University of Pennsylvania(宾夕法尼亚大学)

AI总结 本文提出了一种无参考评估框架,用于评估LLM在高风险领域任务中的特定性、鲁棒性、相关性和上下文利用,通过精心编纂的数据集和人工评估验证了多指标评估的必要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09158 2026-02-11 cs.LG cs.AI

What do Geometric Hallucination Detection Metrics Actually Measure?

几何幻觉检测度量实际上测量什么?

Eric Yeats, John Buckheit, Sarah Scullen, Brendan Kennedy, Loc Truong, Davis Brown, Bill Kay, Cliff Joslyn, Tegan Emerson, Michael J. Henry, John Emanuello, Henry Kvinge

机构 * Pacific Northwest National Laboratory(太平洋西北国家实验室) University of Washington(华盛顿大学) University of Pennsylvania(宾夕法尼亚大学) Colorado State University(科罗拉多州立大学) University of Texas, El Paso(德克萨斯大学埃尔帕索分校) Laboratory for Advanced Cybersecurity Research, National Security Agency(国家安全局高级网络安全研究实验室)

AI总结 本文研究几何统计在检测幻觉中的作用,通过合成数据集分析不同属性对幻觉检测的影响,并提出归一化方法提升多领域检测性能。

Comments Published at the 2025 ICML Workshop on Reliable and Responsible Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19492 2026-02-11 cs.CL

Machine Text Detectors are Membership Inference Attacks

机器文本检测器是成员推断攻击

Ryuto Koike, Liam Dugan, Masahiro Kaneko, Chris Callison-Burch, Naoaki Okazaki

机构 * Institute of Science Tokyo(东京科学研究院) University of Pennsylvania(宾夕法尼亚大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

AI总结 本文研究了成员推断攻击与机器文本检测之间的可转移性,证明了两者在渐近最优性能度量标准上的相同性,并通过实验证明了跨任务性能的强相关性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08892 2026-02-10 stat.ML cs.LG econ.EM

Winner's Curse Drives False Promises in Data-Driven Decisions: A Case Study in Refugee Matching

赢家诅咒驱动数据驱动决策中的虚假承诺:难民匹配案例研究

Hamsa Bastani, Osbert Bastani, Bryce McLaughlin

机构 * Wharton School(沃顿商学院) University of Pennsylvania(宾夕法尼亚大学)

AI总结 本文揭示基于模型的政策评估方法存在赢家诅咒问题,通过难民匹配案例研究证明其在真实效果为零时仍会产生虚假收益。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00989 2026-02-10 stat.ML cs.LG

Optimal Decision-Making Based on Prediction Sets

基于预测集的最优决策制定

Tao Wang, Edgar Dobriban

机构 * University of Pennsylvania(宾夕法尼亚大学)

AI总结 本文提出ROCP算法,通过最小化稳健风险来优化基于预测集的决策制定,适用于医疗诊断等安全关键任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07973 2026-02-10 cs.LG

On Improving Neurosymbolic Learning by Exploiting the Representation Space

通过利用表示空间来改进神经符号学习

Aaditya Naik, Efthymia Tsamoura, Shibo Jin, Mayur Naik, Dan Roth

机构 * University of Pennsylvania(宾夕法尼亚大学) Huawei Labs(华为实验室)

AI总结 CLIPPER通过利用表示空间修剪技术提升神经符号学习性能,显著提高Scallop、Dolphin和ISED等引擎的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18161 2026-02-10 stat.ML cs.LG econ.EM

Beating the Winner's Curse via Inference-Aware Policy Optimization

通过推理感知的策略优化击败赢家诅咒

Hamsa Bastani, Osbert Bastani, Bryce McLaughlin

机构 * Operations, Information, and Decisions Department, The Wharton School(沃顿商学院运营、信息与决策部门) Department of Computer and Information Science, University of Pennsylvania(宾夕法尼亚大学计算机与信息科学系) Wharton Healthcare Analytics Lab, Wharton AI & Analytics Initiative(沃顿健康分析实验室、沃顿人工智能与分析计划)

AI总结 本文提出推理感知的策略优化方法,通过优化策略改进的显著性检验机会,解决赢家诅咒问题,提升策略评估的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06557 2026-02-10 cs.IR cs.LG eess.SP math.MG

Infinity Search: Approximate Vector Search with Projections on q-Metric Spaces

无限搜索:基于q度量空间的向量搜索与投影

Antonio Pariente, Ignacio Hounie, Santiago Segarra, Alejandro Ribeiro

机构 * Department of Electrical and Systems Engineering, University of Pennsylvania, Philadelphia, PA, USA(宾夕法尼亚大学电气与系统工程系) Department of Electrical and Computer Engineering, Rice University, Houston, TX, USA(里士满大学电气与计算机工程系)

AI总结 本文提出基于q度量空间的近似向量搜索方法,通过投影算子将任意不相似函数转换为超度量空间,并利用学习近似以提高搜索效率,实验显示q值增加可提升速度但降低召回率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04836 2026-02-09 cs.AI

Are AI Capabilities Increasing Exponentially? A Competing Hypothesis

人工智能能力是否呈指数增长?一个竞争性假设

Haosen Ge, Hamsa Bastani, Osbert Bastani

机构 * Wharton AI \& Analytics Initiative, The Wharton School, University of Pennsylvania, USA Department of Operations, Information Decisions, The Wharton School, University of Pennsylvania, USA Department of Computer Information Science, University of Pennsylvania, USA

AI总结 本文通过分析数据和提出复杂模型,质疑人工智能能力的指数增长假设,指出拐点已过去并提出未来可能的转折点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03944 2026-02-09 cs.LG

On the Empirical Power of Goodness-of-Fit Tests in Watermark Detection

关于在水印检测中 goodness-of-fit 测试的实证效力

Weiqing He, Xiang Li, Tianqi Shang, Li Shen, Weijie Su, Qi Long

机构 * University of Pennsylvania(宾夕法尼亚大学)

AI总结 本文研究了 goodness-of-fit 测试在水印检测中的实证效力,发现通用 GoF 测试能提升检测能力和鲁棒性,并指出文本重复对 GoF 测试的独特优势。

Comments Accepted at NeurIPS 2025 as a spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏