arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Pennsylvania(宾夕法尼亚大学)

2026-05-13 至 2026-05-13 共收录 7
2605.12421 2026-05-13 cs.AI

Formalize, Don't Optimize: The Heuristic Trap in LLM-Generated Combinatorial Solvers

形式化,而非优化:LLM生成组合求解器中的启发式陷阱

Haoyu Wang, Yuliang Song, Tao Li, Zhiwei Deng, Yaqing Wang, Deepak Ramachandran, Eldan Cohen, Dan Roth

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Toronto(多伦多大学) Google DeepMind(谷歌DeepMind) Oracle AI(Oracle人工智能)

AI总结 本文研究LLM生成组合求解器时形式化与优化的矛盾,通过CP-SynC-XL基准测试,发现Python+OR-Tools在正确性上最优,而MiniZinc+OR-Tools虽使用相同后端但覆盖度较低,启发式优化导致部分问题速度下降和正确性降低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11928 2026-05-13 cs.AI

When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents

当模拟欺骗时:一个仿真到现实的基准和领域随机化的强化学习配方用于工具使用智能体

Xiaolin Zhou, Aojie Yuan, Zheng Luo, Zipeng Ling, Xixiao Pan, Yicheng Gao, Haiyue Zhang, Jiate Li, Shuli Jiang, Prince Zizhuang Wang, Zixuan Zhu, Jinbo Liu, Ryan A. Rossi, Hua Wei, Xiyang Hu

机构 * Arizona State University(亚利桑那州立大学) University of Southern California(南加州大学) Carnegie Mellon University(卡内基梅隆大学) University of Pennsylvania(宾夕法尼亚大学) Adobe Research(Adobe研究)

AI总结 本文研究了工具使用POMDP中的仿真到现实差距,提出了一种领域随机化的强化学习方法,通过扰动增强轨迹提升鲁棒性,缩小了工具使用智能体在现实部署中的性能差距。

Comments Dataset, code, and benchmark leaderboard are available at https://github.com/WillChow66/robustbench-tc-release.git and https://huggingface.co/spaces/willchow66/robustbench-tc-leaderboard

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08813 2026-05-13 cs.LG

Robust Policy Optimization to Prevent Catastrophic Forgetting

鲁棒策略优化以防止灾难性遗忘

Mahdi Sabbaghi, George Pappas, Adel Javanmard, Hamed Hassani

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Southern California(南加州大学)

AI总结 本文提出FRPO框架,通过优化当前策略及下游适应可达的KL限制邻域内策略,提升鲁棒性,减少安全性能退化,同时保持下游任务表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11501 2026-05-13 cs.SE cs.AI cs.CR

Decaf: Improving Neural Decompilation with Automatic Feedback and Search

Decaf:通过自动反馈和搜索改进神经反编译

Alexander Shypula, Osbert Bastani, Edward Schwartz

机构 * University of Pennsylvania(宾夕法尼亚大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 Decaf通过自动反馈和搜索显著提升神经反编译的语义正确性,将反编译率从26.0%提升至83.9%。

Comments 15 pages, 6 figures. Preprint; under review. Code and models available at https://github.com/AlexShypula/decaf

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11265 2026-05-13 cs.CV cs.AI cs.LG

DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction

DenseTRF: 基于纹理的无监督表示适应用于外科场景密集预测

Guiqiu Liao, Matjaž Jogan, Daniel A. Hashimoto

机构 * GRASP Laboratory, University of Pennsylvania(宾夕法尼亚大学GRASP实验室) PCASO Laboratory, Department of Surgery, University of Pennsylvania(宾夕法尼亚大学外科PCASO实验室) Department of Computer and Information Science, University of Pennsylvania(宾夕法尼亚大学计算机与信息科学系)

AI总结 本文提出DenseTRF框架,通过基于纹理的关注机制学习纹理感知的表示,提升外科场景密集预测的跨分布泛化能力。

Comments Accepted to 29th International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11161 2026-05-13 cs.LG cs.AI

Interpretability Can Be Actionable

可解释性可以是可操作的

Hadas Orgad, Fazl Barez, Tal Haklay, Isabelle Lee, Marius Mosbach, Anja Reusch, Naomi Saphra, Byron Wallace, Sarah Wiegreffe, Eric Wong, Ian Tenney, Mor Geva

机构 * Kempner Institute at Harvard University(哈佛大学凯默纳研究所) University of Southern California(美国南加州大学) Mila – Quebec AI Institute(魁北克AI研究所) McGill University(麦吉尔大学) Google DeepMind(谷歌DeepMind) Tel Aviv University(特拉维夫大学) University of Pennsylvania(宾夕法尼亚大学) University of Maryland(马里兰大学) University of Oxford(牛津大学) Northeastern University(东北大学) Boston University(波士顿大学)

AI总结 本文探讨了可解释性研究的核心问题,提出应以可操作性作为评价标准,通过具体性和验证性两个维度分析阻碍实际应用的障碍,并提出五个领域和评估框架。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06870 2026-05-13 cs.LG

Continuous First, Discrete Later: VQ-VAEs Without Dimensional Collapse

先连续后离散:无需维度坍缩的VQ-VAE

Xinyu Zhao, Nikita Karagodin, Hamed Hassani, Sinan Hersek, Paul Pu Liang, Yury Polyanskiy

机构 * MIT(麻省理工学院) University of Pennsylvania(宾夕法尼亚大学) Google(谷歌)

AI总结 本文探讨了VQ-VAE中维度坍缩问题,提出通过预训练无量化自编码器缓解该问题,提升表示维度和重建质量。

详情

展开后加载摘要…

URL PDF HTML 收藏