arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Pennsylvania(宾夕法尼亚大学)

2026-05-04 至 2026-05-04 共收录 3
2605.00300 2026-05-04 cs.AI cs.DC cs.LG cs.PF

Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference

Token Arena:一个统一能量与认知的连续基准测试

Yuxuan Gao, Megan Wang, Yi Ling Yu

机构 * University of Pennsylvania(宾夕法尼亚大学) Columbia University(哥伦比亚大学) OpenMesh AI

AI总结 Token Arena通过端到端的五个核心维度评估AI推理性能,结合能耗和价格指标,揭示不同端点对模型准确性和效率的影响,提供更真实的推理评估标准。

Comments 14 pages, 1 figure, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07630 2026-05-04 cs.CL cs.AI cs.CV

InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information

InterChart:跨分解与分布式图表信息的基准测试

Anirudh Iyengar Kaniyar Narayana Iyengar, Srija Mukhopadhyay, Adnan Qidwai, Shubhankar Singh, Dan Roth, Vivek Gupta

机构 * Arizona State University(亚利桑那州立大学) IIIT, Hyderabad(海得拉巴印度理工学院) Mercer Mettl University(梅森-梅特尔大学) University of Pennsylvania(宾夕法尼亚大学)

AI总结 InterChart通过评估视觉语言模型在多图表间推理的能力,揭示了模型在复杂图表集成中的局限性,推动多模态推理的发展。

Comments 22 pages, 8 figures, 14 tables. Accepted at IJCNLP-AACL 2025

Journal ref Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics 2025, 2046-2067

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10990 2026-05-04 cs.GT cs.LG econ.TH math.ST stat.ML stat.TH

Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium

大语言模型与人类偏好的统计不可能性与可能性:从康德斯悖论到纳什均衡

Kaizhao Liu, Qi Long, Zhekun Shi, Weijie J. Su, Jiancong Xiao

机构 * Massachusetts Institute of Technology(麻省理工学院) University of Pennsylvania(宾夕法尼亚大学) Princeton University(普林斯顿大学)

AI总结 本文探讨了对齐大语言模型与人类偏好的统计限制,证明在一般概率偏好模型下,康德斯循环几乎必然存在,从而表明基于奖励的方法无法完全对齐偏好。同时,研究了非奖励方法下LLM采用混合策略的条件,证明在Luce模型下,少数群体偏好得以保留的统计可能性。

Comments Accepted for publication in the Annals of Statistics

详情

展开后加载摘要…

URL PDF HTML 收藏