Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
跨商业学科的前沿人工智能性能:基于案例的知识工作和分析推理基准
Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, Karim Lakhani
机构
*
The Wharton School, University of Pennsylvania(宾夕法尼亚大学沃顿商学院)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Harvard Business School, Harvard University(哈佛大学哈佛商学院)
Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences
顺序很重要:LVLMs作为图像序列时间推理的评判者
Martina Ianaro, Guilherme Fernandes, Maurizio Gabbrielli, Joao Magalhaes
机构
*
University of Bologna(博洛尼亚大学)
;
NOVA School of Science and Technology(NOVA科技学院)
;
NOVA Laboratory for Computer Science and Informatics(NOVA计算机科学与信息实验室)
OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents
OmnilingualGAIA2:评估前沿AI智能体的多语言差距
Andrea Caciolai, Pere-Lluís Huguet Cabot, Chierh Cheng, Albert Ventayol-Boada, Gabriel Mejia Gonzalez, Christophe Ropers, Lucas Bandarkar, Sebastian Ruder, Darlene Sakakihara, Elliot Yun, Pierre Andrews, Grégoire Mialon, Romain Froger, Marta R. Costa-jussà
Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human Understanding
解释正确性是否重要?将计算XAI评估与人类理解联系起来
Gregor Baer, Chao Zhang, Isel Grau, Pieter Van Gorp
机构
*
Information Systems Group, Eindhoven University of Technology(埃因霍温理工大学信息系统组)
;
Human-Technology Interaction Group, Eindhoven University of Technology(埃因霍温理工大学人机交互组)
机构
*
National Engineering Research Center for Software Engineering, Peking University(北京大学软件工程国家工程研究中心)
;
Institute of Artificial Intelligence, China Telecom (TeleAI)(中国电信人工智能研究院)
;
Tsinghua University(清华大学)
;
Chinese Academy of Sciences(中国科学院)
;
University of British Columbia(不列颠哥伦比亚大学)
;
Renmin University of China(中国人民大学)