机构
*
organization= Department of Geography, Texas A\&M University , city= College Station , country= USA
;
organization= Department of Landscape Architecture \& Urban Planning, Texas A\&M University , city= College Station , country= USA
;
organization= Department of Industrial
;
Systems Engineering, University of Florida , city= Gainesville , country= USA
;
organization= Spatial Sciences Institute, University of Southern California , city= Los Angeles , country= USA
;
organization= Department of Geography
;
Sustainability, University of Tennessee , city= Knoxville , country= USA
专题命中
评测与基准
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)
机构
*
Swiss Institute of Allergy and Asthma Research, Davos, Switzerland(瑞士过敏与哮喘研究所,达沃斯,瑞士)
;
ETH Zurich, Zurich, Switzerland(苏黎世联邦理工学院,苏黎世,瑞士)
;
Swiss Institute of Bioinformatics, Lausanne, Switzerland(瑞士生物信息学研究所,洛桑,瑞士)
One Interaction Is Worth a Thousand Guesses: Benchmarking the Interactive Capabilities of Deep Research Agents
一次交互胜过千次猜测:深度研究代理的交互能力基准测试
Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, Anthony K. H. Tung
机构
*
School of Computing, National University of Singapore(新加坡国立大学计算机学院)
;
School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)智能科学与工程学院)
;
Zhejiang University(浙江大学)
;
State Key Lab of CAD&CG, Zhejiang University(浙江大学计算机辅助设计与图形学国家重点实验室)
专题命中
评测与基准
:LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Comments28 pages, 2 figures, 13 tables. Benchmark, environment spec, and app contract released. First open-weight three-model sweep (k=5) on a 40-task oracle-validated executable suite; frontier-model leaderboard committed in the roadmap
机构
*
Knowlesdge Lab, University of Chicago, Chicago, Illinois, USA(知识实验室,芝加哥大学,芝加哥,伊利诺伊州,美国)
;
Knowledge Lab, Data Science Institute, Department of Sociology, University of Chicago, Chicago, Illinois, USA(知识实验室,数据科学学院,社会学系,芝加哥大学,芝加哥,伊利诺伊州,美国)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
Is Our Benchmark Enough? An Analysis of Continual Learning for MLLMs
我们的基准测试足够吗?多模态大语言模型持续学习分析
Van-Tuan Tran, Shruthi Gowda, Merim Dzaferagic, Marco Ruffini
机构
*
School of Computer Science and Statistics, Trinity College Dublin(都柏林圣三一学院计算机科学与统计学院)
;
Department of Mathematics and Computer Science, Eindhoven University of Technology(埃因霍温理工大学数学与计算机科学系)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation
DoGMaTiQ:面向报告评估的问答片段自动生成
Bryan Li, William Walden, Yu Hou, Gabrielle Kaili-May Liu, Dawn Lawrie, James Mayfield, Eugene Yang, Chris Callison-Burch, Laura Dietz
机构
*
Google Inc.(谷歌公司)
;
Johns Hopkins University(约翰霍普金斯大学)
;
University of Maryland(马里兰大学)
;
Yale University(耶鲁大学)
;
University of Pennsylvania(宾夕法尼亚大学)
;
University of New Hampshire(新罕布什尔大学)