Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse
迈向通用游戏玩家:对游戏多元宇宙中基础模型的调查
Kuan Zhang, Dongchen Liu, Qiyue Zhao, Tianyu Xin, Yue Su, Haisheng Wang, Han Yin, Hongbo Ma, Peize Li, Tianjun Gu, Xiangnan Wu, Xinran Zhang, Yongxuan Li, Zirong Chen, Yiming Li
机构
*
College of AI, Tsinghua University(清华大学人工智能学院)
;
MMLab, The University of Hong Kong(香港大学MMLab)
;
University of Chinese Academy of Sciences(中国科学院大学)
Benchmarking Vision Foundation Models for Input Monitoring in Autonomous Driving
在自动驾驶中基于视觉基础模型的输入监控基准测试
Mert Keser, Halil Ibrahim Orhan, Niki Amini-Naieni, Gesina Schwalbe, Alois Knoll, Matthias Rottmann
机构
*
Continental AG(大陆汽车集团)
;
Technical University of Munich(慕尼黑技术大学)
;
University of Lübeck(吕贝克大学)
;
University of Oxford(牛津大学)
;
University of Wuppertal(伍珀塔尔大学)
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
一念之差:多轮对话中针对隐藏恶意意图的响应感知防御
Xinjie Shen, Rongzhe Wei, Peizhi Niu, Haoyu Wang, Ruihan Wu, Eli Chien, Bo Li, Pin-Yu Chen, Pan Li
机构
*
Georgia Institute of Technology(佐治亚理工学院)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
UCSD(加州大学圣地亚哥分校)
;
National Taiwan University(台湾大学)
;
IBM Research(IBM研究院)
;
Virtue AI
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild
RW-Post:可审计的证据导向多模态事实核查
Danni Xu, Shaojing Fan, Harry Cheng, Mohan Kankanhalli
机构
*
School of Computing (SoC), National University of Singapore (NUS)(新加坡国立大学计算机学院(SoC))
;
National University of Singapore (NUS)(新加坡国立大学)
;
Department of Electrical and Computer Engineering (ECE), National University of Singapore (NUS)(新加坡国立大学电子与计算机工程系(ECE))
Overview of the MedHopQA track at BioCreative IX: track description, participation and evaluation of systems for multi-hop medical question answering
BioCreative IX MedHopQA 轨道概述:轨道描述、参与及多跳医学问答系统评估
Rezarta Islamaj, Joey Chan, Robert Leaman, Jongmyung Jung, Hyeongsoon Hwang, Quoc-An Nguyen, Hoang-Quynh Le, Harikrishnan Gurushankar Saisudha, Ganesh Chandrasekar, Rustam R. Taktashov, Nadezhda Yu. Bizyukova, Sofia I. R. Conceição, Paulo R. C. Lopes, Reem Abdel Salam, Mary Adewunmi, Zhiyong Lu
机构
*
National Library of Medicine (NLM), National Institutes of Health (NIH)(美国国家医学图书馆(NLM)、国家卫生研究院(NIH))
;
University of Illinois at Urbana Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Korea University(韩国大学)
;
VNU University of Engineering and Technology, Hanoi, Vietnam(越南河内工程大学)
;
Concordia University, Montreal, QC, CA(蒙特利尔大学)
;
Institute of Biomedical Chemistry (IBMC), 10 bld. 8, Pogodinskaya str., 119121 Moscow, Russia(俄罗斯生物医学化学研究所(IBMC))
;
LASIGE, Departamento de Informática, Faculdade de Ciências, Universidade de Lisboa, 1749-016 Lisbon, Portugal(葡萄牙里斯本大学 LASIGE 实验室)
;
Faculty of Engineering, Computer Engineering Department Cairo University(埃及开罗大学工程学院)
;
Menzies School of Health Research, Charles Darwin University, NT, Australia(澳大利亚查尔斯达尔文大学梅恩兹健康研究中心)
;
CaresAI, Australia(澳大利亚 CaresAI)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL
AI总结
本文介绍了BioCreative IX MedHopQA共享任务,旨在评估大型语言模型在多跳推理中的表现,通过构建1000个挑战性问题,展示了检索增强生成策略的重要性,并提供了公开数据集和评估结果。