机构
*
Department of Computer Science and Engineering(计算机科学与工程系)
;
Department of Electrical and Electronic Engineering(电气与电子工程系)
;
Islamic University of Technology(伊斯兰技术大学)
An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods
LLM生成代码安全性的提示方法实证评估
Mohammed Kharma, Ahmed Sabbah, Mohammad Alkhanafseh, Mohammad Hammoudeh, David Mohaisen
机构
*
Department of Computer Science, Birzeit University(计算机科学系,巴勒斯坦比泽大学)
;
King Fahd University of Petroleum and Minerals(国王法赫德石油和矿物大学)
;
University of Central Florida(中央佛罗里达大学)
TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation
TRIP-Evaluate: 一个用于评估交通领域大模型的开放多模态基准
Han Gong, Zhen Zhou, Yunyang Shi, Yan Tan, Jinbiao Huo, Qi Hong, Zhiyuan Liu
机构
*
School of Transportation(交通学院)
;
Southeast University(东南大学)
;
School of Artificial Intelligence and Computer Science(人工智能与计算机科学学院)
;
Jiangnan University(江南大学)
;
Department of Civil and Environmental Engineering(土木与环境工程系)
;
Hong Kong Polytechnic University(香港理工大学)
Making AI-Assisted Grant Evaluation Auditable without Exposing the Model
在不暴露模型的情况下使AI辅助的资助评估可审计
Kemal Bicakci
机构
*
Informatics Institute, Istanbul Technical University, Istanbul, Türkiye(伊斯坦布尔技术大学信息学院,伊斯坦布尔,土耳其)
;
Securify Information Technology and Security Training Consulting Inc., Ankara, Türkiye(Securify信息科技与安全培训咨询公司,安卡拉,土耳其)
Navigating Large-Scale Document Collections: MuDABench for Multi-Document Analytical QA
在大规模文档集合中导航:MuDABench用于多文档分析问答
Zhanli Li, Yixuan Cao, Lvzhou Luo, Ping Luo
机构
*
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences (CAS)(人工智能安全国家重点实验室,计算技术研究所,中国科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Wenlan School of Business, Zhongnan University of Economics and Law(中南财经政法大学文澜商学院)
CommentsFindings of ACL 2026. The camera-ready version corrects some labeling errors. The accompanying repository is continuously updated based on community feedback; for the most up-to-date implementation and results, please refer to the repository
RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents
RPA-Check:一种多阶段自动框架,用于评估基于大语言模型的角色扮演代理
Riccardo Rosati, Edoardo Colucci, Massimiliano Bolognini, Adriano Mancini, Paolo Sernani
机构
*
Department of Political Sciences, Communication and International Relations, University of Macerata(马切拉塔大学政治学、传播与国际关系系)
;
Department of Law, University of Macerata(马切拉塔大学法律系)
EVGeoQA: Benchmarking LLMs on Dynamic, Multi-Objective Geo-Spatial Exploration
EVGeoQA:基于动态多目标地理空间探索的LLM基准测试
Jianfei Wu, Zhichun Wang, Zhensheng Wang, Zhiyu He
机构
*
School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)
;
Beijing Key Laboratory of Artificial Intelligence for Education(北京市教育人工智能重点实验室)
;
Engineering Research Center of Intelligent Technology and Educational Application, Ministry of Education(教育部智能技术与教育应用工程研究中心)
;
College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机学院)
Pan Chen, Shaohong Chen, Mark Wang, Shi Xuan Leong, Priscilla Fung, Varinia Bernales, Alan Aspuru-Guzik
机构
*
University of Toronto(多伦多大学)
;
Nanyang Technological University(南洋理工大学)
;
Acceleration Consortium(加速联盟)
;
Vector Institute for Artificial Intelligence(向量人工智能研究所)
;
Canadian Institute for Advanced Research (CIFAR)(加拿大高等研究院(CIFAR))
;
NVIDIA(英伟达)
机构
*
Lehigh University(里海大学)
;
Squirrel Ai Learning(松鼠AI学习)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Michigan State University(密歇根州立大学)
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
CUA-Suite:大规模人工标注的视频演示用于计算机使用代理
Xiangru Jian, Shravan Nayak, Kevin Qinghong Lin, Aarash Feizi, Kaixin Li, Patrice Bechard, Spandana Gella, Sai Rajeswar
机构
*
ServiceNow
;
University of Waterloo(多伦多大学)
;
Mila
;
Université de Montréal(蒙特利尔大学)
;
McGill University(麦吉尔大学)
;
University of Oxford(牛津大学)
;
National University of Singapore(新加坡国立大学)
The PokeAgent Challenge: Competitive and Long-Context Learning at Scale
PokeAgent挑战:在大规模中实现竞争性和长上下文学习
Seth Karten, Jake Grigsby, Tersoo Upaa, Junik Bae, Seonghun Hong, Hyunyoung Jeong, Jaeyoon Jung, Kun Kerdthaisong, Gyungbo Kim, Hyeokgi Kim, Yujin Kim, Eunju Kwon, Dongyu Liu, Patrick Mariglia, Sangyeon Park, Benedikt Schink, Xianwei Shi, Anthony Sistilli, Joseph Twin, Arian Urdu, Matin Urdu, Qiao Wang, Ling Wu, Wenli Zhang, Kunsheng Zhou, Stephanie Milani, Kiran Vodrahalli, Amy Zhang, Fei Fang, Yuke Zhu, Chi Jin
机构
*
Princeton(普林斯顿大学)
;
UT-Austin(得克萨斯大学奥斯汀分校)
;
CMU(卡内基梅隆大学)
;
NYU(纽约大学)
;
Google DeepMind(谷歌DeepMind)
;
Team Heatz(团队Heatz)
;
Team PA-Agent(团队PA-Agent)
;
Team FoulPlay(团队FoulPlay)
;
Team 4thLesson(团队4thLesson)
;
Team Q(团队Q)
;
Team Anthonys(团队Anthonys)
;
Team Hamburg(团队Hamburg)
;
Team Porygon2AI(团队Porygon2AI)
;
Team Deepest(团队Deepest)
;
Team August(团队August)
机构
*
University of North Dakota(北达科他大学)
;
Youngstown State University(青年州大学)
;
University of Toledo(托莱多大学)
;
University of Missouri(密苏里大学)
;
Tribhuvan University(特里布文大学)
Jonathan D. Chang, Andrew Drozdov, Shubham Toshniwal, Owen Oertell, Alexander Trott, Jacob Portes, Abhay Gupta, Pallavi Koppol, Ashutosh Baheti, Sean Kulinski, Ivan Zhou, Irene Dea, Krista Opsahl-Ong, Simon Favreau-Lessard, Sean Owen, Jose Javier Gonzalez Ortiz, Arnav Singhvi, Xabi Andrade, Cindy Wang, Kartik Sreenivasan, Sam Havens, Jialu Liu, Peyton DeNiro, Wen Sun, Michael Bendersky, Jonathan Frankle
Evaluating GPT-5 as a Multimodal Clinical Reasoner: A Landscape Commentary
评估GPT-5作为多模态临床推理者的有效性:领域评论
Alexandru Florea, Shansong Wang, Mingzhe Hu, Qiang Li, Zach Eidex, Luke del Balzo, Mojtaba Safari, Xiaofeng Yang
机构
*
Department of Radiation Oncology, Winship Cancer Institute, Emory University School of Medicine(放射肿瘤科,Winship癌症研究所,埃默里大学医学院)
;
Department of Biomedical Engineering, Georgia Institute of Technology(生物医学工程系,佐治亚理工学院)
Tucano 2 Cool: Better Open Source LLMs for Portuguese
Tucano 2 Cool: 更好的开源葡萄牙语大语言模型
Nicholas Kluge Corrêa, Aniket Sen, Shiza Fatimah, Sophia Falk, Lennard Landgraf, Julia Kastner, Lucie Flek
机构
*
Bonn-Aachen International Center for Information Technology (b-it) / CAISA Lab(波恩-亚琛国际信息科技中心(b-it)/ CAISA实验室)
;
Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习与人工智能研究所)
;
Center for Science and Thought(科学与思想研究中心)
;
Helmholtz-Institut für Strahlen- und Kernphysik(亥姆霍兹辐射与核物理研究所)
;
Bonn Sustainable AI Lab(波恩可持续AI实验室)