机构
*
School of Data Science, Fudan University(复旦大学数据科学学院)
;
Shanghai Innovation Institute(上海创新研究院)
;
Shanghai University of Finance and Economics(上海财经大学)
;
MOE laboratory for National Development and Intelligent Governance, Fudan University(复旦大学国家发展与智能治理实验室)
;
Research Institute of Intelligent Complex Systems, Fudan University(复旦大学智能复杂系统研究院)
专题命中
评测与基准
:LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL
BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science
BioProBench: 一个全面的数据集和生物协议理解与推理基准
Yuyang Liu, Liuzhenghao Lv, Xiancheng Zhang, Jingya Wang Li Yuan, Yonghong Tian
机构
*
School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院)
;
School of Chemical Biology&Biotechnology, Peking University(北京大学化学生物学与生物技术学院)
Comments17 pages, 6 figures, 4 tables. Accepted at AIES 2026 (AAAI/ACM Conference on AI, Ethics, and Society). This version includes the supplementary appendix. Code and data: https://github.com/mbrcic/llm-political-steerability (Zenodo DOI 10.5281/zenodo.21489805)
AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models
AIR-BENCH Live:基础模型不断演进的安全基准测试
Rohan Naphade, Minzhou Pan, Bo Li
机构
*
Virtue AI(美德人工智能公司)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Northeastern University(东北大学)
;
University of Chicago(芝加哥大学)
;
University of Illinois, Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构
*
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件学院)
;
Harbin Institute of Technology Shenzhen(哈尔滨工业大学(深圳))
;
City University of Macau(澳门城市大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback
CodeEvo:通过混合和迭代反馈以交互驱动方式合成以代码为中心的数据
Qiushi Sun, Jinyang Gong, Lei Li, Qipeng Guo, Fei Yuan
机构
*
Shanghai AI Laboratory(上海人工智能实验室)
;
The University of Hong Kong(香港大学)
;
New York University(纽约大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Shanghai Innovation Institute(上海创新研究院)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
E-Bench:在现实世界产品场景中对多步工具使用智能体进行基准测试
Weihuang Zheng, Tianyuan Zou, Eileen Ye, Alphet Liu, Youyong Kong, Ya-Qin Zhang, Duran Zheng, Maxm Pan
机构
*
Hunyuan Team, Tencent(腾讯混元团队)
;
Institute for AI Industry Research, Tsinghua University(清华大学人工智能产业研究院)
;
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI
TLA+-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA Specification Generation
TLA$^{+}$-Bench:用于自然语言到TLA+规范生成的基于执行的基准测试和数据集
Arslan Bisharat, Eric Spencer, Brian Ortiz, Khushboo Bhadauria, Mujtaba Nazari, Beatriz Santos, Anisa Ramos, TaiNing Wang, George K. Thiruvathukal, Konstantin Läufer, Mohammed Abuhamad
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI
Comments17 pages, appendix included. Introduces TLA+-Bench, an execution-grounded benchmark and dataset for natural-language to TLA$^{+}$ specification generation. Dataset, evaluation code, and model outputs available at publication
Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
谁买单?面向真实世界网络代理的以利益相关者为中心的提示注入基准测试
Zihao Wang, Yiming Li, Yutong Wu, Kangjie Chen, Zheyu Liu, Fok Kar Wai, Pin-Yu Chen, Vrizlynn L. L. Thing, Bo Li, Dacheng Tao, Tianwei Zhang
机构
*
Nanyang Technological University, Singapore(南洋理工大学,新加坡)
;
ST Engineering, Singapore(ST工程,新加坡)
;
IBM Research, USA(IBM研究院,美国)
;
University of Illinois Urbana-Champaign, USA(伊利诺伊大学厄巴纳-香槟分校,美国)