机构
*
Beihang University(北京航空航天大学)
;
Centre for Artificial Intelligence and Robotics, HKISI-CAS(香港智能科学与工业研究院人工智能与机器人中心)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院)
;
Beijing Jiaotong University(北京交通大学)
CommentsPublished in Proceedings of the 4th NeurIPS Efficient Natural Language and Speech Processing Workshop (ENLSP-IV), Vancouver, Canada, 2024. 14 pages, 8 figures
CommentsVersion 2 incorporates revisions based on feedback from NeurIPS 2025 reviewers (final score: borderline). We improved clarity in previously complex sections to enhance accessibility for non-expert readers and expanded the experimental evaluation to provide more comprehensive and diverse results
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
利用LLM作为法官/陪审团推进模型响应用户表现精神病的可扩展、临床验证的安全性评估
May Lynn Reese, Markela Zeneli, Mindy Ng, Jacob Haimes, Andreea Damien, Elizabeth Stade
机构
*
Apart Research
;
Odyssean Institute
;
London School of Economics and Political Science(伦敦政治经济学院)
;
Stanford Institute for Human-Centered AI(斯坦福大学以人为本人工智能研究所)
SNEAKDOOR: Stealthy Backdoor Attacks against Distribution Matching-based Dataset Condensation
SNEAKDOOR:针对基于分布匹配的数据集压缩的隐秘后门攻击
He Yang, Dongyi Lv, Song Ma, Wei Xi, Jizhong Zhao
机构
*
School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)
;
National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Xi’an Jiaotong University(西安交通大学人机混合增强智能全国重点实验室)
Comments23 pages (17 pages of main text). See https://github.com/WMD-group/xtalmet for the code. Significantly extended from the early version of this work, which was accepted to the AI4Mat workshop at NeurIPS 2025
Comments6 pages, 2 figures and 2 tables. Version presented at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Machine Learning and the Physical Sciences. 6 December, 2025; San Diego, California, USA
CommentsSubmitted to GenAI4Health@NeurIPS 2025. This was the first version of the LLM-assisted emergency triage benchmark dataset and baseline models. A related but separate benchmark-focused study on emergency triage under constrained sensing has been accepted at the IEEE International Conference on Healthcare Informatics (ICHI) 2026 (see arXiv:2602.20168)
Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks
硅 bureaucracy 与 AI 考试导向教育:LLM 测试基准中的污染敏感性与分数可信度
Yiliang Song, Hongjun An, Jiangan Chen, Xuanchen Yan, Huan Song, Jiawei Shao, Xuelong Li
机构
*
Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究院(TeleAI))
;
Guangxi Normal University(广西师范大学)
;
Northwestern Polytechnical University(西北工业大学)
机构
*
University of Chinese Academy of Sciences(中国科学院大学)
;
University of North Texas(北德克萨斯大学)
;
Institute of Software Chinese Academy of Sciences(中国科学院软件研究所)
;
Rochester Institute of Technology(罗切斯特理工学院)
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
ShotBench:视觉语言模型中的专家级电影叙事理解
Hongbo Liu, Jingwen He, Yi Jin, Dian Zheng, Yuhao Dong, Fan Zhang, Ziqi Huang, Yinan He, Yangguang Li, Weichao Chen, Yu Qiao, Wanli Ouyang, Shengjie Zhao, Ziwei Liu
机构
*
Tongji University(同济大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
S-Lab, Nanyang Technological University(南洋理工大学S-Lab)
CommentsAccepted by NeurIPS 2024. BICCOS is part of the alpha-beta-CROWN verifier, the VNN-COMP 2024 winner; fixed Theorem 3.2 and clarified experimental results
BeetleFlow: An Integrative Deep Learning Pipeline for Beetle Image Processing
BeetleFlow: 一种用于甲虫图像处理的集成深度学习流水线
Fangxun Liu, S M Rayeed, Samuel Stevens, Alyson East, Cheng Hsuan Chiang, Colin Lee, Daniel Yi, Junke Yang, Tejas Naik, Ziyi Wang, Connor Kilrain, Elijah H Buckwalter, Jiacheng Hou, Saul Ibaven Bueno, Shuheng Wang, Xinyue Ma, Yifan Liu, Zhiyuan Tao, Ziheng Zhang, Eric Sokol, Michael Belitz, Sydne Record, Charles V. Stewart, Wei-Lun Chao
机构
*
The Ohio State University(俄亥俄州立大学)
;
Rensselaer Polytechnic Institute(伦斯勒理工学院)
;
The University of Maine(缅因大学)
;
National Ecological Observatory Network (NEON), Battelle(国家生态观测网络(NEON),巴特尔纪念研究所)
;
Michigan State University(密歇根州立大学)
PepThink-R1: LLM for Interpretable Cyclic Peptide Optimization with CoT SFT and Reinforcement Learning
PepThink-R1:基于CoT SFT和强化学习的可解释环状肽优化LLM
Ruheng Wang, Hang Zhang, Trieu Nguyen, Shasha Feng, Hao-Wei Pang, Xiang Yu, Li Xiao, Peter Zhiping Zhang
机构
*
Merck & Co., Inc., Rahway, NJ, USA(默克公司,美国新泽西州拉威)
;
UT Southwestern Medical Center, Dallas, TX, USA(得克萨斯大学西南医学中心,美国得克萨斯州达拉斯)
;
University of Pittsburgh, Pittsburgh, PA, USA(匹兹堡大学,美国宾夕法尼亚州匹兹堡)