Sha Sajadieh, Loredana Fattorini, Raymond Perrault, Yolanda Gil, Vanessa Parli, Lapo Santarlasci, Juan Pava, Nestor Maslej, Russ Altman, Erik Brynjolfsson, Carla Brodley, Jack Clark, Virginia Dignum, Vipin Kumar, James Landay, Terah Lyons, James Manyika, Juan Carlos Niebles, Yoav Shoham, Elham Tabassi, Russell Wald, Toby Walsh, Dan Weld
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
ReasonCLIP-58M: CLIP的视觉基础常识推理监督
Sicheng Zhang, Muzammal Naseer, Binzhu Xie, Naufal Suryanto, Shi Qiu, Jamal Bentahar, Naveed Akhtar, Mubarak Shah
机构
*
Khalifa University(卡利法大学)
;
University of Western Australia(西澳大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
University of Melbourne(墨尔本大学)
;
University of Central Florida(佛罗里达中央大学)
Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models
跨模态对抗扩散:文本、视觉与视觉-语言模型的攻击、防御与评估融合综述
Abrar Alotaibi, Moataz Ahmed
机构
*
Information and Computer Science Department, King Fahd University of Petroleum & Minerals(国王法赫德石油矿物大学信息与计算机科学系)
;
College of Computer Science and Information Technology, Imam Abdulrahman Bin Faisal University(伊玛目阿卜杜勒拉赫曼·本·法伊塞尔大学计算机科学与信息科技学院)
;
SDAIA-KFUPM Joint Research Center for Artificial Intelligence, King Fahd University of Petroleum & Minerals(SDAIA-KFUPM人工智能联合研究中心,国王法赫德石油矿物大学)
STEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity
STEB:用于评估超越翻译保真度的语音到语音翻译表现力基准
Sitong Cheng, Weizhen Bian, Songjun Cao, Jin Li, Bei Liu, Chunyang Jiang, Yike Zhang, Weihao Wu, Yiming Li, Chi-Min Chan, Long Ma, Wei Xue
机构
*
Hong Kong University of Science and Technology(香港科技大学)
;
Tencent Youtu Lab(腾讯优图实验室)
;
Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions
长期模拟揭示AI伴侣的认知发展风险
Kaicheng Shen, Lingyu Li, Wen Wu, Yan Teng, Liang He, Yingchun Wang
机构
*
Shanghai Institute of Artificial Intelligence for Education, East China Normal University(华东师范大学上海人工智能教育研究院)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院)
Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
评估稀疏自编码器与概念标注的可解释性
Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos, Begüm Demir
机构
*
The Berlin Institute for the Foundations of Learning and Data (BIFOLD)(柏林学习与数据基础研究所)
;
Technische Universität Berlin(柏林工业大学)
;
INRAE(法国国家农业、食品与环境研究院)
;
Inria, EVERGREEN(法国国家信息与自动化研究所,EVERGREEN)
;
UMR TETIS, Univ Montpellier(UMR TETIS,蒙彼利埃大学)
EEG Benchmarking Needs a Task Specification Layer: NeuroDoc for Rulebook-Guided, Executable Benchmark Construction
EEG基准测试需要一个任务规范层:NeuroDoc用于基于规则手册的可执行基准构建
Chengxuan Qin, Zhige Chen, Shu Peng, Rui Yang, Jiping Cui, Yikai Dong, Jun Li, Liu Peng, Zhida Shang, Mingze Tang, Kay Chen Tan, Jibin Wu
机构
*
School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西交利物浦大学先进技术学院)
;
School of Electrical Engineering, Electronics and Computer Science, University of Liverpool(利物浦大学电气工程、电子与计算机科学学院)
;
School of Computer Science and Informatics, University of Liverpool(利物浦大学计算机科学与信息学学院)
;
Department of Data Science and Artificial Intelligence, Hong Kong Polytechnic University(香港理工大学数据科学与人工智能系)
SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment
SkillAudit:从固定套件基准测试到以技能为中心的评估
Dexu Yu, Youhua Li, Zhaoyang Guan, Xianhao Lin, Jining Luan, Zihao Rao, Xuanqi Lan, Yang Ran, Bo Lan, Nai-Xin Zhai, Hanwen Du, Junchen Fu, Wenhao Deng, Yongxin Ni, Chunxiao Li
机构
*
Northeastern University(东北大学)
;
City University of Hong Kong(香港城市大学)
;
Northwestern University(西北大学)
;
Fudan University(复旦大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Santa Clara University(圣克拉拉大学)
;
Fenz AI
;
Ohio State University(俄亥俄州立大学)
;
University of Glasgow(格拉斯哥大学)
;
National University of Singapore(新加坡国立大学)
;
DeciLix Lab(DeciLix实验室)
AutoRAS: Learning Robust Agentic Systems with Primitive Representations
AutoRAS: 学习具有原始表示的鲁棒智能系统
Yang Yue, Xuancheng Zhu, Yuyang Ma, Guoshun Nan, Zihan Dou, Jingru Shan, Congyu Guo, Ji Zhang, Hua Wang, Jingfeng Zhang
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Guangxi Transportation Science and Technology Group Co., Ltd.(广西交通科技集团有限公司)
;
Fudan University(复旦大学)
A Validation-Gated Mechanistic Account of Suicidality Detection in LLMs
一种验证门控机制的自杀检测在LLMs中的解释
Nafiz Ahmed, Sarah Sharif, Dingjing Shi, Mike Banad
机构
*
Intelligent Neuromorphic and Quantum Understanding for Innovative Research and Engineering (INQUIRE) Lab(智能神经形态与量子理解创新研究与工程实验室)
;
School of Electrical and Computer Engineering(电气与计算机工程学院)
;
University of Oklahoma(俄克拉荷马大学)
;
School of Psychology(心理学学院)
;
Georgia Institute of Technology(佐治亚理工学院)
Comments24 pages, 9 figures. Conceptual framework paper introducing the AI Evaluability Gap, Evaluability as evidence sufficiency for governance decisions, Operational Certification, Investment Certification, and a six-property evidence lifecycle for AI governance