Comments48 pages, 27 tables, 4 figures. v3 is a correction release: a full source-verification pass over all 83 references corrects 30 defective entries and withdraws two ancillary evaluations run on synthetic stand-in data; no measured numbers changed. Changelog: 10.5281/zenodo.21858218. Earlier versions: Zenodo concept DOI 10.5281/zenodo.19353663, TDCommons dpubs_series/9683
机构
*
The University of Hong Kong(香港大学)
;
Nanjing University(南京大学)
;
University of Science and Technology of China(中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
;
Fudan University(复旦大学)
Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?
我们真的需要参数超过10亿的多模态情感语言模型吗?
Kaiwen Zheng, Junchen Fu, Wenhao Deng, Hu Han, Joemon M. Jose, Xuri Ge
机构
*
University of Glasgow(格拉斯哥大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
School of Artificial Intelligence, Shandong University(山东大学人工智能学院)
Comments62 pages (31-page article and 31-page supplementary information), 8 figures, 4 tables. v3: corrects the author metadata to the sole author, Kwan Soo Shin; revised title and abstract; adds cross-vendor and flagship validation, signal-detection and specified-task controls, and dual-process probes. Reproducibility deposit: doi:10.5281/zenodo.20826823
MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
MENTOR: 一种元认知驱动的自我进化框架,用于发现和缓解大语言模型中的隐式领域风险
Liang Shan, Kaicheng Shen, Wen Wu, Zhenyu Ying, Chaochao Lu, Yan Teng, Jingqi Huang, Qingshan Liu, Guangze Ye, Guoqing Wang, Jie Zhou, Liang He
机构
*
School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院)
;
Shanghai AI Lab, Shanghai Innovation Institute(上海人工智能实验室,上海创新研究院)
TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speech
TANDEM: 面向多模态仇恨言论的时间感知神经检测
Girish A. Koushik, Helen Treharne, Diptesh Kanojia
机构
*
Nature-Inspired Computing & Engineering, University of Surrey(Surrey大学自然启发计算与工程系)
;
Surrey Centre for Cyber Security, University of Surrey(Surrey大学网络安全中心)
Haochen Huang, Yue Su, Xin Sun, Moonisa Ahsan, Mohammad Aliannejadi, Irene Viola, Zhaochun Ren, Chuang Yu, Aneta Lisowska, Artem Belopolsky, Koen Hindriks, Pablo Cesar, Junxiao Wang, Jiahuan Pei
机构
*
Vrije University of Amsterdam University of Amsterdam Centrum Wiskunde \& Informatic
;
University of Amsterdam National Institute of Informatics Centrum Wiskunde \& Informatic
;
Centrum Wiskunde \& Informatic Leiden University University College London
;
Vrije University of Amsterdam
;
Centrum Wiskunde \& Informatica Technische Universiteit Delft Guangzhou University
;
Vrije University of Amsterdam Centrum Wiskunde \& Informatic
Comments20 pages, 1 figure, 9 tables. v2 adds Round 2: Russian-market coding agents (SourceCraft CLI, Koda CLI), Antigravity with Gemini 3.1 Pro / 3.5 Flash, and Codex CLI with GPT-5.6 on the same frozen task set, plus a tool-call contamination re-audit (network + disk layers). Data, full trajectories and harness: https://github.com/eugeneshilow/rubench
机构
*
Department of Chemical and Materials Engineering, New Mexico State University(新墨西哥州立大学化学与材料工程系)
;
Department of Electrical and Communications Engineering, New Jersey Institute of Technology(新泽西理工学院电气与通信工程系)
机构
*
Independent Researchers(独立研究者)
;
Tencent(腾讯)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Tsinghua University(清华大学)