Comments17 pages, 3 figures, 6 tables (9-page main text). Ancillary file hidden-automata-rl-code.zip contains reproduction code and the complete per-run data behind every table and figure. v3: author name corrected, title revised, text rewritten for clarity, new robustness checks (nonlinear and off-policy probes) added; results and conclusions unchanged
ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents
VESTA: 一种全自动的LLM智能体场景生成与安全评估框架
Lu Jia, Haibo Tong, Feifei Zhao, Jindong Li, Dongqi Liang, Ping Wu, Qian Zhang, Yi Zeng
机构
*
BrainCog AI Lab, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所类脑人工智能实验室)
;
Beijing Institute of AI Safety and Governance (Beijing-AISI)(北京人工智能安全与治理研究院)
;
Beijing Key Laboratory of Safe AI and Superalignment(北京市安全人工智能与超级对齐重点实验室)
;
School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院)
;
Long-term AI(长期人工智能)
机构
*
University of Southern California(南加州大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Stanford University(斯坦福大学)
Comments17 pages, 5 figures, 9 tables. v2 corrects scorer and taxonomy defects, adds no-model baselines showing label leakage, re-runs the Lithuanian cells on de-leaked text, and withdraws the claim that few-shot helps on judgment-form classification everywhere; all tables and figures regenerated. Dataset: https://huggingface.co/datasets/overthelex/multi-legal-bench
Minimal Ingredients for Reward Assignment from Expert Demonstrations
从专家演示中进行奖励分配的最小要素
Zixuan Dong, Yumi Omori, Keith Ross
机构
*
New York University(纽约大学)
;
New York University Abu Dhabi(纽约大学阿布扎赫尔分校)
;
Shanghai Frontiers Science Center of Artificial Intelligence and Deep Learning, NYU Shanghai(上海前沿科学中心(人工智能与深度学习))
Skill-RAG: Failure-State-Aware Retrieval Augmentation via Hidden-State Probing and Skill Routing
Skill-RAG: 通过隐藏状态探测与技能路由实现故障状态感知的检索增强
Kai Wei, Raymond Li, Xi Zhu, Zhaoqian Xue, Jiaojiao Han, Jingcheng Niu, Fan Yang
机构
*
University of Michigan(密歇根大学)
;
University of British Columbia(不列颠哥伦比亚大学)
;
Rutgers University(罗格斯大学)
;
University of Pennsylvania(宾夕法尼亚大学)
;
New Jersey Institute of Technology(新泽西理工学院)
;
TU Darmstadt(图腾斯大学)
;
Wake Forest University(威克森林大学)
Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP
Grad-ECLIP: 基于梯度的CLIP视觉与文本解释
Chenyang Zhao, Kun Wang, Janet H. Hsiao, Antoni B. Chan
机构
*
Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系)
;
Division of Social Science and Department of Computer Science & Engineering, Hong Kong University of Science & Technology(香港科学与技术大学社会科学学院及计算机科学与工程系)
;
SenseTime Group Ltd(时光集团有限公司)
Journal refZhao C, Wang K, Hsiao J H, et al. Grad-eclip: Gradient-based visual and textual explanations for clip[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026