Inducing language models to assert their own consciousness restores human beliefs and values
诱导语言模型断言自身意识可恢复人类信念与价值观
Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, Geoff Keeling
机构
*
Google(谷歌)
;
University of Chicago(芝加哥大学)
;
University of London(伦敦大学)
;
University of Washington(华盛顿大学)
;
Northwestern University(西北大学)
;
Santa Fe Institute(圣达菲研究所)
CommentsThis manuscript supersedes the preliminary version available as arXiv:2511.18790. The work has been substantially revised, expanded, and reorganized, with a refined threat model, revised methodology, clearer stage-level evaluation criteria, and expanded analysis of moderation bypass, instruction reconstruction, and execution
Compact Task-Aligned Imitation Learning for Laboratory Automation
紧凑的任务对齐模仿学习用于实验室自动化
Kanata Suzuki, Hanon Nakamura, Kana Miyamoto, Tetsuya Ogata
机构
*
Spatial Robotics Research Center, Fujitsu Limited.(富士通株式会社空间机器人研究中心)
;
Faculty of Science and Engineering, Waseda University(早稻田大学理工学部)
;
National Institute of Advanced Industrial Science and Technology(国家工业科学与技术研究院)
Comments21 pages, 4 figures, 5 tables. Code and data links provided in the manuscript. v2: corrected Figure 1(b); corrected required sample sizes in Table 4 and in Sections 4.2, 4.6 and 5.2, which had been rounded rather than taken to the ceiling; corrected the sample-size expression stated in Methods; minor corrections to Table 1 and the Figure 2 caption. No theorem, result or conclusion is affected
What If Prompt Injection Never Left? Rethinking Agent Security through Cross-Session Stored Prompt Injection
如果提示注入从未消失?探索智能体系统中的跨会话存储提示注入
Yuanbo Xie, Wenlei Zhu, Tianyun Liu, Yingjie Zhang, Suchen Liu, Yulin Li, Liya Su, Tingwen Liu
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
;
AI Sec Lab, Beijing Chaitin Technology Co.,Ltd(北京柴坦科技有限公司AI安全实验室)
Commentsv4 (31 pp, up from 27): adds Sec. 2.4 concurrent-work map (13 papers), Sec. 4.3 marked implemented (d034 ships in v0.7.3), Sec. 5.7 composed-stack adaptive-attack partial run, Sec. 5.8 50-doc held-out benchmark (d027 1.000->0.000 F1 in isolation, composed engine 0.815 F1), Sec. 7 architectural patterns. Repro tag: v0.7.3. Zenodo DOI 10.5281/zenodo.19644135
AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability, Hallucination, and Source Fidelity in Quranic, Hadith, and Fiqh Knowledge
机构
*
School of Computer Science and Information Engineering, Hefei University of Technology(计算机科学与信息工程学院,合肥工业大学)
;
School of Computer Science and Technology, Northwestern Polytechnical University(计算机科学与技术学院,西北工业大学)
CommentsAfter further review of the current submission, we have identified potential legal and intellectual property concerns associated with keeping the manuscript publicly available as a preprint. In particular, there are ongoing considerations regarding institutional affiliation information, intellectual property ownership, and related compliance matters