o3-mini vs DeepSeek-R1: Which One is Safer?
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI
Comments arXiv admin note: substantial text overlap with arXiv:2501.17749
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI
Comments arXiv admin note: substantial text overlap with arXiv:2501.17749
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL
Comments Corrected Version - Solved Some Issues with reference compilation by latex
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI
Comments Paper Under Review
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI
Comments 12 pages, 3 figures
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL
Comments Accepted by ACL 2024. Code: https://github.com/TaoShuchang/CONQORD
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL
Comments ACL 2024
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY
Comments 8 pages, 6 figures, 1 table
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL
Comments Pre-print
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG
Journal ref Amini, MR., Canu, S., Fischer, A., Guns, T., Kralj Novak, P., Tsoumakas, G. (eds) Machine Learning and Knowledge Discovery in Databases. ECML PKDD 2022. Lecture Notes in Computer Science(), vol 13717. Springer, Cham
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI
Comments Experiment code available on TechRxiv: https://www.techrxiv.org/articles/preprint/Computable_Artificial_General_Intelligence/19740190
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG
Comments 7 pages, 12 figures, 3 tables
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG
Comments 21 pages, 10 figures
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG
Comments Appears in SIGKDD 2020 epiDAMIK
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG
Comments 21 pages, extended version with supplementary material
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI
利用自然语言处理对大语言模型生成的反驳进行质量评估自动化
专题命中 其他安全 :safety(abstract,journal_ref);alignment(abstract)
AI总结 研究针对大语言模型生成的反驳质量评估依赖人工且主观的问题,提出结合保证案例图结构特征、语义嵌入和元分类器的自动化评估方法,经案例研究验证,该方法能减少主观差异,为保证案例审查提供决策支持。
Comments 10 pages, 2 figures. Author preprint version of a paper published at ICSRS 2025
Journal ref 2025 9th International Conference on System Reliability and Safety (ICSRS), Turin, Italy, 2025, pp. 101 110
自蒸馏作为大语言模型的性能恢复机制:对抗压缩与灾难性遗忘
机构 * PayPal AI
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出基于自蒸馏微调的性能恢复框架,通过理论分析和实验验证,证明自蒸馏能有效恢复模型能力,揭示了高维流形对齐与性能恢复之间的强相关性。
Comments 18 pages, 8 figures
注意力流动:通过故事摘要追踪LLM概念性参与
机构 * Cornell University(康奈尔大学) ; Aarhus University(奥胡斯大学)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 研究通过比较人类与LLM生成的故事摘要,分析模型在文本中的概念性参与模式,发现模型更关注文本结尾,揭示了摘要生成任务的复杂性。
Comments Error found in data creation pipeline
CForce:通过一致性强制提升扩散大语言模型(dLLMs)的并行解码性能
机构 * Shanghai Jiao Tong University(上海交通大学) ; Ant Group(蚂蚁集团)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出CForce方法,通过一致性强制提升dLLMs的并行解码性能,在LLaDA模型上实验证实其在高并行解码预算下可优化速度-质量权衡。