Probing the Subtle Ideological Manipulation of Large Language Models
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY
机构 * Department of Computer Science University of California, Davis(计算机科学系加州大学戴维斯分校)
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY
Comments Published in the Proceedings of the 40th ACM/SIGAPP Symposium on Applied Computing (SAC'25), March 31--April 4, 2025, Catania, Italy
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
Comments Accepted by WWW2025 Main Track
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY
Comments 7 pages, 2 figures. Presented in SIAM International Conference on Data Mining (SDM25) METACOG-25: 2nd Workshop on Metacognitive Prediction of AI Behavior
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
Comments This work was submitted to the 32nd International Conference on Real-Time Networks and Systems (RTNS) on June 8, 2024
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY
Comments 12 pages
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG
Comments ICLR 2025
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
Comments Code and model checkpoints are available at https://github.com/Jam1ezhang/RankCLIP
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI
Comments Paper accepted for publication in the ACM Transactions on Intelligent Systems and Technology (TIST) journal
专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI
Comments Accepted to LaTeCH-CLfL 2025, held in conjunction with NAACL 2025
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
Comments This paper has been accepted for presentation at the International Conference on Machine Learning (ICML) 2024 and will appear in the conference proceedings