Interpretability Framework for LLMs in Undergraduate Calculus
机构 * University of Texas at Tyler(德克萨斯理工大学) ; Florida Gulf Coast University(佛罗里达盖恩斯维尔大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * University of Texas at Tyler(德克萨斯理工大学) ; Florida Gulf Coast University(佛罗里达盖恩斯维尔大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * University of Peradeniya(珀德尼亚大学) ; RMIT University(皇家墨尔本理工大学)
专题命中 安全评测 :alignment(abstract);trustworthy(abstract)
Comments 10 Pages + 15 Supplementary Material Pages, 5 figures
机构 * Brigthlands Institute for a Smart Society(智能社会研究院)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.CY、cs.LG
机构 * Department of Computer, Control and Management Engineering, Sapienza University of Rome(计算机、控制与管理工程系,罗马萨皮恩扎大学) ; Department of Computer Science, Sapienza University of Rome(计算机科学系,罗马萨皮恩扎大学) ; Department of Legal, Social, and Educational Sciences, Tuscia University(法律、社会与教育科学系,图斯西亚大学) ; Department of Psychology, Sapienza University of Rome(心理学系,罗马萨皮恩扎大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
Comments Please refer to published version: https://doi.org/10.1073/pnas.2518443122
Journal ref Proc. Natl. Acad. Sci. U.S.A. 122 (42) e2518443122, 2025
机构 * Department of Electrical and Computer Engineering, Maroun Semaan Faculty of Engineering and Architecture(电气与计算机工程系,马鲁恩·塞马安工程与建筑学院) ; American University of Beirut(贝鲁特美国大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) ; The Chinese University of Hong Kong(香港中文大学) ; Tsinghua University(清华大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Tsinghua University(清华大学) ; The Ohio State University(俄亥俄州立大学) ; UC Berkeley(加州大学伯克利分校)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Published in ICLR 2024
机构 * University of Virginia(弗吉尼亚大学) ; Dexcom(德科姆公司)
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY、cs.LG
机构 * Fudan University(复旦大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Imperial College London(帝国理工学院) ; University of Cambridge(剑桥大学)
专题命中 安全评测 :alignment(abstract);safety(abstract)
Comments 26 pages, 13 figures
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * KRAFTON ; Seoul National University(首尔国立大学) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ICLR 2025 (spotlight)
机构 * University of Bucharest Faculty of Mathematics and Computer Science(布加勒斯大学数学与计算机科学学院)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.CY、cs.LG
Comments Accepted as long paper @RANLP2025
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 安全评测 :alignment(abstract);safety(abstract)
机构 * Northeastern University(东北大学) ; Princeton University(普林斯顿大学) ; University of Maryland(马里兰大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Bloomberg(彭博)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG
Comments EMNLP 2025 main conference
机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德人工智能大学) ; Khalifa University(卡利法大学)
专题命中 安全评测 :alignment(abstract);trustworthy(abstract)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY
Comments 65 pages, 26 figures, 6 tables
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG
机构 * Andrew Kiruluta and Priscilla Burity(独立研究者)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Oracle AI
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to EMNLP 2025
机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫罕默德·本·扎耶德人工智能大学) ; Swiss Federal Institute of Technology Lausanne (EPFL)(洛桑联邦理工学院) ; Australian National University(澳大利亚国立大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to EMNLP 2025
机构 * Cairo University, Faculty of Engineering, Computer Engineering Department(开罗大学工程学院计算机工程系) ; Menzies School of Health Research, Charles Darwin University, NT, Australia(梅恩兹健康研究学院,查尔斯达尔文大学,澳大利亚NT) ; Department of Biology, Boston College, Massachusetts, USA(生物学系,波士顿学院,马萨诸塞州,美国) ; Peoples' Friendship University of Russia (RUDN University)(俄罗斯人民友谊大学(RUDN大学)) ; Joint Institute for Nuclear Research(联合核研究中心) ; Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) ; University of Skövde(斯德哥尔摩大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Proceedings of the BioCreative IX Challenge and Workshop (BC9): Large Language Models for Clinical and Biomedical NLP at the International Joint Conference on Artificial Intelligence (IJCAI), Montreal, Canada, 2025
机构 * Technological University of Uruguay, UTEC, Uruguay(乌拉圭技术大学)
专题命中 安全评测 :safety(abstract);trustworthy(abstract)
机构 * School of Psychology, South China Normal University(南方科技大学心理学院) ; Information Hub, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)信息中心) ; School of AI, Guangzhou University(广州大学人工智能学院) ; College of Cyber Security, Jinan University(济南大学网络安全学院)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 5 pages, 2 figures
机构 * Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT(宁波空间智能与数字衍生关键实验室,数字孪生研究院,EIT) ; Logic Intelligence Technology(逻辑智能技术) ; BUPT(北京邮电大学) ; Xiamen University(厦门大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Department of Computing Science University of Alberta(计算科学系阿尔伯塔大学) ; ServiceNow Research(ServiceNow研究)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 7 pages