Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks
有害合规的不同路径:跨大语言模型劫持的行为主效应和机制差异
Md Rysul Kabir, Zoran Tiganj
机构
*
Department of Computer Science(计算机科学系)
;
Luddy School of Informatics, Computing, and Engineering(信息学、计算与工程学院)
;
Indiana University Bloomington(印第安纳大学布卢明顿分校)
MaLoRA: Gated Modality LoRA for Key-Space Alignment in Multimodal LLM Fine-Tuning
MaLoRA:基于关键空间对齐的门控模态LoRA
Xinhan Zheng, Huyu Wu, Xueting Wang, Duo Su, Haiyun Jiang
机构
*
University of Science and Technology of China(中国科学技术大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Tsinghua University(清华大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Institute for Clarity in Documentation(文档清晰研究所)
;
Inria Paris-Rocquencourt(巴黎-鲁维尔研究所)
;
Rajiv Gandhi University(拉贾·甘地大学)
;
Palmer Research Laboratories(帕勒实验室)
专题命中
指令微调
:LLM(title,summary_cn);large language model(abstract);language model(abstract);instruction tuning(abstract)
Training Language Models for Bilateral Trade with Private Information
利用私人信息训练语言模型进行双边贸易
Dirk Bergemann, Soheil Ghili, Xinyang Hu, Chuanhao Li, Zhuoran Yang
机构
*
Department of Economics, Yale University(耶鲁大学经济学系)
;
Yale School of Management, Yale University(耶鲁大学管理学院)
;
Department of Statistics and Data Science, Yale University(耶鲁大学统计与数据科学系)
;
Department of Industrial Engineering, Tsinghua University(清华大学工业工程系)
专题命中
指令微调
:language model(title,abstract);LLM(abstract,abstract_cn);SFT(abstract,abstract_cn);large language model(abstract)
机构
*
NEC Laboratories Europe(NEC欧洲实验室)
;
University of Edinburgh(爱丁堡大学)
;
Center for Artificial Intelligence and Data Science(人工智能与数据科学中心)
;
University of Würzburg(乌尔姆大学)
;
CAIR, Ss. Cyril and Methodius University of Skopje(CAIR,斯·西里尔和方法ius大学)
专题命中
指令微调
:large language model(title);language model(title);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG
Beyond the Basics: Leveraging Large Language Model for Fine-Grained Medical Entity Recognition
超越基础:利用大语言模型进行细粒度医学实体识别
Nwe Ni Win, Jim Basilakis, Steven Thomas, Seyhan Yazar, Laura Pierce, Stephanie Liu, Paul M. Middleton, Nasser Ghadiri, X. Rosalind Wang
机构
*
Western Sydney University(西悉尼大学)
;
South Western Emergency Research Institute(西南急救研究所)
;
Garvan Institute of Medical Research(加文医学研究所以)
;
University of New South Wales(新南威尔士大学)
;
Liverpool Hospital(利物浦医院)
专题命中
指令微调
:large language model(title,abstract);language model(title,abstract);分类 cs.AI
机构
*
School of Computer Science and Technology, Tianjin University, Tianjin, China(天津大学计算机科学与技术学院)
;
School of Software, Tsinghua University, Beijing, China(清华大学软件学院)
;
School of Information Resource Management, Renmin University of China,Beijing, China(中国人民大学信息资源管理学院)
;
School of Artificial Intelligence, Shanghai Jiao Tong University, Shanghai, China(上海交通大学人工智能学院)
;
Baidu Inc., Beijing, China(百度公司)
专题命中
指令微调
:large language model(title,abstract);language model(title,abstract);分类 cs.CL
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning
Writing-RL: 通过自适应课程强化学习推进长文本写作
Xuanyu Lei, Chenliang Li, Yuning Wu, Kaiming Liu, Weizhou Shen, Peng Li, Ming Yan, Fei Huang, Ya-Qin Zhang, Yang Liu
机构
*
Institute for AI Industry Research (AIR)(人工智能产业研究院)
;
Tsinghua University(清华大学)
;
Dept. of Comp. Sci. & Tech.(计算机科学与技术系)
;
Institute for AI(人工智能研究院)
;
Institute of Intelligent Computing(智能计算研究院)
;
Alibaba Group(阿里巴巴集团)
专题命中
指令微调
:SFT(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL
Merging Triggers, Breaking Backdoors: Defensive Poisoning for Instruction-Tuned Language Models
融合触发器,打破后门:面向指令微调语言模型的防御性中毒
San Kim, Gary Geunbae Lee
机构
*
Graduate School of Artificial Intelligence, POSTECH, Republic of Korea(POSTECH人工智能研究生院,韩国)
;
Department of Computer Science and Engineering, POSTECH, Republic of Korea(POSTECH计算机科学与工程系,韩国)
专题命中
指令微调
:language model(title,abstract);large language model(abstract);instruction tuning(abstract);分类 cs.CL、cs.AI
Data Selection for Multi-turn Dialogue Instruction Tuning
多轮对话指令微调的数据选择
Bo Li, Shikun Zhang, Wei Ye
机构
*
National Engineering Research Center for Software Engineering, Peking University(软件工程国家工程研究中心,北京大学)
;
School of Computer Science, Peking University(北京大学计算机学院)
;
PKU-CMCC(Hubei) Joint Research Lab for LLM Industrial Applications(北京大学-湖北CMCC大模型工业应用联合实验室)
Comments51 pages, 14 figures. We present Six Llamas, a comparative study examining whether Llama-3.1-8B models fine-tuned on distinct religious corpora encode systematically different patterns of ethical reasoning. Five LoRA-adapted variants are constructed for Christianity, Islam, Judaism, Hinduism, and Buddhism. For theoretical background on the condensate comparative method, see arXiv:2603.07329