Insider Attacks in Multi-Agent LLM Consensus Systems
多智能体大语言模型共识系统中的内部攻击
机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,Location,Country) ; School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,Location,Country) ; Department of Computer Science, Tulane University, New Orleans, United States of America(计算机科学系, Tulane大学,新奥尔良,美国)
专题命中 其他LLM :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
AI总结 研究多智能体大语言模型共识系统中的内部攻击问题,提出基于世界模型的框架,通过学习良性智能体的潜在行为状态并利用强化学习训练攻击者,有效降低共识率并延长分歧时间。