arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Edinburgh(爱丁堡大学)

共收录 812
2601.22311 2026-02-02 cs.AI cs.CL cs.LG

Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents

为何推理无法规划:从规划角度分析LLM代理在长周期决策中的表现

Zehong Wang, Fang Wu, Hongru Wang, Xiangru Tang, Bolian Li, Zhenfei Yin, Yijun Ma, Yiyang Li, Weixiang Sun, Xiusi Chen, Yanfang Ye

机构 * University of Notre Dame(诺丁汉大学) Stanford University(斯坦福大学) Purdue University(普渡大学) University of Edinburgh(爱丁堡大学) University of Oxford(牛津大学) Yale University(耶鲁大学)

AI总结 本文从规划角度分析LLM代理在长周期决策中的失败原因,提出FLARE方法通过前瞻和价值传播提升规划能力,使LLaMA-8B在多个基准中超越GPT-4o。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13036 2026-01-30 cs.AI cs.LG

Repairing Reward Functions with Feedback to Mitigate Reward Hacking

通过反馈修复奖励函数以缓解奖励黑客

Stephane Hatgis-Kessell, Logan Mondal Bhamidipaty, Emma Brunskill

机构 * Computer Science Department, Stanford University(计算机科学系, 斯坦福大学) School of Informatics, The University of Edinburgh(信息学院, 埃迪索恩大学)

AI总结 通过反馈修复奖励函数以缓解奖励黑客,提出PBRR方法,通过学习过渡依赖的修正项来改进代理奖励函数,从而在较少偏好下实现高性能策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10163 2026-01-30 cs.CL

Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions

Compound-QA:一个评估大语言模型在复合问题上的基准

Yutao Hou, Yajing Luo, Zhiwen Ruan, Hongru Wang, Weifeng Ge, Yun Chen, Guanhua Chen

机构 * Shanghai University of Finance and Economics(上海财经大学) University of Edinburgh(爱丁堡大学) Southern University of Science and Technology(南方科技大学) Fudan University(复旦大学)

AI总结 Compound-QA基准通过评估LLM在复合问题上的表现,揭示其在理解和推理方面的不足,并提出改进策略。

Comments Accepted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11232 2026-01-30 eess.IV cs.CV

Scale-Equivariant Imaging: Self-Supervised Learning for Image Super-Resolution and Deblurring

尺度等变成像:用于图像超分辨率和去模糊的自监督学习

Jérémy Scanvic, Mike Davies, Patrice Abry, Julián Tachella

机构 * ENSL, CNRS, Laboratoire de Physique(ENSL、CNRS、物理实验室) School of Engineering, University of Edinburgh(工程学院、爱丁堡大学)

AI总结 本文提出尺度等变成像方法,通过利用图像分布的尺度不变性,提升图像超分辨率和去模糊任务的性能,实现与全监督学习相当的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18944 2026-01-29 cs.AI cs.PL cs.SE

Neural Theorem Proving for Verification Conditions: A Real-World Benchmark

为验证条件进行神经定理证明:一个现实世界的基准测试

Qiyuan Xu, Xiaokun Luan, Renxi Wang, Joshua Ong Jun Leang, Peixin Wang, Haonan Li, Wenda Li, Conrad Watt

机构 * Nanyang Technological University(南洋理工大学) Peking University(北京大学) MBZUAI Imperial College London(帝国理工学院) East China Normal University(华东师范大学) University of Edinburgh(爱丁堡大学)

AI总结 本研究提出NTP4VC,首个现实世界多语言基准测试,评估LLMs在自动验证条件证明中的表现,揭示程序验证中的挑战与未来研究方向。

Comments Accepted in ICLR'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20708 2026-01-29 hep-lat cond-mat.stat-mech cs.LG hep-ph

A scalable flow-based approach to mitigate topological freezing

一种可扩展的基于流的方法用于缓解拓扑冻结

Claudio Bonanno, Andrea Bulgarelli, Elia Cellini, Alessandro Nada, Dario Panfalone, Davide Vadacchino, Lorenzo Verzichelli

机构 * Albert Einstein Center for Fundamental Physics, Institute for Theoretical Physics, University of Bern(阿尔伯特·爱因斯坦基础物理学研究中心,理论物理学研究所,伯尔尼大学) Higgs Centre for Theoretical Physics, School of Physics and Astronomy, The University of Edinburgh(希格斯理论物理学中心,物理与天文学学院,爱丁堡大学) Centre for Mathematical Sciences, University of Plymouth(数学科学中心,普利茅斯大学)

AI总结 本文提出了一种基于流的可扩展方法,通过传输具有OBC缺陷的配置到完全周期性集合,缓解拓扑冻结问题,并在4d SU(3)杨-米尔斯理论中验证其有效性。

Comments 1+9 pages, 3 figures, contribution to the 42nd International Symposium on Lattice Field Theory (Lattice 2025), 2-8 November 2025, Mumbai, India

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20706 2026-01-29 cs.AR cs.AI cs.DC

Beyond GEMM-Centric NPUs: Enabling Efficient Diffusion LLM Sampling

超越以GEMM为中心的NPUs:使高效的扩散LLM采样成为可能

Binglei Lou, Haoran Wu, Yao Lai, Jiayi Nie, Can Xiao, Xuan Guo, Rika Antonova, Robert Mullins, Aaron Zhao

机构 * Imperial College London(伦敦帝国学院) University of Edinburgh(爱丁堡大学) University of Cambridge(剑桥大学)

AI总结 本文提出了一种针对扩散LLM采样的NPU架构优化方案,通过轻量级向量原语、内存重用策略和混合精度内存层次结构,实现了2.53倍的加速性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06165 2026-01-29 cs.LG eess.SP math.ST stat.ML stat.TH

Higher-Order Feature Attribution: Bridging Statistics, Explainable AI, and Topological Signal Processing

高阶特征归因:连接统计学、可解释AI与拓扑信号处理

Kurt Butler, Guanchao Feng, Petar Djuric

机构 * School of Engineering, The University of Edinburgh(爱丁堡大学工程学院) Causality in Healthcare AI Hub (CHAI)(医疗AI因果性研究中心) Department of Electrical and Computer Engineering, Stony Brook University(石溪大学电气与计算机工程系)

AI总结 本文提出高阶特征归因理论,基于集成梯度方法,连接统计学与拓扑信号处理,扩展可解释AI框架并验证其有效性。

Comments 5 pages, 3 figures, to be published in the Proceedings of ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19747 2026-01-28 cs.AR cs.AI cs.SE

Veri-Sure: A Contract-Aware Multi-Agent Framework with Temporal Tracing and Formal Verification for Correct RTL Code Generation

Veri-Sure:一种具有时间追踪和形式验证的合同感知多智能体框架,用于正确RTL代码生成

Jiale Liu, Taiyu Zhou, Tianqi Jiang

机构 * School of Physics and Astronomy, The University of Edinburgh, Edinburgh, UK(物理与天文学院,爱丁堡大学,爱丁堡,英国) State Key Laboratory of Analog and Mixed-Signal VLSI, University of Macau, Macau(模拟与混合信号VLSI国家重点实验室,澳门大学,澳门) School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, Shenzhen, China(科学与工程学院,香港中文大学(深圳))

AI总结 Veri-Sure通过合同感知多智能体框架结合时间追踪与形式验证,实现高正确性的RTL代码生成,优于独立LLMs和传统智能体系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19053 2026-01-28 cs.HC cs.AI

From Answer Givers to Design Mentors: Guiding LLMs with the Cognitive Apprenticeship Model

从答案提供者到设计导师:通过认知 apprenticeship 模型引导 LLMs

Yongsu Ahn, Lejun R Liao, Benjamin Bach, Nam Wook Kim

机构 * Boston College(波士顿学院) Inria(法国国家信息与自动化技术研究院) University of Edinburgh(爱丁堡大学)

AI总结 本文通过认知 apprenticeship 模型引导 LLMs 作为设计导师,提升设计推理和反思反馈质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20691 2026-01-28 cs.AI

Plan Then Retrieve: Reinforcement Learning-Guided Complex Reasoning over Knowledge Graphs

计划后再检索:强化学习引导的知识图谱复杂推理

Yanlin Song, Ben Liu, Víctor Gutiérrez-Basulto, Zhiwei Hu, Qianqian Xie, Min Peng, Sophia Ananiadou, Jeff Z. Pan

机构 * School of Computer Science(计算机科学学院) Wuhan University(武汉大学) School of Computer Science and Informatics(计算机科学与信息学院) Cardiff University(卡迪夫大学) College of Information Science and Engineering(信息科学与工程学院) Shanxi Agricultural University(山西农业大学) School of Artificial Intelligence(人工智能学院) Center for Language and Information Research(语言与信息研究中心) University of Manchester(曼彻斯特大学) ILCC, School of Informatics(信息学院) University of Edinburgh(爱丁堡大学)

AI总结 Graph-RFT通过强化学习引导的知识图谱复杂推理框架,实现自主规划与适应性检索调度,解决KGQA中的冷启动问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23785 2026-01-28 cs.CL cs.AI cs.CY

Meaning Is Not A Metric: Using LLMs to make cultural context legible at scale

意义并非一种度量:利用大语言模型在大规模AI社会技术系统中使文化背景可理解

Cody Kommers, Drew Hemment, Maria Antoniak, Joel Z. Leibo, Hoyt Long, Emily Robinson, Adam Sobey

机构 * The Alan Turing Institute(艾伦·图灵研究所) University of Edinburgh(爱丁堡大学) University of Colorado Boulder(科罗拉多大学丹佛分校) University of Chicago(芝加哥大学) University of Exeter(埃克塞特大学) University of Southampton(南安普顿大学)

AI总结 本文探讨了利用大语言模型在大规模AI系统中通过厚描述使文化背景可理解的挑战与方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04058 2026-01-27 cs.GR cs.CV

SMooGPT: Stylized Motion Generation using Large Language Models

SMooGPT:利用大语言模型进行风格化动作生成

Lei Zhong, Yi Yang, Changjian Li

机构 * University of Edinburgh(爱丁堡大学)

AI总结 SMooGPT通过利用大语言模型的推理、组合和生成能力,实现风格化动作生成,提升动作控制和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11771 2026-01-22 cs.CV cs.AI

Smudged Fingerprints: A Systematic Evaluation of the Robustness of AI Image Fingerprints

模糊指纹:对AI图像指纹鲁棒性的系统评估

Kai Yao, Marc Juarez

机构 * School of Informatics University of Edinburgh(信息学院爱丁堡大学)

AI总结 本文系统评估了AI图像指纹技术的鲁棒性,发现移除攻击效果显著,而伪造攻击则因目标模型不同而成功率各异,揭示了准确性和鲁棒性之间的权衡。

Comments This work has been accepted for publication in the 4th IEEE Conference on Secure and Trustworthy Machine Learning (IEEE SaTML 2026). The final version will be available on IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14894 2026-01-22 cs.AI

To Neuro-Symbolic Classification and Beyond by Compiling Description Logic Ontologies to Probabilistic Circuits

通过将描述逻辑本体编译为概率电路来超越神经符号分类

Nicolas Lazzari, Valentina Presutti, Antonio Vergari

机构 * University of Pisa(比萨大学) University of Bologna(博洛尼亚大学) University of Edinburgh(爱丁堡大学)

AI总结 通过将描述逻辑本体编译为概率电路,实现神经符号分类与知识表示的深度融合,提升演绎推理效率和预测一致性。

Comments Manuscript under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14051 2026-01-21 cs.CL cs.AI cs.LG

Kakugo: Distillation of Low-Resource Languages into Small Language Models

Kakugo:通过低资源语言名称训练通用小型语言模型的蒸馏方法

Peter Devine, Mardhiyah Sanni, Farid Adilazuarda, Julieta Gil Loizaga, Barry Haddow

机构 * School of Informatics, University of Edinburgh(信息学院,爱丁堡大学)

AI总结 Kakugo通过仅使用语言名称作为输入,利用大型教师模型生成合成数据,有效训练小型语言模型以处理低资源语言,从而在多种NLP任务中提升性能,成本低于50美元每种语言。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14027 2026-01-21 cs.AI

Numina-Lean-Agent: An Open and General Agentic Reasoning System for Formal Mathematics

Numina-Lean-Agent: 一种面向形式数学的开放且通用的代理推理系统

Junqi Liu, Zihao Zhou, Zekai Zhu, Marco Dos Santos, Weikun He, Jiawei Liu, Ran Wang, Yunzhou Xie, Junqiao Zhao, Qiufeng Wang, Lihong Zhi, Jia Li, Wenda Li

机构 * Academy of Mathematics and Systems Science, University of Chinese Academy of Sciences(中国科学院数学与系统科学研究院) Tongji University(同济大学) University of Cambridge(剑桥大学) Imperial College London(伦敦帝国学院) University of Edinburgh(爱丁堡大学) University of Liverpool(利物浦大学) Xi'an Jiaotong-Liverpool University(西安交通大学利物浦大学)

AI总结 Numina-Lean-Agent通过通用编码代理实现形式数学推理,解决Putnam 2025全部问题并成功形式化Brascamp-Lieb定理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13987 2026-01-21 eess.IV cs.CV

SHARE: A Fully Unsupervised Framework for Single Hyperspectral Image Restoration

SHARE: 一种完全无监督的单超光谱图像恢复框架

Jiangwei Xie, Zhang Wen, Mike Davies, Dongdong Chen

机构 * School of Mathematical and Computer Sciences, Heriot-Watt University(数学与计算机科学学院,赫里奥特-瓦特大学) School of Engineering, University of Edinburgh(工程学院,爱丁堡大学)

AI总结 SHARE提出了一种完全无监督的单超光谱图像恢复框架,结合几何等变原理和低秩光谱建模,无需地面真实数据,通过自监督信号和动态自适应光谱注意力模块提升恢复性能。

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13697 2026-01-21 cs.CL cs.AI cs.LG

Uncertainty-Aware Gradient Signal-to-Noise Data Selection for Instruction Tuning

考虑不确定性的梯度信号-噪声数据选择用于指令微调

Zhihang Yuan, Chengyu Yue, Long Huang, Litu Ou, Lei Shi

机构 * Alibaba Cloud Computing(阿里巴巴云 computing) The University of Edinburgh(爱丁堡大学)

AI总结 GRADFILTERING通过利用GPT-2代理和LoRA集成,提出了一种考虑不确定性的数据选择方法,以提高指令微调的效率和效果。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13398 2026-01-21 cs.LG cs.AI cs.PL

Can LLMs Compress (and Decompress)? Evaluating Code Understanding and Execution via Invertibility

LLMs能否压缩(和解压)?通过可逆性评估代码理解和执行

Nickil Maveli, Antonio Vergari, Shay B. Cohen

机构 * School of Informatics, University of Edinburgh(爱丁堡大学信息学院)

AI总结 本文通过RTCE基准评估LLMs在代码理解和执行中的回程一致性,发现现有模型在保持一致推理方面存在不足,揭示了新的研究洞察。

Comments 32 pages (preprint)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02134 2026-01-21 eess.SP cs.LG

Robust Channel Estimation for Optical Wireless Communications Using Neural Network

使用神经网络的光学无线通信中鲁棒信道估计

Dianxin Luan, John Thompson

机构 * Institute for Imaging, Data and Communications, School of Engineering, University of Edinburgh(成像、数据与通信研究所,工程学院,爱丁堡大学)

AI总结 本文提出了一种基于神经网络的鲁棒信道估计方法,用于提高光学无线通信系统的可靠性和性能。

Comments Accepted by WCL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21046 2026-01-21 cs.AI

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence

自我进化代理的综述:何时、何地、如何进化以实现人工超级智能

Huan-ang Gao, Jiayi Geng, Wenyue Hua, Mengkang Hu, Xinzhe Juan, Hongzhang Liu, Shilong Liu, Jiahao Qiu, Xuan Qi, Yiran Wu, Hongru Wang, Han Xiao, Yuhang Zhou, Shaokun Zhang, Jiayi Zhang, Jinyu Xiang, Yixiong Fang, Qiwen Zhao, Dongrui Liu, Qihan Ren, Cheng Qian, Zhenhailong Wang, Minda Hu, Huazheng Wang, Qingyun Wu, Heng Ji, Mengdi Wang

机构 * Princeton University(普林斯顿大学) Princeton AI Lab(普林斯顿人工智能实验室) Tsinghua University(清华大学) Carnegie Mellon University(卡内基梅隆大学) University of Sydney(悉尼大学) Shanghai Jiao Tong University(上海交通大学) Pennsylvania State University(宾夕法尼亚州立大学) University of Michigan(密歇根大学) Oregon State University(俄勒冈州立大学) The Chinese University of Hong Kong(香港中文大学) Fudan University(复旦大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The University of Hong Kong(香港大学) University of California, Santa Barbara(加州大学圣芭芭拉分校) University of California San Diego(加州大学圣地亚哥分校) University of Edinburgh(爱丁堡大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 本文综述了自我进化代理的现状,探讨了进化机制、适应方法及挑战,为实现人工超级智能提供路线图。

Comments 77 pages, 9 figures, Transactions on Machine Learning Research (01/2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13137 2026-01-19 cs.AI

Theorem Prover as a Judge for Synthetic Data Generation

定理推理解释器作为合成数据生成的裁判

Joshua Ong Jun Leang, Giwon Hong, Wenda Li, Shay B. Cohen

机构 * School of Informatics, The University of Edinburgh(信息学院,爱丁堡大学)

AI总结 本研究提出TP-as-a-Judge和RLTPF方法,通过定理推理解释器提升LLM的合成数据生成与推理准确性。

Journal ref Proc. ACL 2025, pp. 29941-29977

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10336 2026-01-16 cs.AI cs.CL cs.LG cs.SC

CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning

CoMAT:数学注释的思维链提升数学推理

Joshua Ong Jun Leang, Aryo Pradipta Gema, Shay B. Cohen

机构 * School of Informatics, The University of Edinburgh(信息学院,爱丁堡大学) Imperial College London(伦敦帝国学院)

AI总结 CoMAT通过符号转换和推理执行两个阶段提升数学推理能力,在多个基准测试中超越传统CoT方法。

Comments 9 pages, 12 figures

Journal ref Proc. EMNLP 2025, pp. 20245-20274

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19828 2026-01-15 cs.CL cs.MA

Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning

Memory-R1: 通过强化学习增强大语言模型代理以管理并利用记忆

Sikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma, Jinhe Bi, Kristian Kersting, Jeff Z. Pan, Hinrich Schütze, Volker Tresp, Yunpu Ma

机构 * Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学) Munich Center for Machine Learning(慕尼黑机器学习中心) Technical University of Munich(慕尼黑技术大学) University of Cambridge(剑桥大学) University of Hong Kong(香港大学) Technical University of Darmstadt(达姆施塔特技术大学) University of Edinburgh(爱丁堡大学)

AI总结 Memory-R1通过强化学习框架,使大语言模型具备主动管理与利用外部记忆的能力,有效提升长跨度推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08701 2026-01-15 cs.LO cs.AI cs.CY

A Personalised Formal Verification Framework for Monitoring Activities of Daily Living of Older Adults Living Independently in Their Homes

为独立居家老年人的日常生活活动提供个性化形式验证框架

Ricardo Contreras, Filip Smola, Nuša Farič, Jiawei Zheng, Jane Hillston, Jacques D. Fleuriot

机构 * School of Informatics, The University of Edinburgh(信息学院,爱丁堡大学) University of Edinburgh(爱丁堡大学) DigitLab, University of Exeter(埃克塞特大学数字实验室)

AI总结 本文提出一个个性化形式验证框架,用于独立居家老年人的日常生活活动监测,通过形式化模型和模型检查器确保安全与福祉。

Comments 19 pages, 6 figures

Journal ref For publication in IEEE Sensors Journal 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08840 2026-01-15 cs.CL cs.AI

Consistency-Aware Editing for Entity-level Unlearning in Language Models

面向实体级去学习的一致性感知编辑

Xiaoqi Han, Víctor Gutiérrez-Basulto, Ru Li, Xiaoli Li, Jiye Liang, Jeff Z. Pan

机构 * Shanxi University, China(山西大学) Cardiff University, UK(卡地夫大学) Singapore University of Technology and Design, Singapore(新加坡科技设计大学) ILCC, School of Informatics, University of Edinburgh, Edinburgh, UK(爱丁堡大学信息学院ILCC)

AI总结 本文提出一致性感知编辑框架,用于高效实体级去学习,通过聚合多样化提示并学习低秩更新,提升遗忘准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07765 2026-01-13 cs.CL

Contrastive Learning with Narrative Twins for Modeling Story Salience

基于叙事双胞胎的对比学习用于建模故事显著性

Igor Sterner, Alex Lascarides, Frank Keller

机构 * School of Informatics University of Edinburgh(信息学院爱丁堡大学)

AI总结 本文提出基于叙事双胞胎的对比学习方法,用于建模故事显著性,通过对比学习生成故事嵌入,并验证总结操作在识别显著句子上的有效性。

Comments EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07696 2026-01-13 cs.CL

Exploring the Meta-level Reasoning of Large Language Models via a Tool-based Multi-hop Tabular Question Answering Task

通过基于工具的多跳表格问答任务探索大型语言模型的元级推理

Nick Ferguson, Alan Bundy, Kwabena Nuamah

机构 * School of Informatics University of Edinburgh(信息学院爱丁堡大学)

AI总结 本文通过设计基于工具的多跳表格问答任务,探索大型语言模型的元级推理能力,发现其在任务分解和工具选择方面表现良好,但在数学运算和任务理解上存在不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04736 2026-01-09 cs.CL

AM$^3$Safety: Towards Data Efficient Alignment of Multi-modal Multi-turn Safety for MLLMs

AM$^3$Safety: 向多模态多轮安全对齐的数据高效方法

Han Zhu, Jiale Chen, Chengkun Cai, Shengjie Sun, Haoran Li, Yujin Zhou, Chi-Min Chan, Pengcheng Wen, Lei Li, Sirui Han, Yike Guo

机构 * Hong Kong University of Science and Technology(香港科技大学) Zhongshan School of Medicine, SUN YAT-SEN UNIVERSITY(中山医学院,孙中山大学) University of Edinburgh(爱丁堡大学) University of Washington(华盛顿大学)

AI总结 AM$^3$Safety通过结合冷启动拒绝阶段和组相对策略优化,有效提升多模态多轮对话的安全性,降低攻击成功率并增强模型的无害与帮助维度。

详情

展开后加载摘要…

URL PDF HTML 收藏