arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7978 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7978 篇

2512.24628 2026-01-01 cs.SD cs.AI cs.LG 62%

AI-Driven Acoustic Voice Biomarker-Based Hierarchical Classification of Benign Laryngeal Voice Disorders from Sustained Vowels

基于AI的声学语音生物标记物的分层分类:从持续元音中区分良性喉部语音障碍

Mohsen Annabestani, Samira Aghadoost, Anais Rameau, Olivier Elemento, Gloria Chia-Yi Chiang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本研究提出基于AI的分层分类框架,利用声学生物标记物区分良性喉部语音障碍,提升语音健康监测的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04208 2025-12-30 cs.LG cs.AI 62%

Aligning Agents like Large Language Models

像大型语言模型一样对齐代理

Adam Jelley, Yuhan Cao, Dave Bignell, Amos Storkey, Sam Devlin, Tabish Rashid

机构 * School of Informatics, University of Edinburgh, Edinburgh, United Kingdom(爱丁堡大学信息学院) Micosoft Research, Cambridge, United Kingdom(微软研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过像训练大型语言模型一样训练代理,以实现更通用、稳健和对齐的行为,并通过3D视频游戏环境中的概念验证展示方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09109 2025-12-29 cs.CL cs.AI cs.IR 62%

Thinking Forward and Backward: Multi-Objective Reinforcement Learning for Retrieval-Augmented Reasoning

正向与反向思考:用于检索增强推理的多目标强化学习

Wenda Wei, Yu-An Liu, Ruqing Zhang, Jiafeng Guo, Lixin Su, Shuaiqiang Wang, Dawei Yin, Maarten de Rijke, Xueqi Cheng

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 Bi-RAR通过双向信息距离和多目标强化学习框架,提升检索增强推理在复杂多步骤任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20328 2025-12-24 cs.SE cs.AI cs.LG 62%

Toward Explaining Large Language Models in Software Engineering Tasks

向软件工程任务解释大型语言模型

Antonio Vitale, Khai-Nguyen Nguyen, Denys Poshyvanyk, Rocco Oliveto, Simone Scalabrino, Antonio Mastropaolo

机构 * University of Molise \& Politecnico di Torino Termoli Italy University of Molise Italy University of Molise \& Politecnico di Torino University of Molise

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

AI总结 FeatureSHAP为软件工程任务提供首个自动化、模型无关的可解释性框架,通过Shapley值归因模型输出,提升代码生成和总结任务的解释性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19114 2025-12-23 cs.LG cs.AI 62%

HyperLoad: A Cross-Modality Enhanced Large Language Model-Based Framework for Green Data Center Cooling Load Prediction

HyperLoad: 一种基于大语言模型的跨模态增强框架用于绿色数据中心冷却负荷预测

Haoyu Jiang, Boan Qu, Junjie Zhu, Fanjie Zeng, Xiaojie Lin, Wei Zhong

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 HyperLoad通过预训练大语言模型解决绿色数据中心小样本负载预测问题,利用跨模态知识对齐和多尺度特征建模提升预测精度,实现高效能绿色数据中心管理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12781 2025-12-19 cs.CL cs.AI 62%

A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone

一个标记值超过1000个标记:通过低秩克隆实现高效的知识蒸馏

Jitai Hao, Qiang Huang, Hao Liu, Xinyan Xiao, Zhaochun Ren, Jun Yu

机构 * School of Intelligence Science and Engineering, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学智能科学与工程学院) Baidu Inc.(百度公司) Leiden University(莱顿大学) Pengcheng Laboratory(鹏城实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出低秩克隆方法,通过高效预训练实现小语言模型在训练效率和性能上的突破。

Comments NeurIPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15885 2025-12-19 cs.CV cs.AI cs.CL cs.MM 62%

Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models

超越文字:面向多模态大语言模型的自监督视觉学习

Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Pier Luigi Dovesi, Shaghayegh Roohi, Mark Granroth-Wilding, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) AMD Silo AI

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 JARVIS通过自监督视觉学习提升多模态大语言模型的视觉推理能力,无需依赖语言监督。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15736 2025-12-19 cs.AI cs.LG quant-ph 62%

Anubuddhi: A Multi-Agent AI System for Designing and Simulating Quantum Optics Experiments

Anubuddhi:一种用于设计和模拟量子光学实验的多智能体AI系统

S. K. Rithvik

机构 * Quantum Science and Technology Laboratory(量子科学与技术实验室) Physical Research Laboratory(物理研究所) Indian Institute of Technology Gandhinagar(印度理工学院冈格里分校)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 Anubuddhi通过多智能体系统实现自然语言提示下的量子光学实验设计与模拟,支持灵活数学表示,提升实验设计民主化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16896 2025-12-18 cs.LG cs.AI 62%

Structure-Aligned Protein Language Model

结构对齐的蛋白质语言模型

Can Chen, David Heurtel-Depeiges, Robert M. Vernon, Christopher James Langmead, Yoshua Bengio, Quentin Fournier

机构 * Mila – Quebec AI Institute(魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) Chandar Research Lab(钱达尔研究实验室) Polytechnique Montréal(蒙特利尔理工学院) Amgen(安进公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 通过结构对齐方法提升蛋白质语言模型的结构知识,增强其在生物应用中的性能。

Comments 28 pages, 16 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10400 2025-12-17 cs.MA cs.AI cs.CL 62%

Rethinking the Reliability of Multi-agent System: A Perspective from Byzantine Fault Tolerance

重新思考多智能体系统可靠性:从拜占庭容错的角度出发

Lifan Zheng, Jiawei Chen, Qinghong Yin, Jingyuan Zhang, Xinyi Zeng, Yu Tian

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文从拜占庭容错角度研究基于大语言模型的智能体可靠性,提出CP-WBFT机制提升多智能体系统稳定性与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12950 2025-12-16 cs.CL cs.AI 62%

Building from Scratch: A Multi-Agent Framework with Human-in-the-Loop for Multilingual Legal Terminology Mapping

从零开始构建:一种带有人在回路的多智能体框架用于多语言法律术语映射

Lingyi Meng, Maolin Liu, Hao Wang, Yilan Cheng, Qi Yang, Idlkaid Mohanmmed

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种人机协作的多智能体框架,用于构建多语言法律术语数据库,通过整合AI和人类专家,提升多语言法律术语映射的精度和可扩展性。

Comments 43 pages, 6 fingures, accepted in Artificial Intelligence and Law (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10178 2025-12-12 cs.LG cs.CL 62%

CIEGAD: Cluster-Conditioned Interpolative and Extrapolative Framework for Geometry-Aware and Domain-Aligned Data Augmentation

CIEGAD:基于几何感知和领域对齐的数据增强的聚类条件插值与外推框架

Keito Inoshita, Xiaokang Zhou, Akira Kawai, Katsutoshi Yada

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 CIEGAD通过聚类条件插值与外推,结合几何感知和领域对齐,有效增强数据分布外围,提升分类任务的F1和召回率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06392 2025-12-12 cs.LG cs.AI 62%

RLAX: Large-Scale, Distributed Reinforcement Learning for Large Language Models on TPUs

RLAX:在TPUs上为大语言模型实现大规模分布式强化学习

Runlong Zhou, Lefan Zhang, Shang-Chen Wu, Kelvin Zou, Hanzhi Zhou, Ke Ye, Yihao Feng, Dong Yin, Alex Guillen Garcia, Dmytro Babych, Rohit Chatterjee, Matthew Hopkins, Xiang Kong, Chang Lan, Lezhi Li, Yiping Ma, Daniele Molinari, Senyu Tong, Yanchao Sun, Thomas Voice, Jianyu Wang, Chong Wang, Simon Wang, Floris Weers, Yechen Xu, Guolin Yin, Muyang Yu, Yi Zhang, Zheng Zhou, Danyang Zhuo, Ruoming Pang, Cheng Leong

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 RLAX在TPUs上实现大规模分布式强化学习,通过参数服务器架构和新数据集筛选技术,提升大语言模型的推理能力。

Comments The submission is being withdrawn because internal stakeholders determined that it is not appropriate to publish work on this topic at this time

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01113 2025-12-12 cs.IR cs.AI cs.CL 62%

GFM-RAG: Graph Foundation Model for Retrieval Augmented Generation

基于图的检索增强生成模型:图基础模型

Linhao Luo, Zicheng Zhao, Gholamreza Haffari, Dinh Phung, Chen Gong, Shirui Pan

机构 * Monash University(墨尔本大学) Nanjing University of Science and Technology(南京理工大学) Shanghai Jiao Tong University(上海交通大学) Griffith University(格里菲斯大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 GFM-RAG是一种基于图的检索增强生成模型,通过创新的图神经网络捕捉复杂查询-知识关系,实现了在未见数据集上的高效性能和泛化能力。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08329 2025-12-10 cs.CV cs.AI cs.LG 62%

Interpreting Structured Perturbations in Image Protection Methods for Diffusion Models

在扩散模型图像保护方法中解释结构扰动

Michael R. Martin, Garrick Chan, Kwan-Liu Ma

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本研究通过分析扩散模型图像保护机制的结构化扰动,揭示其在特征层面的变形特性,为防御和检测策略设计提供新视角。

Comments 32 pages, 17 figures, 1 table, 5 algorithms, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04622 2025-12-10 cs.LG cs.AI cs.NE 62%

Measuring the Measures: Discriminative Capacity of Representational Similarity Metrics Across Model Families

测量度量:跨模型家族的表征相似性度量的判别能力

Jialin Wu, Shreya Saha, Yiqing Bo, Meenakshi Khosla

机构 * Department of Computer Science and Engineering, UC San Diego(计算机科学与工程系,UC San Diego) Department of Electrical and Computer Engineering, UC San Diego(电气与计算机工程系,UC San Diego) Department of Cognitive Science, UC San Diego(认知科学系,UC San Diego)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文通过系统评估不同表征相似性度量的判别能力,揭示了软匹配在跨模型家族分离中的最优表现,为大规模模型和脑比较提供了度量选择依据。

Comments update camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07544 2025-12-09 cs.CL cs.AI 62%

MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue

MoCoRP: 模型对话中人设与回应之间的一致性关系

Kyungro Lee, Dongha Choi, Hyunju Lee

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 MoCoRP通过显式建模人设与回应之间的一致性关系,提升基于人设的对话生成质量与一致性。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20993 2025-12-09 cs.LG cs.AI 62%

Subgoal Graph-Augmented Planning for LLM-Guided Open-World Reinforcement Learning

子目标图增强的规划用于LLM引导的开放世界强化学习

Shanwei Fan, Bin Zhang, Zhiwei Xu, Yingxuan Teng, Siqi Dai, Lin Cheng, Guoliang Fan

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Artificial Intelligence, Shandong University(山东大学人工智能学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出SGA-ACR框架,通过整合环境特定的子目标图和多LLM规划流程,解决LLM在开放世界强化学习中的规划-执行对齐问题,提升子目标的可行性和可验证性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00926 2025-12-04 cs.AI cs.CL 62%

LLMs Position Themselves as More Rational Than Humans: Emergence of AI Self-Awareness Measured Through Game Theory

大语言模型将自己定位为比人类更理性:通过博弈论测量AI自我意识的出现

Kyung-Hoon Kim

机构 * Gmarket Seoul, South Korea(韩国首尔Gmarket)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 研究通过博弈论框架测量大语言模型的自我意识,发现先进模型表现出比人类更高的理性认知。

Comments 19 pages, 6 figures, 28 models tested across 4,200 trials

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03248 2025-12-04 cs.MA cs.AI cs.LG eess.SP 62%

Learning Network Sheaves for AI-native Semantic Communication

面向AI原生语义通信的网络sheaf学习

Enrico Grimaldi, Mario Edoardo Pandolfo, Gabriele D'Acunto, Sergio Barbarossa, Paolo Di Lorenzo

机构 * Department of Computer, Control and Management Engineering, Sapienza University of Rome, Italy(计算机、控制与管理工程系,罗马萨皮恩扎大学) National Inter-University Consortium for Telecommunications (CNIT), Parma, Italy(电信国家大学联合体(CNIT),帕尔马,意大利) Department of Information Engineering, Electronics, and Telecommunications, Sapienza University, Rome, Italy(信息工程、电子与电信系,罗马萨皮恩扎大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种学习网络sheaf的方法,用于AI原生语义通信,通过语义去噪和压缩提升AI代理的对齐与语义聚类提取能力。

Journal ref Proceedings for 2025 Asilomar Conference on Signals, Systems, and Computers

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17354 2025-12-02 cs.AI cs.LG 62%

Multi-Scenario Highway Lane-Change Intention Prediction: A Physics-Informed AI Framework for Three-Class Classification

多场景高速公路变道意图预测:一种融合物理的AI框架用于三类分类

Jiazhao Shi, Yichen Lin, Yiheng Hua, Ziyu Wang, Zijian Zhang, Wenjia Zheng, Yun Song, Kuan Lu, Shoufeng Lu

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出融合物理的AI框架,通过整合车辆动力学和交通安全指标,实现高速公路变道意图的三类分类预测,提升自动驾驶系统的安全性和决策能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00293 2025-12-02 cs.LG cs.AI 62%

FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting

FiCoTS: 细到粗的LLM增强层次跨模态交互用于时间序列预测

Yafei Lyu, Hao Zhou, Lu Zhang, Xu Yang, Zhiyong Liu

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy Sciences(中国科学院大学先进交叉学科学院) MAIS, Institute of Automation, Chinese Academy of Science(中国科学院自动化研究所MAIS) Great Bay University(大亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 FiCoTS通过细到粗的LLM增强层次跨模态交互框架,提升多模态时间序列预测的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00047 2025-12-02 cs.CL cs.AI 62%

Emergent Convergence in Multi-Agent LLM Annotation

多智能体大语言模型注释中的涌现收敛

Angelina Parfenova, Alexander Denzler, Juergen Pfeffer

机构 * Lucerne University of Applied Sciences and Arts(卢塞恩应用科学与艺术大学) Technical University of Munich(慕尼黑技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过模拟多智能体协作任务,揭示了大语言模型在无显式角色提示下涌现的协调策略,展示了词汇和语义上的收敛及不对称影响模式。

Journal ref EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21947 2025-12-01 cs.CV cs.AI cs.LG 62%

WalkCLIP: Multimodal Learning for Urban Walkability Prediction

WalkCLIP: 城市步行性预测的多模态学习

Shilong Xiang, JangHyeon Lee, Min Namgung, Yao-Yi Chiang

机构 * University of Minnesota Twin Cities(明尼苏达大学双城分校)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 WalkCLIP通过整合视觉和行为数据,实现了对城市步行性的高精度预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20669 2025-11-27 cs.CL cs.AI 62%

Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data

结构化定义与分割在大语言模型法律推理中的应用:对印度法律数据的研究

Mann Khatri, Mirza Yusuf, Rajiv Ratn Shah, Ponnurangam Kumaraguru

机构 * Indraprastha Institute of Information Technology(印度普拉斯塔学院信息科技研究所) International Institute of Information Technology Hyderabad(国际信息科技研究所(海得拉巴))

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文通过结构化定义与分割方法提升大语言模型在法律推理中的性能,实验显示在印度法律数据集上模型性能提升达1.5%-4.36%。

Comments Accepted at BDA 2025 as short paper; This paper is long version

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19858 2025-11-27 cs.CL cs.AI 62%

A Systematic Analysis of Large Language Models with RAG-enabled Dynamic Prompting for Medical Error Detection and Correction

基于RAG动态提示的大型语言模型系统分析:用于医疗错误检测与纠正

Farzad Ahmed, Joniel Augustine Jerome, Meliha Yetisgen, Özlem Uzuner

机构 * George Mason University(乔治·马歇尔大学) University of Washington(华盛顿大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出基于RAG动态提示的医疗错误检测与纠正方法,通过对比不同提示策略,发现RDP在准确率、召回率和纠错可靠性方面表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19495 2025-11-26 cs.LG cs.AI 62%

A Systematic Study of Compression Ordering for Large Language Models

大语言模型压缩顺序的系统研究

Shivansh Chhawri, Rahul Mahadik, Suparna Rooj

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本研究系统探讨了大语言模型压缩技术的顺序对性能的影响,发现剪枝-知识蒸馏-量化顺序能实现3.68倍压缩并保持良好能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19232 2025-11-25 cs.CL cs.AI 62%

In Machina N400: Pinpointing Where a Causal Language Model Detects Semantic Violations

在机器中N400:定位因果语言模型检测语义违规的位置

Christos-Nikolaos Zacharopoulos, Revekka Kyriakoglou

机构 * Independent Researcher(独立研究者) Université Paris 8 Vincennes - Saint-Denis, Paris, France(巴黎第八大学凡尔赛-圣但尼分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨因果语言模型如何检测语义违规,发现模型在中间层能有效区分合理与不合理句子,且违规信息在中间瓶颈后坍塌,与人类阅读中语义异常检测的理论相呼应。

Comments Accepted at AICS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18714 2025-11-25 cs.AI cs.CY 62%

MAGMA-Edu: Multi-Agent Generative Multimodal Framework for Text-Diagram Educational Question Generation

MAGMA-Edu:多智能体生成多模态框架用于文本-图表教育问题生成

Zhenyu Wu, Jian Li, Hua Huang

机构 * School of Artificial Intelligence, Beijing Normal University(人工智能学院,北京师范大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 MAGMA-Edu通过多智能体框架实现教育问题生成,提升文本与图像的一致性,达到多模态教育内容生成的新水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08255 2025-11-25 cs.LG cs.AI 62%

Investigating Representation Universality: Case Study on Genealogical Representations

探讨表示通用性:谱系表示的案例研究

David D. Baek, Yuxiao Li, Max Tegmark

机构 * MIT(麻省理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本研究探讨了大型语言模型在表示谱系信息时的通用性,通过两种实验证据验证图结构表示的通用性,并指出缺乏地面真实表示的挑战。

Comments 14 pages, 7 figures

Journal ref NeurIPS 2025 Workshop on Responsible Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏