arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9311 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9311 篇

2512.14019 2025-12-17 cs.LG q-bio.QM 79%

EXAONE Path 2.5: Pathology Foundation Model with Multi-Omics Alignment

EXAONE Path 2.5:多组学对齐的病理基础模型

Juseung Yun, Sunwoo Yu, Sumin Ha, Jonghyun Kim, Janghyeon Lee, Jongseong Jang, Soonyoung Lee

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

AI总结 EXAONE Path 2.5通过多组学对齐构建病理基础模型,实现更全面的肿瘤生物学建模,展现高效率和适应性,推动精准肿瘤学发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11421 2025-12-15 cs.AI 79%

Towards Trustworthy Multi-Turn LLM Agents via Behavioral Guidance

通过行为指导实现可信的多轮大语言模型代理

Gonca Gürsun

机构 * Gonca Gürsun(独立研究者)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

AI总结 本文提出一种框架,通过行为指导使大语言模型代理在多轮任务中实现可靠和可验证的行为。

Comments Accepted to AAAI 2026 Workshop on Trust and Control in Agentic AI (TrustAgent)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10470 2025-12-15 cs.CL 79%

Statistical Analysis of Sentence Structures through ASCII, Lexical Alignment and PCA

通过ASCII、词性对齐和PCA进行句子结构的统计分析

Abhijeet Sahdev

机构 * New Jersey Institute of Technology(新泽西理工学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

AI总结 本研究通过ASCII、词性对齐和PCA分析句子结构的平衡,提出一种资源高效的统计方法,用于评估文本平衡并补充传统语法工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10304 2025-12-12 cs.AI cs.ET 79%

Trustworthy Orchestration Artificial Intelligence by the Ten Criteria with Control-Plane Governance

基于十项准则的可信 orchestration 人工智能与控制平面治理

Byeong Ho Kang, Wenli Yang, Muhammad Bilal Amin

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

AI总结 本文提出基于十项准则的可信 orchestration 人工智能框架,通过整合人类输入、语义一致性等要素,系统性地提升人工智能系统的可信度与可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19980 2025-12-12 cs.LG 79%

RAD: Towards Trustworthy Retrieval-Augmented Multi-modal Clinical Diagnosis

RAD:迈向可信的检索增强多模态临床诊断

Haolin Li, Tianjie Dai, Zhe Chen, Siyuan Du, Jiangchao Yao, Ya Zhang, Yanfeng Wang

机构 * College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) Shanghai AI Laboratory(上海人工智能实验室) CMIC, Shanghai Jiao Tong University(上海交通大学计算机学院) School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) Institute of Artificial Intelligence for Medicine, Shanghai Jiao Tong University(上海交通大学医学人工智能研究所)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

AI总结 RAD通过检索增强多模态模型,提升临床诊断的可信度与准确性,实现任务特定知识的显式注入。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15456 2025-12-12 cs.CL 79%

Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized Alignment

教导语言模型与用户共同进化:面向个性化对齐的动态资料模型

Weixiang Zhao, Xingyu Sui, Yulin Hu, Jiahe Guo, Haixiao Liu, Biye Li, Yanyan Zhao, Bing Qin, Ting Liu

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

AI总结 本研究提出RLPA框架,通过动态资料推断提升个性化对话性能,Qwen-RLPA在多个基准测试中超越现有方法。

Comments NeurIPS 2025 Camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11963 2025-12-11 cs.RO cs.LG 79%

Learning-based social coordination to improve safety and robustness of cooperative autonomous vehicles in mixed traffic

基于学习的社会协调以提高混合交通中协作自动驾驶车辆的安全性和鲁棒性

Rodolfo Valiente, Behrad Toghi, Mahdi Razzaghpour, Ramtin Pedarsani, Yaser P. Fallah

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

AI总结 本文提出基于多智能体强化学习的方法,通过量化自动驾驶车辆的社会偏好并引入利他主义,提升其在混合交通中与人类驾驶车辆协作的安全性和鲁棒性。

Comments arXiv admin note: substantial text overlap with arXiv:2202.00881

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06917 2025-12-09 cs.LG 79%

Know your Trajectory -- Trustworthy Reinforcement Learning deployment through Importance-Based Trajectory Analysis

了解你的轨迹 -- 通过基于重要性的轨迹分析实现可信的强化学习部署

Clifford F, Devika Jay, Abhishek Sarkar, Satheesh K Perepu, Santhosh G S, Kaushik Dey, Balaraman Ravindran

机构 * Clifford F(未知) Devika Jay(未知) Abhishek Sarkar(未知) Satheesh K Perepu(未知) Santhosh G S(未知) Kaushik Dey(未知) Balaraman Ravindran(未知)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

AI总结 本文提出基于重要性的轨迹分析框架,通过评估轨迹级状态重要性,提升强化学习部署的可信度和可解释性。

Comments Accepted at 4th Deployable AI Workshop at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01080 2025-12-02 cond-mat.mtrl-sci cs.LG 79%

Building Trustworthy AI for Materials Discovery: From Autonomous Laboratories to Z-scores

构建可信的人工智能用于材料发现:从自主实验室到Z分数

Benhour Amirian, Ashley S. Dale, Sergei Kalinin, Jason Hattrick-Simpers

机构 * University of Toronto(多伦多大学) University of Tennessee(田纳西大学) Vector Institute for Artificial Intelligence(人工智能矢量研究所) Schwartz Reisman Institute for Technology and Society(技术与社会斯瓦茨-雷曼研究所)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

AI总结 本文提出GIFTERS框架,用于评估材料发现中AI方法的信任度,强调信任原则和改进方法,以确保AI加速发现的同时符合科学规范。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21920 2025-12-01 cs.SE cs.AI 79%

Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code

迈向基于LLM生成代码的自动化和可信的科学分析与可视化

Apu Kumar Chakroborti, Yi Ding, Lipeng Wan

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

AI总结 本文探讨了LLM生成代码在科学分析与可视化中的可信度与自动化潜力,提出三种策略提升代码执行成功率与质量,强调需进一步优化以构建更可靠的AI辅助研究工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21799 2025-12-01 cs.LG 79%

The Double-Edged Nature of the Rashomon Set for Trustworthy Machine Learning

拉索蒙集在可信机器学习中的双刃剑性质

Ethan Hsu, Harry Chen, Chudi Zhong, Lesia Semenova

机构 * Duke University(杜克大学) MIT(麻省理工学院) UNC-Chapel Hill(北卡罗来纳大学教堂山分校) Rutgers University(罗格斯大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

AI总结 研究揭示了拉索蒙集在可信机器学习中的双刃剑性质,指出其在提升鲁棒性的同时也带来隐私风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21568 2025-11-27 cs.CL 79%

RoParQ: Paraphrase-Aware Alignment of Large Language Models Towards Robustness to Paraphrased Questions

RoParQ:面向抗 paraphrased 问题的大型语言模型对齐

Minjoon Choi

机构 * Seoul National University(首尔国立大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

AI总结 RoParQ 通过引入 XParaCon 指标和基于推理的 SFT 策略,提升 LLM 在 paraphrased 问题上的鲁棒性与一致性。

Comments 12 pages, 9 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21561 2025-11-27 cs.LG 79%

Machine Learning Approaches to Clinical Risk Prediction: Multi-Scale Temporal Alignment in Electronic Health Records

基于机器学习的临床风险预测方法:电子健康记录中的多尺度时间对齐

Wei-Chen Chang, Lu Dai, Ting Xu

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

AI总结 本研究提出基于多尺度时间对齐网络的临床风险预测方法,通过多尺度卷积和注意力机制提升EHR中异步时间序列的建模能力。

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17041 2025-11-27 cs.IR cs.AI 79%

CLLMRec: LLM-powered Cognitive-Aware Concept Recommendation via Semantic Alignment and Prerequisite Knowledge Distillation

CLLMRec: 基于大语言模型的认知感知概念推荐:通过语义对齐和先决知识蒸馏

Xiangrui Xiong, Yichuan Lu, Zifei Pan, Chang Sun

机构 * Xihua University, Chengdu 610039, China(西华大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 CLLMRec通过语义对齐和先决知识蒸馏,利用大语言模型实现认知感知的概念推荐,无需依赖结构化知识图谱。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18597 2025-11-25 cs.CL 79%

Toward Trustworthy Difficulty Assessments: Large Language Models as Judges in Programming and Synthetic Tasks

迈向可信的难度评估:大型语言模型作为编程和合成任务的裁判

H. M. Shadman Tabib, Jaber Ahmed Deedar

机构 * Department of Computer Science, Bangladesh University of Engineering and Technology(计算机科学系,孟加拉国工程与技术大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL

AI总结 本文研究了大型语言模型在评估编程问题难度时的可靠性,发现LightGBM在准确性上显著优于GPT-4o,揭示了模型在处理数字约束方面的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11500 2025-11-25 cs.LG 79%

Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation

诚实胜过准确:通过强化犹豫实现可信语言模型

Mohamad Amin Mohamadi, Tianhao Wang, Zhiyuan Li

机构 * Toyota Technological Institute at Chicago(芝加哥丰田技术研究所) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

AI总结 通过强化犹豫方法,将回避作为核心训练目标,使语言模型在风险可控下提升可信度,降低计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22300 2025-11-24 cs.CR cs.AI cs.CV 79%

T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model

T2I-RiskyPrompt:用于评估、攻击和防御文本到图像模型安全性的基准

Chenyu Zhang, Tairen Zhang, Lanjun Wang, Ruidong Chen, Wenhui Li, Anan Liu

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

AI总结 T2I-RiskyPrompt提出了一种用于评估文本到图像模型安全性的综合基准,通过层次化风险分类和原因驱动的检测方法,全面评估了多种模型和防御策略的安全性能。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01418 2025-11-20 cs.CL 79%

On the Alignment of Large Language Models with Global Human Opinion

Yang Liu, Masahiro Kaneko, Chenhui Chu

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments 28 pages, 26 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04759 2025-11-20 cs.AI 79%

Driving with Regulation: Trustworthy and Interpretable Decision-Making for Autonomous Driving with Retrieval-Augmented Reasoning

Tianhui Cai, Yifan Liu, Zewei Zhou, Haoxuan Ma, Seth Z. Zhao, Zhiwen Wu, Xu Han, Zhiyu Huang, Jiaqi Ma

专题命中 安全评测 :trustworthy(title);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02175 2025-11-19 cs.SD cs.CL eess.AS 79%

Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment through Latent Acoustic Pattern Triggers

Liang Lin, Miao Yu, Kaiwen Luo, Yibo Zhang, Lilan Peng, Dexian Wang, Xuehai Tang, Yuanhe Zhang, Xikang Yang, Zhenhong Zhou, Kun Wang, Yang Liu

专题命中 安全评测 :alignment(title);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10810 2025-11-17 cs.AI 79%

HARNESS: Human-Agent Risk Navigation and Event Safety System for Proactive Hazard Forecasting in High-Risk DOE Environments

Ran Elgedawy, Sanjay Das, Ethan Seefried, Gavin Wiggins, Ryan Burchfield, Dana Hewit, Sudarshan Srinivasan, Todd Thomas, Prasanna Balaprakash, Tirthankar Ghosal

机构 * Oak Ridge National Laboratory(奥克勒斯国家实验室)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09964 2025-11-14 cs.SE cs.AI cs.PL 79%

EnvTrace: Simulation-Based Semantic Evaluation of LLM Code via Execution Trace Alignment -- Demonstrated at Synchrotron Beamlines

Noah van der Vleuten, Anthony Flores, Shray Mathur, Max Rakitin, Thomas Hopkins, Kevin G. Yager, Esther H. R. Tsai

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07871 2025-11-14 cs.CL 79%

AlignSurvey: A Comprehensive Benchmark for Human Preferences Alignment in Social Surveys

Chenxi Lin, Weikang Yuan, Zhuoren Jiang, Biao Huang, Ruitao Zhang, Jianan Ge, Yueqian Xu, Jianxing Yu

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15434 2025-11-14 cs.CV cs.LG 79%

Semantic4Safety: Causal Insights from Zero-shot Street View Imagery Segmentation for Urban Road Safety

Huan Chen, Ting Han, Siyu Chen, Zhihao Guo, Yiping Chen, Meiliu Wu

机构 * School of Geospatial Engineering and Science, Sun Yat-sen University(地理空间工程与科学学院,中山大学) School of Geographical and Earth Sciences, University of Glasgow(地理与地球科学学院,格拉斯哥大学) School of Economics and Management, Shanxi University(经济学与管理学院,山西大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

Comments 11 pages, 10 figures, The 8th ACM SIGSPATIAL International Workshop on AI for Geographic Knowledge Discovery (GeoAI '25), November 3--6, 2025, Minneapolis, MN, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20810 2025-11-12 cs.CL 79%

Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enriching

Songze Li, Zhiqiang Liu, Zhengke Gui, Huajun Chen, Wen Zhang

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团) ZJU-Ant Group Joint Lab of Knowledge Graph(浙江大学-蚂蚁集团知识图谱联合实验室)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Journal ref EMNLP 2025, pages 7683-7703

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11881 2025-11-11 cs.CL 79%

Evaluating Human-LLM Representation Alignment: A Case Study on Affective Sentence Generation for Augmentative and Alternative Communication

Shadab Choudhury, Asha Kumar, Lara J. Martin

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments Published at IJCNLP-AACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04312 2025-11-07 cs.AI 79%

Probing the Probes: Methods and Metrics for Concept Alignment

Jacob Lysnæs-Larsen, Marte Eggen, Inga Strümke

机构 * Department of Computer Science NTNU - Norwegian University of Science and Technology(计算机科学系挪威科学技术大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments 29 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20621 2025-11-03 cs.AI 79%

Towards the Formalization of a Trustworthy AI for Mining Interpretable Models explOiting Sophisticated Algorithms

Riccardo Guidotti, Martina Cinquini, Marta Marchiori Manerba, Mattia Setzu, Francesco Spinnato

机构 * University of Pisa(比萨大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12667 2025-11-03 cs.AI cs.LO 79%

Building Trustworthy AI by Addressing its 16+2 Desiderata with Goal-Directed Commonsense Reasoning

Alexis R. Tudor, Yankai Zeng, Huaduo Wang, Joaquin Arias, Gopal Gupta

机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06497 2025-10-30 cs.CV cs.AI 79%

Evaluation of Safety Cognition Capability in Vision-Language Models for Autonomous Driving

Enming Zhang, Peizhe Gong, Xingyuan Dai, Min Huang, Yisheng Lv, Qinghai Miao

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏