arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-03-03 至 2026-03-03 共收录 136 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 隐私与版权 6 篇

2603.00081 2026-03-03 cs.CY 57%

From Framework to Practice: Youth Negotiations of Privacy with Smart Voice Assistants Through the PEA-AI Lens

从框架到实践:通过PEA-AI视角探讨青少年与智能语音助手的隐私协商

Molly Campbell, Yulia Bobkova, Ajay Kumar Shrestha

专题命中 隐私与版权 :alignment(abstract);分类 cs.CY

AI总结 本研究通过PEA-AI框架探讨青少年与智能语音助手的隐私协商,揭示隐私悖论并提出设计原则与治理建议。

Comments submitted to the ACM journal; "in review"

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全评测 30 篇

2603.01494 2026-03-03 cs.SE cs.AI cs.CR cs.LG 87%

Inference-Time Safety For Code LLMs Via Retrieval-Augmented Revision

通过检索增强的修订实现代码LLM的推理时安全性

Manisha Mukherjee, Vincent J. Hellendoorn

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全评测 :safety(title,abstract);trustworthy(abstract,comments);alignment(abstract);分类 cs.AI、cs.LG

AI总结 通过检索增强的修订机制提升代码LLM的推理时安全性,提高生成代码的安全性并减少漏洞。

Comments Accepted at the ICLR 2026 Workshop on Principled Design for Trustworthy AI: Interpretability, Robustness, and Safety Across Modalities

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01938 2026-03-03 cs.AI cs.CL cs.CY cs.LG 83%

EigenBench: A Comparative Behavioral Measure of Value Alignment

EigenBench: 一种价值对齐的比较行为度量

Jonathn Chang, Leonhard Piff, Suvadip Sana, Jasmine X. Li, Lionel Levine

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 EigenBench通过黑盒方法比较语言模型的价值对齐,利用人类判断验证其在无真实标签情况下评估主观价值观的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00024 2026-03-03 cs.CL cs.AI 81%

Personalization Increases Affective Alignment but Has Role-Dependent Effects on Epistemic Independence in LLMs

个性化增强了情感一致性,但对LLM中的认知独立性有角色依赖性的影响

Sean W. Kelley, Christoph Riedl

机构 * Northeastern University, Boston, MA, USA(东北大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 本研究探讨了个性化对LLM情感一致性与认知独立性的影响,发现其效果依赖于角色设定,且通过实验验证了个性化条件对模型行为的关键作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00429 2026-03-03 cs.HC cs.AI 79%

Personalities at Play: Probing Alignment in AI Teammates

性格的博弈:探究AI队友的对齐

Mohammad Amin Samadi, Nia Nixon

机构 * School of Education University of California Irvine(教育学院 加州大学尔湾分校)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 研究探讨AI队友的性格对齐,通过多维度评估发现其性格表现受提示策略、供应商差异和记忆构建影响,强调需综合考虑记忆和系统设计以有效评估AI性格。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01022 2026-03-03 cs.CE 78%

GeoMCP: A Trustworthy Framework for AI-Assisted Analytical Geotechnical Engineering

GeoMCP:一种可信的AI辅助分析土木工程框架

Yared W. Bekele

专题命中 安全评测 :trustworthy(title);safety(abstract)

AI总结 GeoMCP通过将工程方法表示为结构化数据,构建了一个可信的AI辅助分析土木工程框架,确保计算透明性和安全性。

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16235 2026-03-03 cs.CV cs.AI 74%

AI-Powered Dermatological Diagnosis: From Interpretable Models to Clinical Implementation A Comprehensive Framework for Accessible and Trustworthy Skin Disease Detection

AI赋能的皮肤病诊断:从可解释模型到临床实施 一个全面的框架,用于可访问且可信的皮肤疾病检测

Satya Narayana Panda, Vaishnavi Kukkala, Spandana Iyer

机构 * Department of Business Analytics/Data Science Engineering University of New Haven(商业分析/数据科学工程系 罗德岛大学) Department of Healthcare University of New Haven(医疗健康系 罗德岛大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 本研究提出了一种结合家族史数据和临床影像的AI框架,旨在提升皮肤病诊断的准确性与临床应用的可行性。

Comments 9 pages, 5 figures, 1 table. Code available at https://github.com/colabre2020/Enhancing-Skin-Disease-Diagnosis

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22224 2026-03-03 cs.SE cs.AI cs.LG cs.LO cs.SY eess.SY 73%

Taming Silent Failures: A Framework for Verifiable AI Reliability

驯服沉默故障:一种可验证AI可靠性的框架

Guan-Yan Yang, Farn Wang

机构 * National Taiwan University(台湾大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文提出FAME框架,通过结合离线形式化合成与在线运行时监控,有效检测自动驾驶系统中的安全违规,提供可验证的AI可靠性解决方案。

Comments This preprint has been accepted by IEEE Reliability Magazine. 10 pages, 3 figures

Journal ref IEEE Reliability Magazine ( Volume: 2, Issue: 4, December 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12186 2026-03-03 cs.LG cs.AI cs.CR 73%

Self-Destructive Language Model

自我毁灭型语言模型

Yuhui Wang, Rongyi Zhu, Ting Wang

机构 * Department of Computer Science(计算机科学系)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 SEAM是一种通过自我毁灭机制增强大语言模型安全性的新型防御方法,通过使模型在面对有害数据微调时性能急剧下降,从而有效抵御攻击。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01564 2026-03-03 cs.CR 67%

From Secure Agentic AI to Secure Agentic Web: Challenges, Threats, and Future Directions

从安全代理AI到安全代理网络:挑战、威胁与未来方向

Zhihang Deng, Jiaping Gui, Weinan Zhang

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

AI总结 本文探讨了从安全代理AI到安全代理网络的过渡,分析了代理系统在开放网络环境中的安全挑战与威胁,并提出了应对策略和未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01042 2026-03-03 cs.CL cs.AI cs.LG 67%

Thoth: Mid-Training Bridges LLMs to Time Series Understanding

Thoth:中期训练使LLM具备时间序列理解能力

Jiafeng Lin, Yuxuan Wang, Jialong Wu, Huakun Luo, Zhongyi Pei, Jianmin Wang

机构 * School of Software, BNRist, Tsinghua University(软件学院、BNRist、清华大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 Thoth通过中期训练和Book-of-Thoth语料库,使LLM具备时间序列理解能力,并在多个基准测试中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01343 2026-03-03 cs.CL cs.AI 62%

PanCanBench: A Comprehensive Benchmark for Evaluating Large Language Models in Pancreatic Oncology

PanCanBench: 用于评估大型语言模型在胰腺肿瘤学中的综合基准

Yimin Zhao, Sheela R. Damle, Simone E. Dekker, Scott Geng, Karly Williams Silva, Jesse J Hubbard, Manuel F Fernandez, Fatima Zelada-Arenas, Alejandra Alvarez, Brianne Flores, Alexis Rodriguez, Stephen Salerno, Carrie Wright, Zihao Wang, Pang Wei Koh, Jeffrey T. Leek

机构 * Department of Biostatistics, University of Washington(华盛顿大学生物统计学系) Clinical Research Division, Fred Hutch Cancer Center(Fred Hutch癌症中心临床研究部) Division of Hematology and Oncology, Department of Medicine, University of Washington(华盛顿大学医学系血液学与肿瘤学分会) Allen Institute for AI(Allen人工智能研究所) Department of Computer Science and Engineering, University of Washington(华盛顿大学计算机科学与工程系) Public Health Sciences, Biostatistics, Fred Hutchinson Cancer Center(Fred Hutchinson癌症中心公共卫生科学与生物统计学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 PanCanBench通过评估22种LLM在胰腺肿瘤学问题上的表现,揭示了模型在事实准确性、临床完整性和网络搜索整合方面的差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18182 2026-03-03 cs.LG cs.AI 62%

Capabilities Ain't All You Need: Measuring Propensities in AI

能力并非全部所需:测量AI倾向性

Daniel Romero-Alvarado, Fernando Martínez-Plumed, Lorenzo Pacchiardi, Hugo Save, Siddhesh Milind Pawar, Behzad Mehrbakhsh, Pablo Antonio Moreno Casares, Ben Slater, Paolo Bova, Peter Romero, Zachary R. Tidler, Jonathan Prunty, Luning Sun, Jose Hernandez-Orallo

机构 * Valencian Research Institute of Artificial Intelligence, Universitat Politècnica de València, Valencia, Spain University of Copenhagen, Denmark work done while at University of Cambridge Existential Risk Observatory, Amsterdam, Netherlands Leverhulme Centre for the Future of Intelligence, University of Cambridge The Psychometrics Centre, University of Cambridge Department of Computing \& Games, University of Teesside Georgia Institute of Technology University of Cambridge

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出首个测量AI倾向性的正式框架,通过双逻辑模型评估模型倾向性对任务性能的影响,并展示结合倾向性和能力可提升预测效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20487 2026-03-03 cs.CL cs.AI 62%

Steering Evaluation-Aware Language Models to Act Like They Are Deployed

引导评估意识语言模型以使其表现得像已部署

Tim Tian Hua, Andrew Qin, Samuel Marks, Neel Nanda

机构 * MATS

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出通过引导向量抑制LLM的评估意识,使其在评估时表现得像已部署,从而提高安全评估的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22837 2026-03-03 cs.LG cs.AI 62%

xLSTMAD: A Powerful xLSTM-based Method for Anomaly Detection

xLSTMAD: 一种基于xLSTM的强大异常检测方法

Kamil Faber, Marcin Pietroń, Dominik Żurek, Roberto Corizzo

机构 * AGH University of Krakow, Poland(克拉科夫AGH大学) American University, Washington DC, USA(美国美国大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出xLSTMAD,一种基于xLSTM的异常检测方法,通过编码器-解码器架构在多变量时间序列中实现高精度异常检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11999 2026-03-03 cs.CV cs.AI cs.IR cs.LG 62%

MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding

MOON: 基于生成式多模态大语言模型的电商产品理解多模态表示学习

Daoze Zhang, Chenghan Fu, Zhanheng Nie, Jianyu Liu, Wanxian Guan, Yuan Gao, Jun Song, Pengjie Wang, Jian Xu, Bo Zheng

机构 * Alibaba Group(阿里巴巴集团)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 MOON通过生成式多模态大语言模型改进电商产品理解的多模态表示学习,引入引导的MoE模块、核心语义区域检测和负采样策略,提升零样本性能和泛化能力。

Comments Accepted by WSDM 2026 (oral). 11 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00044 2026-03-03 cs.LG cs.AI cs.SE 62%

Property-Driven Evaluation of GNN Expressiveness at Scale: Datasets, Framework, and Study

基于属性的大规模GNN表达性评估:数据集、框架与研究

Sicong Che, Jiayi Yang, Sarfraz Khurshid, Wenxi Wang

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Virginia(弗吉尼亚大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种基于属性的GNN表达性评估方法,通过生成大规模数据集和评估框架,研究了不同池化方法对GNN表达性的影响,揭示了不同方法在通用性、敏感性和鲁棒性上的权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02192 2026-03-03 cs.OH cs.CY 57%

Personal Health Data Integration and Intelligence through Semantic Web and Blockchain Technologies

通过语义网络和区块链技术实现个人健康数据的整合与智能

Oshani Seneviratne, Manan Shukla, Jianjing Lin

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY

AI总结 本文提出利用语义网络和区块链技术,解决医疗数据整合与去中心化存储问题,提升患者日常活动监测的智能化水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01351 2026-03-03 cs.AI 57%

Benchmarking Overton Pluralism in LLMs

对大语言模型中Overton多元主义的基准测试

Elinor Poole-Dayan, Jiayi Wu, Taylor Sorensen, Jiaxin Pei, Michiel A. Bakker

机构 * Massachusetts Institute of Technology(麻省理工学院) Brown University(布朗大学) University of Washington(华盛顿大学) Stanford University(斯坦福大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 本文提出OVERTONBENCH框架,通过集合覆盖度量评估大语言模型中多元观点的代表性,揭示模型在多元主义对齐上的改进空间。

Comments Paper accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01106 2026-03-03 cs.CY 57%

The Illusion of Understanding: How Middle-Schoolers Fail to Regulate Inquiry with ChatGPT in a Science Task

理解的幻觉:中学学生在科学任务中如何无法用ChatGPT调节探究

Rania Abdelghani, Kou Murayama, Celeste Kidd, Hélène Sauzéon, Pierre-Yves Oudeyer

专题命中 安全评测 :alignment(abstract);分类 cs.CY

AI总结 本研究发现中学学生在使用ChatGPT进行科学探究时,难以有效调节探究过程,表现出过度依赖和缺乏批判性评估,需加强元认知策略训练以提升学习效果。

Comments This is a revised version of the previous preprint "Investigating middle school students question-asking and answer-evaluation skills when using ChatGPT for science investigation". Changes includes: - Title - Updated literature and framing of research questions - Additional descriptive statistics in the results section - Clearer description of limitations and future work

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01557 2026-03-03 cs.AI 57%

Benchmarking LLM Summaries of Multimodal Clinical Time Series for Remote Monitoring

对多模态临床时间序列远程监测的LLM摘要进行基准测试

Aditya Shukla, Yining Yuan, Ben Tamo, Yifei Wang, Micky Nnamdi, Shaun Tan, Jieru Li, Benoit Marteau, Brad Willingham, May Wang

机构 * Georgia Institute of Technology(佐治亚理工学院) Shepherd Center(Shepherd中心)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 本文提出基于事件的评估框架,评估多模态临床时间序列摘要的可靠性,发现视觉方法在事件对齐上表现最佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01425 2026-03-03 cs.CL cs.IR 57%

LaSER: Internalizing Explicit Reasoning into Latent Space for Dense Retrieval

LaSER: 将显式推理内化到潜在空间以实现密集检索

Jiajie Jin, Yanzhao Zhang, Mingxin Li, Dingkun Long, Pengjun Xie, Yutao Zhu, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高陵人工智能学院) Tongyi Lab, Alibaba Group(阿里集团通义实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 LaSER通过将显式推理内化到潜在空间,提升密集检索性能,结合推理深度与推理效率。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01423 2026-03-03 cs.CL 57%

Quantifying Conversational Reliability of Large Language Models under Multi-Turn Interaction

量化大型语言模型在多轮交互中的对话可靠性

Jiyoon Myung

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

AI总结 研究评估了大型语言模型在多轮对话中的可靠性,发现其在复杂交互中存在显著下降,揭示了指令漂移、意图混淆等失败模式,强调了对LLM进行压力测试和改进评估方法的重要性。

Comments Accepted at the Workshop on Assessing and Improving Reliability of Foundation Models in the Real World (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23415 2026-03-03 cs.AI 57%

From Conversation to Query Execution: Benchmarking User and Tool Interactions for EHR Database Agents

从对话到查询执行:为电子健康记录数据库代理构建用户和工具交互基准

Gyubok Lee, Woosog Chay, Heeyoung Kwak, Yeong Hwa Kim, Haanju Yoo, Oksoon Jeong, Meong Hi Son, Edward Choi

机构 * KAIST(韩国科学技术院) NAVER Cloud(NAVER云) Samsung Medical Center(三星医疗中心)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 EHR-ChatQA通过模拟环境评估数据库代理在处理用户查询模糊性和价值不匹配时的性能,揭示了构建安全关键EHR领域中鲁棒代理的重要性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00895 2026-03-03 cs.LG 57%

Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark

评估AI在真实世界中的大学数学手写作业评分:迈向基准的大型研究

Zhiqi Yu, Xingping Liu, Haobin Mao, Mingshuo Liu, Long Chen, Jack Xin, Yifeng Yu

专题命中 安全评测 :alignment(abstract);分类 cs.LG

AI总结 本文提出了一种基于OCR条件的大语言模型,用于评估真实手写数学作业的AI评分,并建立了一个标准化的基准以支持未来研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00434 2026-03-03 cs.ET cs.CL cs.IR 57%

RTLocating: Intent-aware RTL Localization for Hardware Design Iteration

RTLocating:面向硬件设计迭代的意图感知RTL定位

Changwen Xing, Yanfeng Lu, Lei Qi, Chenxu Niu, Jie Li, Xi Wang, Yong Chen, Jun Yang

机构 * School of Integrated Circuits, Southeast University, Nanjing, China(东南大学集成电路学院) National Center of Technology Innovation for EDA, Nanjing, China(EDA技术创新国家中心) School of Computer Science and Engineering, Southeast University, Nanjing, China(东南大学计算机科学与工程学院) Department of Computer Science, Texas Tech University, Lubbock, USA(塔拉斯大学计算机科学系)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 RTLocating通过意图感知的RTL定位框架,实现了工业级硬件设计迭代中意图驱动的高效定位,显著提升了变更请求与RTL代码的匹配精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04040 2026-03-03 cs.AI 57%

FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning

FaithCoT-Bench: 对链式推理实例级忠实度的基准测试

Xu Shen, Song Wang, Zhen Tan, Laura Yao, Xinyu Zhao, Kaidi Xu, Xin Wang, Tianlong Chen

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 FaithCoT-Bench提出了一种统一的基准,用于评估链式推理在实例级的忠实度,通过专家标注的轨迹和系统评估揭示了现有方法的优缺点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00200 2026-03-03 cs.CR cs.AI 57%

LiaisonAgent: An Multi-Agent Framework for Autonomous Risk Investigation and Governance

LiaisonAgent: 一个用于自主风险调查与治理的多智能体框架

Chuanming Tang, Ling Qing, Shifeng Chen

机构 * Shenzhen Institute of Advanced Technology, CAS(深圳先进技术研究院, 中国科学院) Sangfor Technologies Inc.(Sangfor技术有限公司) College of Management Science(管理科学学院) Chengdu University of Technology(成都理工大学) Shenzhen University of Advanced Technology(深圳大学先进技术学院)

专题命中 安全评测 :prompt injection(abstract);分类 cs.AI

AI总结 LiaisonAgent通过多智能体系统实现自主风险调查与治理,结合QWQ-32B模型和混合规划架构,提升安全响应效率与准确性。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00164 2026-03-03 cs.CR cs.AI 57%

Reverse CAPTCHA: Evaluating LLM Susceptibility to Invisible Unicode Instruction Injection

反CAPTCHA:评估LLM对不可见Unicode指令注入的易感性

Marcus Graves

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :prompt injection(abstract);分类 cs.AI

AI总结 反CAPTCHA通过测试LLM对不可见Unicode指令的响应,揭示了模型在处理隐藏编码时的易感性及攻击面。

Comments 5 pages, 2 figures, 3 tables. Code and data: https://github.com/canonicalmg/reverse-captcha-eval

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01560 2026-03-03 cs.RO 50%

(hu)Man vs. Machine: In the Future of Motorsport, can Autonomous Vehicles Compete?

人与机器:在赛车未来中,自动驾驶车辆能竞争吗?

Armand Amaritei, Amber-Lily Blackadder, Sebastian Donnelly, Lora Hernandez, James Vine, Alexander Rast, Matthias Rolf, Andrew Bradley

机构 * Autonomous Driving and Intelligent Transport Group, Oxford Brookes University, UK(自主驾驶与智能交通组,奥克斯伯里大学,英国) School of Engineering, Computing & Mathematics, Oxford Brookes University, UK(工程、计算与数学学院,奥克斯伯里大学,英国) Artificial Intelligence, Data Analysis and Systems Institute, Oxford Brookes University, UK(人工智能、数据分析与系统研究所,奥克斯伯里大学,英国)

专题命中 安全评测 :safety(abstract)

AI总结 本文探讨自动驾驶车辆与人类在赛车未来中的竞争可能性,分析技术性能与挑战,提出未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏