arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9324 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9324 篇

2601.08331 2026-01-14 cs.CL 74%

CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark

CLaS-Bench: 一种跨语言对齐与引导基准

Daniil Gurgurov, Yusser Al Ghussin, Tanja Baeumel, Cheng-Ting Chou, Patrick Schramowski, Marius Mosbach, Josef van Genabith, Simon Ostermann

机构 * Saarland University(萨尔兰大学) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) Centre for European Research in Trusted AI (CERTAIN)(可信AI欧洲研究中心(CERTAIN)) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) TU Darmstadt(德累斯顿技术大学) Mila - Quebec Artificial Intelligence Institute(魁北克人工智能研究所(Mila)) McGill University(麦吉尔大学)

专题命中 安全评测 :alignment(title);分类 cs.CL

AI总结 CLaS-Bench是首个多语言引导基准,通过评估32种语言中的语言强制行为,验证了残差基于DiffMean方法在跨语言引导中的有效性。

Comments pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02364 2026-01-07 cs.IR cs.AI 74%

Towards Trustworthy LLM-Based Recommendation via Rationale Integration

通过推理整合实现可信的基于大语言模型的推荐

Chung Park, Taesan Kim, Hyeongjun Yun, Dongjoon Hong, Junui Hong, Kijung Park, MinCheol Cho, Mira Myong, Jihoon Oh, Min sung Choi

机构 * SK Telecom(SK电信)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 本文提出基于大语言模型的推荐系统,通过生成逻辑严谨的理由提升推荐的可解释性和性能。

Comments Accepted at RS4SD'25 (CIKM'25 Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24040 2026-01-01 cs.AI 74%

ROAD: Reflective Optimization via Automated Debugging for Zero-Shot Agent Alignment

ROAD: 通过自动调试实现零样本智能体对齐的反思优化

Natchaya Temyingyong, Daman Jain, Neeraj Kumarsahu, Prabhat Kumar, Rachata Phondi, Wachiravit Modecrua, Krittanon Kaewtawee, Krittin Pachtrachai, Touchapon Kraisingkorn

专题命中 安全评测 :alignment(title);分类 cs.AI

AI总结 ROAD通过自动调试实现零样本智能体对齐的反思优化,利用多智能体架构将无结构故障日志转化为结构化决策树协议,提升智能体性能和样本效率。

Comments 22 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20018 2025-12-12 cs.AI cs.AR 74%

Achieving Trustworthy Real-Time Decision Support Systems with Low-Latency Interpretable AI Models

实现低延迟可解释AI模型的可信实时决策支持系统

Zechun Deng, Ziwei Liu, Ziqian Bi, Junhao Song, Chia Xin Liang, Joe Yeong, Xinyuan Song, Junfeng Hao

机构 * School of Physics and Astronomy(物理与天文学系) University of Edinburgh(爱丁堡大学) Siebel School of Computing and Data Science(计算与数据科学学院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) AI Agent Lab(AI代理实验室) Vokram Group(Vokram集团) Department of Anatomical Pathology(解剖病理学部) Singapore General Hospital(新加坡中央医院) Department of Computer Science(计算机科学系) Emory University(埃默里大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 本文提出通过低延迟可解释AI模型实现可信实时决策支持系统,探讨资源受限环境下的人机协作方法及边缘计算优化策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02603 2025-12-11 eess.AS cs.CL cs.SD 74%

SEAL: Speech Embedding Alignment Learning for Speech Large Language Model with Retrieval-Augmented Generation

SEAL:用于带有检索增强生成的语音大语言模型的语音嵌入对齐学习

Chunyu Sun, Bingyu Liu, Zhichao Cui, Junhan Shi, Anbin Qi, Tian-hao Zhang, Dinghao Zhou, Lewei Lu

机构 * SenseTime Research(商汤科技研究院)

专题命中 安全评测 :alignment(title);分类 cs.CL

AI总结 SEAL通过统一的语音和文本嵌入框架,提升语音大语言模型的检索效率与准确性,减少延迟并增强多模态检索能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06902 2025-12-09 cs.SE cs.AI 74%

BabelCoder: Agentic Code Translation with Specification Alignment

BabelCoder: 基于规范对齐的代理代码翻译

Fazle Rabbi, Soumit Kanti Saha, Tri Minh Triet Pham, Song Wang, Jinqiu Yang

机构 * Concordia University(康科迪亚大学) York University(约克大学)

专题命中 安全评测 :alignment(title);分类 cs.AI

AI总结 BabelCoder通过代理协作框架提升代码翻译准确性,实现94.16%的平均准确率。

Comments 21 pages, 8 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00142 2025-12-02 cs.CR cs.AI q-fin.CP q-fin.GN 74%

DeFi TrustBoost: Blockchain and AI for Trustworthy Decentralized Financial Decisions

去中心化金融信任提升:区块链与人工智能用于可信的去中心化金融决策

Swati Sachan, Dale S. Fickett

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 该研究提出一种结合区块链与可解释人工智能的去中心化金融信任提升框架,旨在解决贷款审批中的信任问题,并通过防篡改审计和数据存储策略提升监管合规性。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08135 2025-11-27 cs.SE cs.AI cs.DC cs.PF 74%

Leveraging AI for Productive and Trustworthy HPC Software: Challenges and Research Directions

利用AI实现高效且可信的HPC软件:挑战与研究方向

Keita Teranishi, Harshitha Menon, William F. Godoy, Prasanna Balaprakash, David Bau, Tal Ben-Nun, Abhinav Bhatele, Franz Franchetti, Michael Franusich, Todd Gamblin, Giorgis Georgakoudis, Tom Goldstein, Arjun Guha, Steven Hahn, Costin Iancu, Zheming Jin, Terry Jones, Tze Meng Low, Het Mankad, Narasinga Rao Miniskar, Mohammad Alaul Haque Monil, Daniel Nichols, Konstantinos Parasyris, Swaroop Pophale, Pedro Valero-Lara, Jeffrey S. Vetter, Samuel Williams, Aaron Young

机构 * Oak Ridge National Laboratory(奥克荷厄斯国家实验室) Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室) Lawrence Berkeley National Laboratory(劳伦斯伯克利国家实验室) Carnegie Mellon University(卡内基梅隆大学) Northeastern University(东北大学) University of Maryland(马里兰大学) SpiralGen Inc.(SpiralGen公司)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 本文探讨了利用AI改进HPC软件的挑战与研究方向,提出通过Ellora和Durban项目推动AI在HPC软件发展中的应用。

Comments 12 pages, 1 Figure, Accepted at "The 1st International Workshop on Foundational Large Language Models Advances for HPC" LLM4HPC to be held in conjunction with ISC High Performance 2025

Journal ref In: Neuwirth, S., Paul, A.K., Weinzierl, T., Carson, E.C. (eds) High Performance Computing. ISC High Performance 2025. Lecture Notes in Computer Science, vol 16091. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01668 2025-11-18 cs.AI 74%

Hybrid Retrieval-Augmented Generation Agent for Trustworthy Legal Question Answering in Judicial Forensics

Yueqing Xi, Yifan Bai, Huasen Luo, Weiliang Wen, Hui Liu, Haoliang Li

机构 * Department of Electronic Engineering, City University of Hong Kong(DongGuan)(香港城市大学(东莞)电子工程系) Department of Electronic Engineering, City University of Hong Kong(香港城市大学电子工程系)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12014 2025-11-18 cs.CL cs.HC 74%

CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs

Truong Vo, Sanmi Koyejo

专题命中 安全评测 :alignment(title);分类 cs.CL

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09855 2025-11-14 cs.LG 74%

Unlearning Imperative: Securing Trustworthy and Responsible LLMs through Engineered Forgetting

James Jin Kang, Dang Bui, Thanh Pham, Huo-Chong Ling

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments 14 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.01446 2025-11-06 cs.LG 74%

Trustworthy Representation Learning via Information Funnels and Bottlenecks

João Machado de Freitas, Bernhard C. Geiger

机构 * Christian Doppler Laboratory for Dependable Intelligent Systems in Harsh Environments(可信智能系统在恶劣环境中的克里斯蒂安·多普勒实验室) Graz University of Technology(格拉茨技术大学) Know Center Research GmbH(Know Center研究有限责任公司)

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments Published in Machine Learning (Springer), vol. 114, no. 12, Article 267, 2025

Journal ref Mach Learn 114, 267 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00606 2025-11-05 cs.CL 74%

SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding

Jameson Sandler, Jacob K. Christopher, Thomas Hartvigsen, Ferdinando Fioretto

专题命中 安全评测 :alignment(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02148 2025-11-05 cs.LG 74%

CFL: On the Use of Characteristic Function Loss for Domain Alignment in Machine Learning

Abdullah Almansour, Ozan Tonguz

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全评测 :alignment(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14681 2025-10-31 cs.CL 74%

Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality

Yuto Harada, Yusuke Yamauchi, Yusuke Oda, Yohei Oseki, Yusuke Miyao, Yu Takagi

机构 * NII LLMC(日本信息处理学会大语言模型中心) The University of Tokyo(东京大学) NAIST(日本科学技术大学) Nagoya Institute of Technology(名古屋技术大学)

专题命中 安全评测 :alignment(title);分类 cs.CL

Comments Accepted to EMNLP 2025 (Main Conference). Models and evaluation results available at: https://github.com/llm-jp/massive-sft

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01128 2025-10-29 cs.CV cs.AI 74%

RipVIS: Rip Currents Video Instance Segmentation Benchmark for Beach Monitoring and Safety

Andrei Dumitriu, Florin Tatui, Florin Miron, Aakash Ralhan, Radu Tudor Ionescu, Radu Timofte

机构 * Computer Vision Lab, CAIDAS & IFI, University of Würzburg, Germany(计算机视觉实验室,CAIDAS与IFI,乌尔姆大学,德国) University of Bucharest, Romania(布加勒斯特大学,罗马尼亚)

专题命中 安全评测 :safety(title);分类 cs.AI

Comments Accepted at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03815 2025-10-07 eess.SY cs.LG cs.SY eess.SP 74%

A Trustworthy Industrial Fault Diagnosis Architecture Integrating Probabilistic Models and Large Language Models

Yue wu

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments 1tables,6 figs,11pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14744 2025-10-07 cs.LG 74%

Beyond the Single-Best Model: Rashomon Partial Dependence Profile for Trustworthy Explanations in AutoML

Mustafa Cavus, Jan N. van Rijn, Przemysław Biecek

机构 * Department of Statistics, Eskisehir Technical University, Turkiye(埃斯基谢普大学统计系) Leiden Institute of Advanced Computer Science, Leiden University, the Netherlands(莱顿大学高级计算机科学研究所) Faculty of Mathematics and Information Science, Warsaw University of Technology, Poland(华沙理工大学数学与信息科学学院) Informatics and Mechanics, University of Warsaw, Faculty of Mathematics, Poland(华沙大学信息技术与力学系)

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments Accepted at 28th International Conference on Discovery Science 2025

Journal ref In: Džeroski, S., Levatić, J., Pio, G., Simidjievski, N. (eds) Discovery Science. DS 2025. Lecture Notes in Computer Science, vol 16090. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17559 2025-09-23 cs.CL 74%

Specification-Aware Machine Translation and Evaluation for Purpose Alignment

Yoko Kayano, Saku Sugawara

机构 * The Graduate University for Advanced Studies (SOKENDAI)(高级研究大学(SOKENDAI)) National Institute of Informatics(信息研究所)

专题命中 安全评测 :alignment(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17514 2025-09-18 cs.AI 74%

TAI Scan Tool: A RAG-Based Tool With Minimalistic Input for Trustworthy AI Self-Assessment

Athanasios Davvetas, Xenia Ziouvelou, Ypatia Dami, Alexios Kaponis, Konstantina Giouvanopoulou, Michael Papademas

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments 9 pages, 1 figure, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08912 2025-09-12 cs.CY cs.HC 74%

Towards Trustworthy AI: Characterizing User-Reported Risks across LLMs "In the Wild"

Lingyao Li, Renkai Ma, Zhaoqian Xue, Junjie Xiong

专题命中 安全评测 :trustworthy(title);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18827 2025-09-09 cs.SE cs.AI 74%

Test It Before You Trust It: Applying Software Testing for Trustworthy In-context Learning

Teeradaj Racharak, Chaiyong Ragkhitwetsagul, Chommakorn Sontesadisai, Thanwadee Sunetnanta

机构 * Advanced Institute of So-Go-Chi (Convergence Knowledge) Informatics(融合知识研究院) Tohoku University(东北大学) Japan Advanced Institute of Science and Technology(日本先进科学研究院) Faculty of Information and Communication Technology(信息与通信技术学院) Mahidol University(玛希敦大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Journal ref Natural Language Processing and Information Systems (NLDB 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00673 2025-08-04 cs.CL 74%

MELAC: Massive Evaluation of Large Language Models with Alignment of Culture in Persian Language

Farhan Farsi, Farnaz Aghababaloo, Shahriar Shariati Motlagh, Parsa Ghofrani, MohammadAli SadraeiJavaheri, Shayan Bali, Amirhossein Shabani, Farbod Bijary, Ghazal Zamaninejad, AmirMohammad Salehoof, Saeedeh Momtazi

机构 * Amirkabir University of Technology(阿姆irkabir技术大学) Part AI Research Center(Part人工智能研究中心) University of Mazandaran(马赞德兰大学) King’s College London(伦敦国王学院)

专题命中 安全评测 :alignment(title);分类 cs.CL

Comments Preprint. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12384 2025-07-17 cs.LG cs.ET 74%

Trustworthy Tree-based Machine Learning by $MoS_2$ Flash-based Analog CAM with Inherent Soft Boundaries

Bo Wen, Guoyun Gao, Zhicheng Xu, Ruibin Mao, Xiaojuan Qi, X. Sharon Hu, Xunzhao Yin, Can Li

机构 * Department of Electrical and Electronic Engineering, The University of Hong Kong, Hong Kong SAR, China(香港大学电子与电气工程系) College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou, China(浙江大学信息科学与电子工程学院) Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, IN, USA(Notre Dame 大学计算机科学与工程系) Center for Advanced Semiconductor and Integrated Circuit, The University of Hong Kong, Hong Kong SAR, China(香港大学先进半导体与集成电路中心)

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15821 2025-06-23 cs.GR cs.AI cs.CV eess.IV 74%

VEIGAR: View-consistent Explicit Inpainting and Geometry Alignment for 3D object Removal

Pham Khai Nguyen Do, Bao Nguyen Tran, Nam Nguyen, Duc Dung Nguyen

机构 * AITech Lab(AITech实验室) Computer Science and Engineering Faculty(计算机科学与工程学院) Ho Chi Minh City University of Technology(胡志明市技术大学) VNUHCM

专题命中 安全评测 :alignment(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02621 2025-06-18 cs.NE cs.AI 74%

LLMs Help Alleviate the Cross-Subject Variability in Brain Signal and Language Alignment

Yifei Liu, Hengwei Ye, Shuhang Li

专题命中 安全评测 :alignment(title);分类 cs.AI

Comments The result is no longer believeable. Teaching force issue exists in the infer time of LLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17934 2025-06-06 cs.HC cs.CL cs.CR 74%

Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents

Chaoran Chen, Zhiping Zhang, Ibrahim Khalilov, Bingcan Guo, Simret A Gebreegziabher, Yanfang Ye, Ziang Xiao, Yaxing Yao, Tianshi Li, Toby Jia-Jun Li

机构 * University of Notre Dame(圣约翰大学) Northeastern University(东北大学) Virginia Tech(弗吉尼亚理工大学) University of Washington(华盛顿大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 安全评测 :trustworthy(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23813 2025-06-02 cs.CR cs.AI 74%

DP-RTFL: Differentially Private Resilient Temporal Federated Learning for Trustworthy AI in Regulated Industries

Abhijit Talluri

机构 * RTFL Project Contributor(RTFL项目贡献者)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments 6 pages (IEEE conference format), 10 figures. Source code available at https://github.com/abhitall/federated-credit-risk-rtfl.git

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21115 2025-05-28 cs.CL 74%

Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA

Sergey Pletenev, Maria Marina, Nikolay Ivanov, Daria Galimzianova, Nikita Krayko, Mikhail Salnikov, Vasily Konovalov, Alexander Panchenko, Viktor Moskvoretskii

专题命中 安全评测 :trustworthy(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20882 2025-05-28 cs.LG cs.SI 74%

Fedivertex: a Graph Dataset based on Decentralized Social Networks for Trustworthy Machine Learning

Marc Damie, Edwige Cyffers

机构 * University of Twente(代尔夫特理工大学) Inria(法国国家信息与自动化技术研究所) Institute of Science and Technology Austria(奥地利科学与技术研究所)

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏