arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9324 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9324 篇

2502.11066 2025-05-21 cs.CL 74%

CARMA: Enhanced Compositionality in LLMs via Advanced Regularisation and Mutual Information Alignment

Nura Aljaafari, Danilo S. Carvalho, André Freitas

机构 * Department of Computer Science, University of Manchester(曼彻斯特大学计算机科学系) Idiap Research Institute(Idiap研究机构) National Biomarker Centre, CRUK-MI, Univ. of Manchester(国家生物标记中心,CRUK-MI,曼彻斯特大学)

专题命中 安全评测 :alignment(title);分类 cs.CL

Comments 19 pages, 8 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09083 2025-05-15 econ.GN cs.CL q-fin.EC 74%

Ornithologist: Towards Trustworthy "Reasoning" about Central Bank Communications

Dominic Zaun Eu Jones

机构 * Reserve Bank of Australia(澳大利亚储备银行)

专题命中 安全评测 :trustworthy(title);分类 cs.CL

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02843 2025-05-07 eess.IV cs.AI cs.CV physics.med-ph 74%

Physical foundations for trustworthy medical imaging: a review for artificial intelligence researchers

Miriam Cobo, David Corral Fontecha, Wilson Silva, Lara Lloret Iglesias

机构 * IFCA.unican.es(IFCA大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments 17 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15496 2025-04-29 cs.LG 74%

Verification and Validation for Trustworthy Scientific Machine Learning

John D. Jakeman, Lorena A. Barba, Joaquim R. R. A. Martins, Thomas O'Leary-Roseberry

机构 * Sandia National Laboratories(桑迪亚国家实验室) The George Washington University(乔治·华盛顿大学) University of Michigan(密歇根大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13399 2025-04-21 cs.CV cs.AI 74%

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety

Shashank Shriram, Srinivasa Perisetla, Aryan Keskar, Harsha Krishnaswamy, Tonko Emil Westerhof Bossen, Andreas Møgelmose, Ross Greer

机构 * Machine Intelligence, Interaction, and Imagination (Mi 3 ) Laboratory(机器智能、交互与想象(Mi 3)实验室) University of California, Merced(加州大学默塞德分校) Aalborg Universitet(奥胡斯大学)

专题命中 安全评测 :safety(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17139 2025-04-18 cs.AI 74%

Trustworthy XAI and Application

MD Abdullah Al Nasim, A. S. M Anas Ferdous, Abdur Rashid, Fatema Tuj Johura Soshi, Parag Biswas, Angona Biswas, Kishor Datta Gupta

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20472 2025-03-27 cs.CV cs.AI 74%

From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment

Yucheng Suo, Fan Ma, Linchao Zhu, Tianyi Wang, Fengyun Rao, Yi Yang

专题命中 安全评测 :alignment(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00519 2025-02-05 cs.SE cs.LG 74%

CoDocBench: A Dataset for Code-Documentation Alignment in Software Maintenance

Kunal Pai, Premkumar Devanbu, Toufique Ahmed

专题命中 安全评测 :alignment(title);分类 cs.LG

Comments Accepted at the 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR) - Data and Tool Showcase Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11909 2025-01-22 cs.AI 74%

Bridging the Communication Gap: Evaluating AI Labeling Practices for Trustworthy AI Development

Raphael Fischer, Magdalena Wischnewski, Alexander van der Staay, Katharina Poitz, Christian Janiesch, Thomas Liebig

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10379 2025-01-22 cs.CY 74%

What Information Should Be Shared with Whom "Before and During Training"?

Haydn Belfield

专题命中 安全评测 :safety(abstract,comments);AI safety(abstract,comments);分类 cs.CY

Comments To be published in the proceedings of the 2024 Conference on Frontier AI Safety Commitments. 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09218 2025-01-17 q-bio.QM cs.AI 74%

Interpretable Droplet Digital PCR Assay for Trustworthy Molecular Diagnostics

Yuanyuan Wei, Yucheng Wu, Fuyang Qu, Yao Mu, Yi-Ping Ho, Ho-Pui Ho, Wu Yuan, Mingkun Xu

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20674 2024-12-31 cs.DC cs.CR cs.LG 74%

Blockchain-Empowered Cyber-Secure Federated Learning for Trustworthy Edge Computing

Ervin Moore, Ahmed Imteaj, Md Zarif Hossain, Shabnam Rezapour, M. Hadi Amini

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18971 2024-12-30 cs.LG 74%

Adopting Trustworthy AI for Sleep Disorder Prediction: Deep Time Series Analysis with Temporal Attention Mechanism and Counterfactual Explanations

Pegah Ahadian, Wei Xu, Sherry Wang, Qiang Guan

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Journal ref IEEE Bigdata 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17527 2024-12-24 cs.AI 74%

Enhancing Cancer Diagnosis with Explainable & Trustworthy Deep Learning Models

Badaru I. Olumuyiwa, The Anh Han, Zia U. Shamszaman

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11713 2024-12-17 cs.CL cs.SE 74%

Seeker: Towards Exception Safety Code Generation with Intermediate Language Agents Framework

Xuanming Zhang, Yuxuan Chen, Yiming Zheng, Zhexin Zhang, Yuan Yuan, Minlie Huang

专题命中 安全评测 :safety(title);分类 cs.CL

Comments 30 pages, 9 figures, submitted to ARR Dec

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18222 2024-12-09 cs.AI 74%

Trustworthy AI: Securing Sensitive Data in Large Language Models

Georgios Feretzakis, Vassilios S. Verykios

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments 40 pages, 1 figure

Journal ref AI 5(4), 2773-2800 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08889 2024-11-15 cs.HC cs.AI cs.SD eess.AS 74%

Multilingual Standalone Trustworthy Voice-Based Social Network for Disaster Situations

Majid Behravan, Elham Mohammadrezaei, Mohamed Azab, Denis Gracanin

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments Accepted for publication in IEEE UEMCON 2024, to appear in December 2024. 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03376 2024-11-07 cs.DC cs.AI cs.NI 74%

An Open API Architecture to Discover the Trustworthy Explanation of Cloud AI Services

Zerui Wang, Yan Liu, Jun Huang

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments Published in: IEEE Transactions on Cloud Computing ( Volume: 12, Issue: 2, April-June 2024)

Journal ref IEEE Transactions on Cloud Computing ( Volume: 12, Issue: 2, April-June 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01335 2024-11-04 cs.CV cs.AI 74%

BehAVE: Behaviour Alignment of Video Game Encodings

Nemanja Rašajski, Chintan Trivedi, Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis

专题命中 安全评测 :alignment(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15706 2024-09-25 cs.HC cs.AI 74%

Improving Emotional Support Delivery in Text-Based Community Safety Reporting Using Large Language Models

Yiren Liu, Yerong Li, Ryan Mayfield, Yun Huang

专题命中 安全评测 :safety(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04407 2024-07-08 cs.LG 74%

Trustworthy Classification through Rank-Based Conformal Prediction Sets

Rui Luo, Zhixin Zhou

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01862 2024-07-02 cs.CL cs.DB 74%

$R^3$-NL2GQL: A Model Coordination and Knowledge Graph Alignment Approach for NL2GQL

Yuhang Zhou, Yu He, Siyu Tian, Yuchen Ni, Zhangyue Yin, Xiang Liu, Chuanjun Ji, Sen Liu, Xipeng Qiu, Guangnan Ye, Hongfeng Chai

专题命中 安全评测 :alignment(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11845 2024-06-21 cs.CY cs.HC 74%

Decoding the Digital Fine Print: Navigating the potholes in Terms of service/ use of GenAI tools against the emerging need for Transparent and Trustworthy Tech Futures

Sundaraparipurnan Narayanan

专题命中 安全评测 :trustworthy(title);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07820 2024-06-13 cs.CV cs.LG 74%

Are Objective Explanatory Evaluation metrics Trustworthy? An Adversarial Analysis

Prithwijit Chowdhury, Mohit Prabhushankar, Ghassan AlRegib, Mohamed Deriche

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07080 2024-06-12 cs.CL 74%

DARA: Decomposition-Alignment-Reasoning Autonomous Language Agent for Question Answering over Knowledge Graphs

Haishuo Fang, Xiaodan Zhu, Iryna Gurevych

专题命中 安全评测 :alignment(title);分类 cs.CL

Comments Accepted by ACL2024 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11698 2024-04-19 cs.AI cs.DC 74%

A Secure and Trustworthy Network Architecture for Federated Learning Healthcare Applications

Antonio Boiano, Marco Di Gennaro, Luca Barbieri, Michele Carminati, Monica Nicoli, Alessandro Redondi, Stefano Savazzi, Albert Sund Aillet, Diogo Reis Santos, Luigi Serio

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09668 2024-03-18 cs.CV cs.AI 74%

Trustworthy Automated Driving through Qualitative Scene Understanding and Explanations

Nassim Belmecheri, Arnaud Gotlieb, Nadjib Lazaar, Helge Spieker

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments Transport Research Arena (TRA) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00446 2024-02-02 cs.CL 74%

Improving Dialog Safety using Socially Aware Contrastive Learning

Souvik Das, Rohini K. Srihari

专题命中 安全评测 :safety(title);分类 cs.CL

Comments SCI-CHAT@EACL2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04231 2023-12-08 cs.CV cs.AI 74%

Adventures of Trustworthy Vision-Language Models: A Survey

Mayank Vatsa, Anubhooti Jain, Richa Singh

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments Accepted in AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01213 2023-12-05 cs.AR cs.AI 74%

Recent Advances in Scalable Energy-Efficient and Trustworthy Spiking Neural networks: from Algorithms to Technology

Souvik Kundu, Rui-Jie Zhu, Akhilesh Jaiswal, Peter A. Beerel

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏