arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1832 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1832 篇

2503.04785 2025-05-06 cs.CL cs.CY 62%

Mapping Trustworthiness in Large Language Models: A Bibliometric Analysis Bridging Theory to Practice

José Siqueira de Cerqueira, Kai-Kristian Kemell, Rebekah Rousi, Nannan Xi, Juho Hamari, Pekka Abrahamsson

机构 * Tampere University(塔尔基耶大学) University of Vaasa(瓦萨大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00339 2025-05-02 cs.CL cs.AI 62%

Enhancing AI-Driven Education: Integrating Cognitive Frameworks, Linguistic Feedback Analysis, and Ethical Considerations for Improved Content Generation

Antoun Yaacoub, Sansiri Tarnpradab, Phattara Khumprom, Zainab Assaghir, Lionel Prevost, Jérôme Da-Rugna

机构 * Learning, Data and Robotics (LDR) ESIEA Lab(学习、数据与机器人(LDR)ESIEA实验室) Department of Computer Engineering(计算机工程系) Graduate School of Management and Innovation(管理与创新研究生院) Faculty of Science(科学学院) Lebanese University(黎巴嫩大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments This article will be presented in IJCNN 2025 "AI Innovations for Education: Transforming Teaching and Learning through Cutting-Edge Technologies" workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09495 2025-05-01 cs.LG cs.AI 62%

FADE: Towards Fairness-aware Generation for Domain Generalization via Classifier-Guided Score-based Diffusion Models

Yujie Lin, Dong Li, Minglai Shao, Guihong Wan, Chen Zhao

机构 * School of New Media and Communication, Tianjin University(天津大学新媒体与传播学院) School of Informatics, Xiamen University(厦门大学信息学院) Department of Computer Science, Baylor University(贝勒大学计算机科学系) Departments of Biostatistics and Epidemiology, Harvard University(哈佛大学生物统计学与流行病学系)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19255 2025-04-29 cs.AI cs.CY 62%

The Convergent Ethics of AI? Analyzing Moral Foundation Priorities in Large Language Models with a Multi-Framework Approach

Chad Coleman, W. Russell Neuman, Ali Dasdan, Safinah Ali, Manan Shah

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 25 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17401 2025-04-29 cs.CY cs.AI cs.HC 62%

AIJIM: A Scalable Model for Real-Time AI in Environmental Journalism

Torsten Tiltack

机构 * Torsten Tiltack(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 22 pages, 10 figures, 5 tables. Keywords: Artificial Intelligence, Environmental Journalism, Real-Time Reporting, Vision Transformers, Image Recognition, Crowdsourced Validation, GPT-4, Automated News Generation, GIS Integration, Data Privacy Compliance, Explainable AI (XAI), AI Ethics, Sustainable Development

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04476 2025-04-28 cs.CY cs.AI 62%

The Moral Mind(s) of Large Language Models

Avner Seror

机构 * Aix Marseille Univ, CNRS, AMSE, Marseille, France(阿维尼翁-马赛大学,国家科学研究中心,AMSE,马赛,法国)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16948 2025-04-25 cs.CY cs.AI cs.ET 62%

Intrinsic Barriers to Explaining Deep Foundation Models

Zhen Tan, Huan Liu

机构 * Arizona State University(亚利桑那州立大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16622 2025-04-24 cs.AI cs.CY 62%

Cognitive Silicon: An Architectural Blueprint for Post-Industrial Computing Systems

Christoforus Yoga Haryanto, Emily Lomempow

机构 * ZipThought

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Working Paper, 37 pages, 1 figure, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13957 2025-04-22 cs.CY cs.AI cs.CR 62%

Naming is framing: How cybersecurity's language problems are repeating in AI governance

Lianne Potter

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 20 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09946 2025-04-21 cs.CY cs.CL 62%

Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Qian Wang, Zhanzhi Lou, Zhenheng Tang, Nuo Chen, Xuandong Zhao, Wenxuan Zhang, Dawn Song, Bingsheng He

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12358 2025-04-18 cs.CY cs.AI physics.soc-ph 62%

Towards an AI Observatory for the Nuclear Sector: A tool for anticipatory governance

Aditi Verma, Elizabeth Williams

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments Presented at the Sociotechnical AI Governance Workshop at CHI 2025, Yokohama

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11501 2025-04-17 cs.CY cs.AI 62%

A Framework for the Private Governance of Frontier Artificial Intelligence

Dean W. Ball

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00041 2025-04-15 eess.SP cs.AI cs.LG 62%

Needles in Needle Stacks: Meaningful Clinical Information Buried in Noisy Waveform Data

Sujay Nagaraj, Andrew J. Goodwin, Dmytro Lopushanskyy, Danny Eytan, Robert W. Greer, Sebastian D. Goodfellow, Azadeh Assadi, Anand Jayarajan, Anna Goldenberg, Mjaye L. Mazwi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Comments Machine Learning For Health Care 2024 (MLHC)

Journal ref PMLR (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08640 2025-04-14 cs.AI cs.CY cs.GT nlin.CD 62%

Do LLMs trust AI regulation? Emerging behaviour of game-theoretic LLM agents

Alessio Buscemi, Daniele Proverbio, Paolo Bova, Nataliya Balabanova, Adeela Bashir, Theodor Cimpeanu, Henrique Correia da Fonseca, Manh Hong Duong, Elias Fernandez Domingos, Antonio M. Fernandes, Marcus Krellner, Ndidi Bianca Ogbo, Simon T. Powers, Fernando P. Santos, Zia Ush Shamszaman, Zhao Song, Alessandro Di Stefano, The Anh Han

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07516 2025-04-11 cs.CY cs.AI cs.HC 62%

Enhancements for Developing a Comprehensive AI Fairness Assessment Standard

Avinash Agarwal, Mayashankar Kumar, Manisha J. Nene

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 5 pages. Published in 2025 17th International Conference on COMmunication Systems and NETworks (COMSNETS). Access: https://ieeexplore.ieee.org/abstract/document/10885551

Journal ref 2025 17th International Conference on COMmunication Systems and NETworks (COMSNETS), Bengaluru, India, 2025, pp. 1216-1220

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07118 2025-04-11 cs.CY cs.AI cs.ET 62%

Sacred or Secular? Religious Bias in AI-Generated Financial Advice

Muhammad Salar Khan, Hamza Umer

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12848 2025-04-10 cs.CY cs.AI cs.SI 62%

ClarityEthic: Explainable Moral Judgment Utilizing Contrastive Ethical Insights from Large Language Models

Yuxi Sun, Wei Gao, Jing Ma, Hongzhan Lin, Ziyang Luo, Wenxuan Zhang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments We have noticed that this version of our experiment and method description isn't quite complete or accurate. To make sure we present our best work, we think it would be a good idea to withdraw the manuscript for now and take some time to revise and reformat it

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02917 2025-04-07 cs.CL cs.AI 62%

Bias in Large Language Models Across Clinical Applications: A Systematic Review

Thanathip Suenghataiphorn, Narisara Tribuddharat, Pojsakorn Danpanichkul, Narathorn Kulthamrongsri

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01810 2025-04-03 cs.CY cs.AI 62%

Propaganda is all you need

Paul Kronlund-Drouault

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00652 2025-04-02 cs.CY cs.AI cs.ET 62%

Towards Adaptive AI Governance: Comparative Insights from the U.S., EU, and Asia

Vikram Kulothungan, Deepti Gupta

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Accepted at IEEE BigDataSecurity 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00241 2025-04-02 cs.CL cs.AI 62%

Synthesizing Public Opinions with LLMs: Role Creation, Impacts, and the Future to eDemorcacy

Rabimba Karanjai, Boris Shor, Amanda Austin, Ryan Kennedy, Yang Lu, Lei Xu, Weidong Shi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24228 2025-04-01 cs.AI cs.CL cs.MA 62%

PAARS: Persona Aligned Agentic Retail Shoppers

Saab Mansour, Leonardo Perelli, Lorenzo Mainetti, George Davidson, Stefano D'Amato

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17756 2025-03-25 cs.LG cs.AI cs.CR cs.NE 62%

Bandwidth Reservation for Time-Critical Vehicular Applications: A Multi-Operator Environment

Abdullah Al-Khatib, Abdullah Ahmed, Klaus Moessner, Holger Timinger

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06795 2025-03-13 cs.CL cs.AI 62%

Bridging the Fairness Gap: Enhancing Pre-trained Models with LLM-Generated Sentences

Liu Yu, Ludie Guo, Ping Kuang, Fan Zhou

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Journal ref ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14210 2025-03-13 cs.LG cs.AI 62%

Fair Overlap Number of Balls (Fair-ONB): A Data-Morphology-based Undersampling Method for Bias Reduction

José Daniel Pascual-Triana, Alberto Fernández, Paulo Novais, Francisco Herrera

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 14 pages, 5 tables, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06523 2025-03-11 cs.CY cs.AI 62%

Generative AI as Digital Media

Gilad Abiri

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Journal ref Harv. J. Sports & Ent. L. 15 (2024): 279

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05937 2025-03-11 cs.CY cs.AI 62%

The Unified Control Framework: Establishing a Common Foundation for Enterprise AI Governance, Risk Management and Regulatory Compliance

Ian W. Eisenberg, Lucía Gamboa, Eli Sherman

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02255 2025-03-05 cs.CL cs.LG 62%

AxBERT: An Interpretable Chinese Spelling Correction Method Driven by Associative Knowledge Network

Fanyu Wang, Hangyu Zhu, Zhenping Xie

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16696 2025-02-25 cs.LG cs.AI 62%

Dynamic LLM Routing and Selection based on User Preferences: Balancing Performance, Cost, and Ethics

Deepak Babu Piskala, Vijay Raajaa, Sachin Mishra, Bruno Bozza

专题命中 AI治理与伦理 :harmlessness(abstract);分类 cs.AI、cs.LG

Journal ref International Journal of Computer Applications, Vol. 186, No. 51, November 2024, pp. 1-7

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15357 2025-02-24 cs.CY cs.AI 62%

Integrating Generative AI in Cybersecurity Education: Case Study Insights on Pedagogical Strategies, Critical Thinking, and Responsible AI Use

Mahmoud Elkhodr, Ergun Gide

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏