arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1832 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1832 篇

2503.08931 2025-03-13 cs.CY 57%

ARCHED: A Human-Centered Framework for Transparent, Responsible, and Collaborative AI-Assisted Instructional Design

Hongming Li, Yizirui Fang, Shan Zhang, Seiyon M. Lee, Yiming Wang, Mark Trexler, Anthony F. Botelho

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments Accepted to the iRAISE Workshop at AAAI 2025. To be published in PMLR Volume 273

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08849 2025-03-11 cs.AI 57%

Path To Gain Functional Transparency In Artificial Intelligence With Meaningful Explainability

Md. Tanzib Hosain, Mehedi Hasan Anik, Sadman Rafi, Rana Tabassum, Khaleque Insia, Md. Mehrab Siddiky

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments Hosain, M. T., Anik, M. H., Rafi, S., Tabassum, R., Insia, K., & Sıddıky, M. M. (2023). Path to gain functional transparency in artificial intelligence with meaningful explainability. Journal of Metaverse, 3(2), 166-180

Journal ref Journal of Metaverse, 3(2), 166-180 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02865 2025-03-06 cs.CL 57%

FairSense-AI: Responsible AI Meets Sustainability

Shaina Raza, Mukund Sayeeganesh Chettiar, Matin Yousefabadi, Tahniat Khan, Marcelo Lotif

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05364 2025-03-06 cs.CR cs.AI 57%

Is On-Device AI Broken and Exploitable? Assessing the Trust and Ethics in Small Language Models

Kalyan Nakka, Jimmy Dani, Nitesh Saxena

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments 26 pages, 31 figures and 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21250 2025-03-03 cs.AI 57%

Towards Developing Ethical Reasoners: Integrating Probabilistic Reasoning and Decision-Making for Complex AI Systems

Nijesh Upreti, Jessica Ciupa, Vaishak Belle

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18359 2025-02-26 cs.CY 57%

Responsible AI Agents

Deven R. Desai, Mark O. Riedl

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11031 2025-02-18 cs.LG 57%

A Critical Review of Predominant Bias in Neural Networks

Jiazhi Li, Mahyar Khayatkhoei, Jiageng Zhu, Hanchen Xie, Mohamed E. Hussein, Wael AbdAlmageed

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

Comments 31 pages, 8 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09628 2025-02-18 cs.AI 57%

Artificial Intelligence-Driven Clinical Decision Support Systems

Muhammet Alkan, Idris Zakariyya, Samuel Leighton, Kaushik Bhargav Sivangi, Christos Anagnostopoulos, Fani Deligianni

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments Added acknowledgements for the corresponding author, updated Figure 4, 5 & 6

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10400 2025-02-18 cs.CL 57%

Self-Reflection Makes Large Language Models Safer, Less Biased, and Ideologically Neutral

Fengyuan Liu, Nouar AlDahoul, Gregory Eady, Yasir Zaki, Talal Rahwan

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06472 2025-02-14 cs.RO cs.AI cs.HC 57%

Enabling Novel Mission Operations and Interactions with ROSA: The Robot Operating System Agent

Rob Royce, Marcel Kaufmann, Jonathan Becktor, Sangwoo Moon, Kalind Carpenter, Kai Pak, Amanda Towler, Rohan Thakker, Shehryar Khattak

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Preprint. Accepted at IEEE Aerospace Conference 2025, 16 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14040 2025-02-10 cs.CY 57%

Global Perspectives of AI Risks and Harms: Analyzing the Negative Impacts of AI Technologies as Prioritized by News Media

Mowafak Allaham, Kimon Kieslich, Nicholas Diakopoulos

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03472 2025-02-07 cs.CY 57%

Powering LLM Regulation through Data: Bridging the Gap from Compute Thresholds to Customer Experiences

Wesley Pasfield

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Presented at the 2nd Workshop on Regulatable ML at NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04683 2025-02-07 cs.AI 57%

From Principles to Practice: A Deep Dive into AI Ethics and Regulations

Nan Sun, Yuantian Miao, Hao Jiang, Ming Ding, Jun Zhang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Submitted to JAIR

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15985 2025-01-28 cs.CY 57%

Demographic Benchmarking: Bridging Socio-Technical Gaps in Bias Detection

Gemma Galdon Clavell, Rubén González-Sendino, Paola Vazquez

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15571 2025-01-28 cs.CL 57%

Cross-Cultural Fashion Design via Interactive Large Language Models and Diffusion Models

Spencer Ramsey, Amina Grant, Jeffrey Lee

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20130 2025-01-28 cs.HC cs.CY 57%

The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI Relationships

Renwen Zhang, Han Li, Han Meng, Jinyuan Zhan, Hongyuan Gan, Yi-Chieh Lee

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01957 2025-01-24 cs.AI 57%

Usage Governance Advisor: From Intent to AI Governance

Elizabeth M. Daly, Sean Rooney, Seshu Tirupathi, Luis Garces-Erice, Inge Vejsbjerg, Frank Bagehorn, Dhaval Salwala, Christopher Giblin, Mira L. Wolf-Bauwens, Ioana Giurgiu, Michael Hind, Peter Urbanetz

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 9 pages, 8 figures, AAAI workshop submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12521 2025-01-23 cs.SE cs.AI 57%

An Empirically-grounded tool for Automatic Prompt Linting and Repair: A Case Study on Bias, Vulnerability, and Optimization in Developer Prompts

Dhia Elhaq Rzig, Dhruba Jyoti Paul, Kaiser Pister, Jordan Henkel, Foyzul Hassan

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.06687 2025-01-22 cs.CL 57%

Hire Me or Not? Examining Language Model's Behavior with Occupation Attributes

Damin Zhang, Yi Zhang, Geetanjali Bihani, Julia Rayz

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments COLING 2025

Journal ref Proceedings of the 31st International Conference on Computational Linguistics (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14062 2025-01-20 cs.HC cs.CY 57%

Understanding and Evaluating Trust in Generative AI and Large Language Models for Spreadsheets

Simon Thorne

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

Journal ref Proceedings of the EuSpRIG 2024 Conference "Spreadsheet Productivity & Risks" ISBN : 978-1-905404-59-9

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06695 2025-01-14 cs.AI 57%

DVM: Towards Controllable LLM Agents in Social Deduction Games

Zheng Zhang, Yihuai Lan, Yangsen Chen, Lei Wang, Xiang Wang, Hao Wang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17114 2025-01-14 cs.AI cs.ET 57%

Decentralized Governance of Autonomous AI Agents

Tomer Jordi Chaffer, Charles von Goins, Bayo Okusanya, Dontrail Cotlage, Justin Goldston

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05617 2025-01-13 cs.CY cs.DL 57%

Datasheets for Healthcare AI: A Framework for Transparency and Bias Mitigation

Marjia Siddik, Harshvardhan J. Pandit

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments Irish Conference on Artificial Intelligence and Cognitive Science (AICS), December 2024, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04437 2025-01-09 eess.SY cs.AI cs.ET cs.SY 57%

Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions

Doaa Mahmud, Hadeel Hajmohamed, Shamma Almentheri, Shamma Alqaydi, Lameya Aldhaheri, Ruhul Amin Khalil, Nasir Saeed

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Accepted for publication in IEEE Transactions on Intelligent Transportation Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19915 2025-01-08 econ.GN cs.AI q-fin.EC 57%

AI-Driven Scenarios for Urban Mobility: Quantifying the Role of ODE Models and Scenario Planning in Reducing Traffic Congestion

Katsiaryna Bahamazava

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02368 2025-01-07 cs.AI cs.HC 57%

Enhancing Workplace Productivity and Well-being Using AI Agent

Ravirajan K, Arvind Sundarajan

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18588 2024-12-25 cs.RO cs.AI cs.SY eess.SY 57%

A Paragraph is All It Takes: Rich Robot Behaviors from Interacting, Trusted LLMs

OpenMind, Shaohong Zhong, Adam Zhou, Boyuan Chen, Homin Luo, Jan Liphardt

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 10 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17505 2024-12-24 stat.ML cs.LG 57%

More is Less? A Simulation-Based Approach to Dynamic Interactions between Biases in Multimodal Models

Mounia Drissi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00469 2024-12-23 cs.CY 57%

Beyond Incompatibility: Trade-offs between Mutually Exclusive Fairness Criteria in Machine Learning and Law

Meike Zehlike, Alex Loosley, Håkan Jonsson, Emil Wiedemann, Philipp Hacker

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04310 2024-12-19 cs.CL 57%

Montague semantics and modifier consistency measurement in neural language models

Danilo S. Carvalho, Edoardo Manino, Julia Rozanova, Lucas Cordeiro, André Freitas

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏