arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1832 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1832 篇

2510.14443 2025-10-17 cs.SD cs.AI eess.AS 57%

Big Data Approaches to Bovine Bioacoustics: A FAIR-Compliant Dataset and Scalable ML Framework for Precision Livestock Welfare

Mayuri Kate, Suresh Neethirajan

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 40 pages, 14 figures, 9 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10623 2025-10-17 cs.LG cs.CV 57%

Flows and Diffusions on the Neural Manifold

Daniel Saragih, Deyu Cao, Tejas Balaji

机构 * Queen’s University and Vector Institute(女王大学和向量研究所) University of Tokyo(东京大学) University of Toronto(多伦多大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 43 pages, 11 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13699 2025-10-17 cs.IR cs.AI 57%

A Comprehensive Review of Recommender Systems: Transitioning from Theory to Practice

Shaina Raza, Mizanur Rahman, Safiullah Kamawal, Armin Toroghi, Ananya Raval, Farshad Navah, Amirmohammad Kazemeini

机构 * Vector Institute(向量研究所) Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments we quarterly update of this literature

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13195 2025-10-16 cs.AI 57%

Emotional Cognitive Modeling Framework with Desire-Driven Objective Optimization for LLM-empowered Agent in Social Simulation

Qun Ma, Xiao Xue, Xuwen Zhang, Zihan Zhao, Yuwei Guo, Ming Zhang

机构 * College of Intelligence and Computing(智能与计算学院) Faculty of Environment, Science and Economy(环境、科学与经济学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16152 2025-10-16 cs.CL 57%

Towards Region-aware Bias Evaluation Metrics

Angana Borah, Aparna Garimella, Rada Mihalcea

机构 * University of Michigan(密歇根大学) Adobe Research(Adobe研究)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted to Cross-Cultural Considerations in NLP (C3NLP Workshop at NAACL 2025) -- Outstanding Paper Award

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18943 2025-10-15 cs.CL 57%

MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems

Xuanming Zhang, Yuxuan Chen, Samuel Yeh, Sharon Li

机构 * Tsinghua University(清华大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments NeurIPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11595 2025-10-14 cs.AI cs.GL 57%

Reproducibility: The New Frontier in AI Governance

Israel Mason-Williams, Gabryel Mason-Williams

机构 * Imperial College London(帝国理工学院) King's College London(国王学院) Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments 12 pages,6 figures,Workshop on Technical AI Governance at ICML

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11079 2025-10-14 cs.AI 57%

Argumentation-Based Explainability for Legal AI: Comparative and Regulatory Perspectives

Andrada Iulia Prajescu, Roberto Confalonieri

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08086 2025-10-10 cs.AI 57%

From Ethical Declarations to Provable Independence: An Ontology-Driven Optimal-Transport Framework for Certifiably Fair AI Systems

Sukriti Bhattacharya, Chitro Majumdar

机构 * Senior Scientist, Trustworthy AI, Luxembourg Institute of Science \& Technology, Maison de l'innovation 5, L-4362, Luxembourg Chief Investment Risk Strategist for Sovereign Institutions Founder, RsRL, Jumeirah Beach Residence (JBR), P.O.Box 29215, Dubai, United Arab Emirates

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments 19 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06623 2025-10-09 cs.LG 57%

DPA-Net: A Dual-Path Attention Neural Network for Inferring Glycemic Control Metrics from Self-Monitored Blood Glucose Data

Canyu Lei, Benjamin Lobo, Jianxin Xie

机构 * Binghamton University, Department of Computer Science(宾夕法尼亚州立大学比恩代特分校计算机科学系) University of Virginia, School of Data Science(弗吉尼亚大学数据科学学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04577 2025-10-07 cs.SD cs.LG cs.MM eess.AS 57%

Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

Juncheng Wang, Chao Xu, Cheng Yu, Zhe Hu, Haoyu Xie, Guoqi Yu, Lei Shang, Shujun Wang

机构 * The Hong Kong Polytechnic University(香港理工大学) Alibaba Group(阿里巴巴集团)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01963 2025-10-03 cs.SD cs.LG 57%

Bias beyond Borders: Global Inequalities in AI-Generated Music

Ahmet Solak, Florian Grötschla, Luca A. Lanzendörfer, Roger Wattenhofer

机构 * ETH Zurich(苏黎世联邦理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01192 2025-10-03 cs.HC cs.CY cs.RO 57%

Better Than "Better Than Nothing": Design Strategies for Enculturated Empathetic AI Robot Companions for Older Adults

Isabel Pedersen, Andrea Slane

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

Comments 26 pages, 6 figures, version submitted to journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12344 2025-10-02 cs.LG cs.DC 57%

CYCle: Choosing Your Collaborators Wisely to Enhance Collaborative Fairness in Decentralized Learning

Nurbek Tastan, Samuel Horvath, Karthik Nandakumar

机构 * Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(莫德尔·本·扎耶德人工智能大学) Michigan State University (MSU)(密歇根州立大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments Published in TMLR 08/2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13758 2025-09-30 cs.CY 57%

Towards Evaluting Fake Reasoning Bias in Language Models

Qian Wang, Zhenheng Tang, Zhanzhi Lou, Nuo Chen, Wenxuan Wang, Bingsheng He

专题命中 AI治理与伦理 :DPO(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21542 2025-09-29 cs.HC cs.AI 57%

Psychological and behavioural responses in human-agent vs. human-human interactions: a systematic review and meta-analysis

Jianan Zhou, Fleur Corbett, Joori Byun, Talya Porat, Nejra van Zalk

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21207 2025-09-26 cs.LG 57%

From Physics to Machine Learning and Back: Part II - Learning and Observational Bias in PHM

Olga Fink, Ismail Nejjar, Vinay Sharma, Keivan Faghih Niresi, Han Sun, Hao Dong, Chenghao Xu, Amaury Wei, Arthur Bizzi, Raffael Theiler, Yuan Tian, Leandro Von Krannichfeldt, Zhan Ma, Sergei Garmaev, Zepeng Zhang, Mengjie Zhao

机构 * Intelligent Maintenance and Operations Systems Lab, EPFL, Lausanne, Switzerland(智能维护与运营系统实验室,EPFL,拉沃斯纳,瑞士)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16151 2025-09-22 cs.LG cs.CR 57%

Automated Cyber Defense with Generalizable Graph-based Reinforcement Learning Agents

Isaiah J. King, Benjamin Bowman, H. Howie Huang

机构 * Cybermonic LLC(Cybermonic公司) The George Washington University(乔治华盛顿大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15803 2025-09-22 cs.CV cs.AI 57%

CIDER: A Causal Cure for Brand-Obsessed Text-to-Image Models

Fangjian Shen, Zifeng Liang, Chao Wang, Wushao Wen

机构 * School of Computer Science(计算机科学学院) Engineering, Sun Yat-sen University, Guangzhou, China(工程学院,中山大学,广州,中国)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments 5 pages, 7 figures, submitted to ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12566 2025-09-22 cs.AI 57%

Exploring the Impact of Personality Traits on LLM Bias and Toxicity

Shuo Wang, Renhao Li, Xi Chen, Yulin Yuan, Derek F. Wong, Min Yang

机构 * University of Macau(澳门大学) Shenzhen Key Laboratory for High Performance Data Mining, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳高性能数据挖掘重点实验室,深圳先进技术研究院,中国科学院) Nanyang Technological University(南洋理工大学) Department of Chinese Language and Literature, University of Macau(中文语言文学系,澳门大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12104 2025-09-16 cs.AI 57%

JustEva: A Toolkit to Evaluate LLM Fairness in Legal Knowledge Inference

Zongyue Xue, Siyuan Zheng, Shaochun Wang, Yiran Hu, Shenran Wang, Yuxin Yao, Haitao Li, Qingyao Ai, Yiqun Liu, Yun Liu, Weixing Shen

机构 * Tsinghua University(清华大学) Yale Law School(耶鲁法学院) Shanghai Jiaotong University(上海交通大学) University of Waterloo(滑铁卢大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments This paper has been accepted at CIKM 2025 (Demo Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05929 2025-09-11 cs.CY 57%

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support

Keyang Qian, Shiqi Liu, Tongguang Li, Mladen Raković, Xinyu Li, Rui Guan, Inge Molenaar, Sadia Nawaz, Zachari Swiecki, Lixiang Yan, Dragan Gašević

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Journal ref Computers & Education, Volume 240, 2026, 105448

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10160 2025-09-11 cs.CV cs.AI 57%

PriorCLIP: Visual Prior Guided Vision-Language Model for Remote Sensing Image-Text Retrieval

Jiancheng Pan, Muyuan Ma, Qing Ma, Cong Bai, Shengyong Chen

机构 * IEEE Publication Technology Department(IEEE出版技术部)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04993 2025-09-08 cs.MA cs.AI 57%

LLM Enabled Multi-Agent System for 6G Networks: Framework and Method of Dual-Loop Edge-Terminal Collaboration

Zheyan Qu, Wenbo Wang, Zitong Yu, Boquan Sun, Yang Li, Xing Zhang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments This paper has been accepted by IEEE Communications Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20201 2025-09-08 cs.CL 57%

Social Bias in Multilingual Language Models: A Survey

Lance Calvin Lim Gamboa, Yue Feng, Mark Lee

机构 * School of Computer Science, University of Birmingham(伯明翰大学计算机科学学院) Department of Information Systems and Computer Science, Ateneo de Manila University(马尼拉大学信息系统与计算机科学系)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted into EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02025 2025-09-05 cs.DC cs.AI 57%

Evaluating the Efficacy of LLM-Based Reasoning for Multiobjective HPC Job Scheduling

Prachi Jadhav, Hongwei Jin, Ewa Deelman, Prasanna Balaprakash

机构 * University of Tennessee, Knoxville\ Ridge National Laboratory Oak Ridge, TN USA Argonne National Laboratory Lemont, IL USA University of Southern California Los Angeles, CA USA Oak Ridge National Laboratory Oak Ridge, TN USA University of Tennessee, Knoxville\ Ridge National Laboratory Argonne National Laboratory University of Southern California Oak Ridge National Laboratory

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 10 pages, 6 figures, work under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02007 2025-09-03 cs.AI 57%

mFARM: Towards Multi-Faceted Fairness Assessment based on HARMs in Clinical Decision Support

Shreyash Adappanavar, Krithi Shailya, Gokul S Krishnan, Sriraam Natarajan, Balaraman Ravindran

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02080 2025-09-01 eess.AS cs.AI 57%

Enhancing GOP in CTC-Based Mispronunciation Detection with Phonological Knowledge

Aditya Kamlesh Parikh, Cristian Tejedor-Garcia, Catia Cucchiarini, Helmer Strik

机构 * Centre for Language Studies(语言研究中心)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments Accepted to Interspeech 2025. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants which is financed by the Dutch Research Council (NWO)

Journal ref https://www.isca-archive.org/interspeech_2025/parikh25_interspeech.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18765 2025-08-28 cs.LG 57%

Governance-as-a-Service: A Multi-Agent Framework for AI System Compliance and Policy Enforcement

Suyash Gaurav, Jukka Heikkonen, Jatin Chaudhary

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13042 2025-08-27 cs.CY 57%

How Do AI Companies "Fine-Tune" Policy? Examining Regulatory Capture in AI Governance

Kevin Wei, Carson Ezell, Nick Gabrieli, Chinmay Deshpande

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 39 pages (14 pages main text), 3 figures, 9 tables. To be published in the Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, & Society (AIES)

Journal ref Proc. AAAI/ACM Conf. AI, Ethics & Soc., 7 (2024) 1539-1555

详情

展开后加载摘要…

URL PDF HTML 收藏