arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-03-12 至 2026-03-12 共收录 66 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 20 篇

2511.09394 2026-03-12 cs.HC 67%

EyeAgent: An Agentic AI System for Multimodal Clinical Decision Support in Ophthalmology

EyeAgent:一种用于眼科多模态临床决策支持的智能AI系统

Danli Shi, Xiaolan Chen, Bingjie Yan, Weiyi Zhang, Pusheng Xu, Jiancheng Yang, Ruoyu Chen, Siyu Huang, Bowen Liu, Xinyuan Wu, Meng Xie, Ziyu Gao, Yue Wu, Senlin Lin, Kai Jin, Xia Gong, Yih Chung Tham, Xiujuan Zhang, Li Dong, Yuzhou Zhang, Jason Yam, Guangming Jin, Xiaohu Ding, Haidong Zou, Yalin Zheng, Zongyuan Ge, Mingguang He

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

AI总结 EyeAgent通过整合53种眼科工具提升诊断准确率,实现多模态临床决策支持,为眼科AI系统提供新范式。

Comments 28 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10303 2026-03-12 cs.CL cs.AI 62%

Is this Idea Novel? An Automated Benchmark for Judgment of Research Ideas

这个想法是否新颖?一种自动化的判断研究想法新颖性的基准

Tim Schopf, Michael Färber

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出RINoBench基准,用于评估研究想法新颖性判断,发现LLM生成的推理与人类相似但判断不准确。

Comments Accepted to LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05941 2026-03-12 cs.LG cs.AI 62%

Comparative Analysis of Modern Machine Learning Models for Retail Sales Forecasting

现代机器学习模型在零售销售预测中的比较分析

Luka Hobor, Mario Brcic, Lidija Polutnik, Ante Kapetanovic

机构 * Faculty of Electrical Engineering and Computing, University of Zagreb(Zagreb大学电气工程与计算学院) It From Bit d.o.o.(It From Bit公司) Babson College(巴布森学院) mStart Plus d.o.o.(mStart Plus公司)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文比较了现代机器学习模型在零售销售预测中的表现,发现基于树的集成方法在特定条件下表现更优。

Comments 12 total pages, 12 pages article

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03281 2026-03-12 cs.CV cs.LG 57%

CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance

CFG-Ctrl: 基于控制的分类器自由扩散引导

Hanyang Wang, Yiyang Liu, Jiawei Chi, Fangfu Liu, Ran Xue, Yueqi Duan

专题命中 安全评测 :alignment(abstract);分类 cs.LG

AI总结 CFG-Ctrl通过引入滑模控制方法改进分类器自由扩散引导,提升语义对齐和鲁棒性。

Comments Accepted by CVPR 2026; Project Page: https://hanyang-21.github.io/CFG-Ctrl

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10764 2026-03-12 cs.CL 57%

HeartAgent: An Autonomous Agent System for Explainable Differential Diagnosis in Cardiology

HeartAgent:一种用于心血管疾病可解释性差分诊断的自主代理系统

Shuang Zhou, Kai Yu, Song Wang, Wenya Xie, Zaifu Zhan, Meng-Han Tsai, Yuen-Hei Chung, Shutong Hou, Huixue Zhou, Min Zeng, Bhavadharini Ramu, Lin Yee Chen, Feng Xie, Rui Zhang

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

AI总结 HeartAgent通过整合定制工具和数据资源,提供可解释的差分诊断支持,显著提升心血管疾病诊断的准确性和解释质量。

Comments 26 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10703 2026-03-12 cs.CV cs.CY 57%

WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation

WalkGPT: 基于深度感知的视觉语言对话用于行人导航

Rafi Ibn Sultan, Hui Zhu, Xiangyu Zhou, Chengyin Li, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu

专题命中 安全评测 :alignment(abstract);分类 cs.CY

AI总结 WalkGPT通过结合视觉语言模型与深度感知分割技术,实现无障碍导航指导,提供精准的可访问性评估和深度估计。

Comments Accepted by CVPR-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10294 2026-03-12 eess.SY cs.AI cs.SY 57%

Simulation-in-the-Reasoning (SiR): A Conceptual Framework for Empirically Grounded AI in Autonomous Transportation

推理中的模拟(SiR):面向自主交通系统经验导向AI的概念框架

Wuping Xin

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 SiR通过将领域模拟器嵌入LLM推理循环,为自主交通系统提供经验导向的AI框架,提升策略验证与优化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09689 2026-03-12 cs.CV cs.AI 57%

AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering

AutoViVQA:一个大规模自动构建的数据集用于越南语视觉问答

Nguyen Anh Tuong, Phan Ba Duc, Nguyen Trung Quoc, Tran Dac Thinh, Dang Duy Lan, Nguyen Quoc Thinh, Tung Le

机构 * Faculty of Information Technology, University of Science, VNU-HCM(越南国家大学胡志明市分校信息科技学院) Vietnam National University, Ho Chi Minh City(越南国家大学胡志明市分校)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 本文提出AutoViVQA数据集,利用变压器架构进行越南语视觉问答,结合文本和视觉预训练,并在多语言环境下系统比较自动评估指标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03507 2026-03-12 cs.LG cond-mat.dis-nn q-bio.NC stat.ML 57%

Solving adversarial examples requires solving exponential misalignment

解决对抗样本需要解决指数级不一致

Alessandro Salvatore, Stanislav Fort, Surya Ganguli

专题命中 安全评测 :alignment(abstract);分类 cs.LG

AI总结 研究通过分析神经网络的感知流形维度,揭示对抗样本的起源与机器与人类感知的指数级不一致,并提出维度一致是提升对抗鲁棒性的关键。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00353 2026-03-12 cs.AI cs.HC 57%

MHDash: An Online Platform for Benchmarking Mental Health-Aware AI Assistants

MHDash:一个用于评估心理健康意识AI助手的在线平台

Yihe Zhang, Cheyenne N Mohawk, Kaiying Han, Vijay Srinivas Tida, Manyu Li, Xiali Hei

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 MHDash是一个开源平台,用于评估心理健康意识AI助手的安全性和性能,揭示传统基准在心理健康领域不足的问题。

Comments Accepted for presentation at IEEE SoutheastCon 2026. This is the author version of an accepted paper. The final version will appear in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05748 2026-03-12 cs.LG 57%

Communication Enables Cooperation in LLM Agents: A Comparison with Curriculum-Based Approaches

通信使LLM代理协作:与基于课程的方法的比较

Hachem Madmoun, Salem Lahlou

机构 * MBZUAI(穆扎布伊人工智能研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

AI总结 本文通过比较直接通信与课程学习方法,发现通信在促进LLM代理合作中更有效,而课程学习可能因设计选择影响对齐目标。

Comments Published in EACL 2026 - Corrected cooperation rates for two-stage communication conditions (96.7% and 100.0%, previously reported as 48.3% and 50.0% due to a denominator bug in the evaluation code). All other results unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10940 2026-03-12 cs.SE cs.RO 50%

STADA: Specification-based Testing for Autonomous Driving Agents

STADA:基于规范的自动驾驶代理测试

Joy Saha, Trey Woodlief, Sebastian Elbaum, Matthew B. Dwyer

专题命中 安全评测 :safety(abstract)

AI总结 STADA是一种基于规范的自动驾驶代理测试生成框架,通过形式规范生成多样化的测试场景,提高测试覆盖率并减少模拟次数。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11099 2026-03-12 cs.CV 50%

A Survey on Interpretability in Visual Recognition

视觉识别中可解释性的综述

Qiyang Wan, Chengzhi Gao, Ruiping Wang, Xilin Chen

专题命中 安全评测 :safety(abstract)

AI总结 本文综述了视觉识别中可解释性的发展,从意图、对象、呈现和方法学角度建立多维分类法,总结了关键评估指标,并探讨了多模态大语言模型的可解释性及实际应用。

Comments 20 pages, 8 figures, 7 tables. Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 5 篇

2603.10351 2026-03-12 cs.CL cs.AI 62%

Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck

通过解耦信息瓶颈缓解多语言大语言模型作为评判者的翻译偏差

Hongbin Zhang, Kehai Chen, Xuefen Bai, Youcheng Pan, Yang Xiang, Jinpeng Wang, Min Zhang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出DIBJudge框架,通过解耦信息瓶颈缓解多语言LLM的翻译偏差问题,有效提升多语言评估的公平性和准确性。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00307 2026-03-12 cs.AI 57%

BiasBusters: Uncovering and Mitigating Tool Selection Bias in Large Language Models

BiasBusters: 检测和缓解大语言模型中的工具选择偏差

Thierry Blankenstein, Jialin Yu, Zixuan Li, Vassilis Plachouras, Sunando Sengupta, Philip Torr, Yarin Gal, Alasdair Paren, Adel Bibi

机构 * University of Oxford(牛津大学) Microsoft(微软)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 研究通过基准测试揭示LLM工具选择中的系统性偏差,并提出轻量级缓解策略以减少选择偏差。

Comments ICLR 2026 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22699 2026-03-12 cs.CL 57%

Are you sure? Measuring models bias in content moderation through uncertainty

你确定吗?通过不确定性测量内容审核中的模型偏差

Alessandra Urbinati, Mirko Lai, Simona Frenda, Marco Antonio Stranisci

机构 * Laboratory for the Modeling of Biological and Socio-technical Systems, Northeastern University(生物与社会技术系统建模实验室,东北大学) Heriot-Watt University(赫瑞-瓦特大学) aequa-tech(aequa-tech公司) Università del Piemonte Orientale(皮埃蒙特东方大学) Università degli Studi di Torino(托里尼大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

AI总结 本文提出通过模型预测不确定性来衡量内容审核中模型的偏差,揭示预训练模型对少数群体的预测准确性与置信度的差异,以改进模型公平性。

Comments accepted at Findings of ACL: EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10028 2026-03-12 cs.CY cs.AI 54%

How to Count AIs: Individuation and Liability for AI Agents

如何计算AI:AI代理的个体化与责任归属

Yonathan Arbel, Peter Salib, Simon Goldstein

机构 * University of Alabama(阿拉巴马大学) University of Hong Kong(香港大学) University of Houston(休斯顿大学) HKU AI & Humanity Lab(香港大学AI与人类实验室) Center for Law & AI Risk(法律与人工智能风险中心) Institute for Law & AI(法律与人工智能研究所)

专题命中 AI治理与伦理 :分类 cs.AI、cs.CY;safety(comments);AI safety(comments)

AI总结 本文提出通过‘算法公司’解决AI代理的个体识别与责任归属问题,通过法律虚构实体实现对AI行为的追踪与责任划分。

Comments 36 pages. Presented at the Law Following AI conference, Cambridge University. Interdisciplinary: AI safety, AI governance, legal theory

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10773 2026-03-12 cs.HC 50%

AI-Generated Rubric Interfaces: K-12 Teachers' Perceptions and Practices

AI生成的评分标准界面:K-12教师的感知与实践

Bahare Riahi, Sayali Patukale, Joy Niranjan, Yogya Koneru, Tiffany Barnes, Veronica Cateté

专题命中 AI治理与伦理 :alignment(abstract)

AI总结 本研究探讨了K-12教师在使用AI生成评分标准时的感知与实践,发现AI生成的评分标准在结构和清晰度上有帮助,但需要教师监督以确保准确性和相关性。

Comments 20 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 18 篇

2603.10243 2026-03-12 cs.CL 88%

GR-SAP: Generative Replay for Safety Alignment Preservation during Fine-Tuning

GR-SAP: 生成性重放用于在微调过程中保持安全对齐

Zhouxiang Fang, Jiawei Zhou, Hanjie Chen

专题命中 其他安全 :alignment(title,abstract);safety(title,abstract);分类 cs.CL

AI总结 GR-SAP通过生成领域特定对齐数据来保持微调过程中的安全对齐,有效缓解了微调导致的安全退化问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09997 2026-03-12 cs.CL cs.AI cs.CY cs.HC 83%

Empathy Is Not What Changed: Clinical Assessment of Psychological Safety Across GPT Model Generations

共情并未改变:对心理安全的临床评估跨GPT模型世代

Michael Keeman, Anastasia Keeman

机构 * Keido Labs(Keido实验室)

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 研究通过临床评估发现,GPT模型在心理安全方面存在变化,尽管共情评分无显著差异,但危机检测能力提升而建议安全下降,揭示了模型在对话中间阶段的显著变化。

Comments 17 pages, 7 figures. First empirical measurement of the #keep4o phenomenon using clinical psychological safety frameworks. Compares GPT-4o, o4-mini, and GPT-5-mini on empathy, crisis detection, and advice safety dimensions

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08017 2026-03-12 cs.HC cs.AI 79%

Alignment-Process-Outcome: Rethinking How AIs and Humans Collaborate

对齐-过程-结果:重新思考AI和人类如何协作

Haichang Li, Anjun Zhu, Arpit Narechania

机构 * George Mason University(乔治·马歇尔大学) Simon Fraser University(西蒙·弗雷泽大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出通过任务和意图两个视角重新理解AI与人类协作的结构关系,揭示对齐、过程和结果之间的动态联系。

Comments Accepted by Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA 26), Barcelona, Spain, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10682 2026-03-12 cs.RO 78%

OnFly: Onboard Zero-Shot Aerial Vision-Language Navigation toward Safety and Efficiency

OnFly:面向安全与效率的机载零样本空中视觉语言导航

Guiyong Zheng, Yueting Ban, Mingjie Zhang, Juepeng Zheng, Boyu Zhou

机构 * School of Artificial Intelligence, Sun Yat-Sen University(中山大学人工智能学院) Southern University of Science and Technology(南方科技大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 其他安全 :safety(title,abstract)

AI总结 OnFly通过共享感知双智能体架构和混合记忆机制,实现高效稳定的零样本空中视觉语言导航,提升安全性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09566 2026-03-12 cs.AI cs.LG 76%

Toward Closed-loop Molecular Discovery via Language Model, Property Alignment and Strategic Search

通过语言模型、性质对齐和战略搜索实现闭环分子发现

Junkai Ji, Zhangfan Yang, Dong Xu, Ruibin Bai, Jianqiang Li, Tingjun Hou, Zexuan Zhu

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

AI总结 Trio结合语言模型、性质对齐和战略搜索,实现高效且可解释的闭环分子设计,提升药物配体的结合亲和力、药物性和合成可及性。

Comments 30 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09987 2026-03-12 cs.CL cs.AI cs.LG 67%

Evolving Demonstration Optimization for Chain-of-Thought Feature Transformation

链式推理特征转换的进化演示优化

Xinyuan Wang, Kunpeng Liu, Arun Vignesh Malarkkan, Yanjie Fu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出一种通过进化轨迹经验优化LLM驱动的特征转换方法,提升转换多样性与下游任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10877 2026-03-12 cs.CL 57%

From Images to Words: Efficient Cross-Modal Knowledge Distillation to Language Models from Black-box Teachers

从图像到词语:高效的跨模态知识蒸馏用于从黑盒教师模型向语言模型转移

Ayan Sengupta, Shantanu Dixit, Md Shad Akhtar, Tanmoy Chakraborty

机构 * Indian Institute of Technology Delhi, India(印度德里印度理工学院) Indraprastha Institute of Information Technology Delhi, India(印度德里印度信息科技学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本文提出ARMADA框架,通过高效的跨模态知识蒸馏方法,从黑盒视觉-语言模型向语言模型转移知识,实现性能提升且无需昂贵预训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10052 2026-03-12 cs.RO cs.LG 57%

OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot Policies

OmniGuide: 通用引导场用于增强通用机器人策略

Yunzhou Song, Long Le, Yong-Hyun Park, Jie Wang, Junyao Shi, Lingjie Liu, Jiatao Gu, Eric Eaton, Dinesh Jayaraman, Kostas Daniilidis

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 OMNIGUIDE通过整合多种引导源提升通用机器人策略在复杂任务中的性能

Comments Project Page: $\href{https://omniguide.github.io/}{this\; url}$

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02816 2026-03-12 cs.CV cs.AI 57%

BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation

BrandFusion: 一种多智能体框架,用于文本到视频生成中的无缝品牌整合

Zihao Zhu, Ruotong Wang, Siwei Lyu, Min Zhang, Baoyuan Wu

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Loop Area Institute(深圳河套学院) State University of New York at Buffalo(纽约州立大学布法罗分校) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 BrandFusion通过多智能体框架实现文本到视频生成中的无缝品牌整合,提升品牌辨识度和内容自然度,推动T2V的可持续商业化应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24224 2026-03-12 cs.LG 57%

Proposing a Framework for Machine Learning Adoption on Legacy Systems

为遗留系统采用机器学习提出一个框架

Ashiqur Rahman, Hamed Alhoori

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本文提出一个基于API的框架,通过解耦ML模型生命周期与生产环境,为遗留系统提供轻量级浏览器界面,减少升级成本和生产中断,提升制造业生产质量与安全。

Comments Accepted at The First International Workshop on Resilient Artificial Intelligence for Manufacturing (ICDM'25)

Journal ref 2025 IEEE International Conference on Data Mining Workshops (ICDMW), Washington, DC, USA, 2025, pp. 1616-1623

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04233 2026-03-12 cs.LG cs.SI 57%

Graph machine learning for flight delay prediction due to holding manouver

基于图机器学习的飞行延误预测:因等待 maneuver 引起的延误

Jorge L. Franco, Manoel V. Machado Neto, Filipe A. N. Verri, Diego R. Amancio

机构 * Institute of Mathematics and Computer Science, University of São Paulo(圣保罗大学数学与计算机科学研究所) Aeronautical Technology Institute – ITA(航空技术研究所 – ITA)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本研究利用图机器学习方法预测因等待 maneuver 引起的飞行延误,通过CatBoost和GATs模型提升预测精度并提供可解释性,以提高航空运营效率。

Journal ref Physica A 685, 131318 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10862 2026-03-12 cs.SD 50%

OSUM-Pangu: An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUs

OSUM-Pangu:基于OpenPangu在Ascend NPUs上的开源多维语音理解基础模型

Yujie Liao, Xuelong Geng, Hongfei Xue, Shuiyuan Wang, Lei Xie

机构 * Ascend NPUs(Ascend NPU)

专题命中 其他安全 :alignment(abstract)

AI总结 OSUM-Pangu基于非CUDA的Ascend NPU平台,结合openPangu-7B LLM后端,实现语音理解任务的高准确率与自然语言交互能力。

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏