arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9324 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9324 篇

2605.12729 2026-06-17 cs.NI cs.AI cs.CR 版本更新 74%

Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

用于代理网络运维和AI运维的大型语言模型:架构、评估与安全

Muhammad Bilal, Jon Crowcroft, Ruizhi Wang, Xiaolong Xu, Schahram Dustdar

机构 * School of Computing and Communications(计算与通信学院) University of Cambridge(剑桥大学) School of Software(软件学院) Nanjing University of Information Science and Technology(南京信息科技大學) TU Wien(维也纳技术大学) ICREA

专题命中 安全评测 :safety(title);分类 cs.AI

AI总结 本文探讨了大型语言模型在网络运维和AI运维中的应用,分析了代理架构、评估方法及安全挑战,强调系统可靠性依赖于模型周边机制,而非模型本身。

Comments 49 pages, 15 figures, 6 tables; survey article

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19662 2026-06-09 cs.AI 版本更新 74%

When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment Approach

当表格基础模型遇见策略性表格数据:一种先验对齐方法

Xinpeng Lv, Yunxin Mao, Renzhe Xu, Chunyuan Zheng, Yikai Chen, Haoxuan Li, Jinxuan Yang, Kun Kuang, Yuanlong Chen, Mingyang Geng, Wanrong Huang, Shixuan Liu, Shaowu Yang, Wenjing Yang, Zhouchen Lin, Haotian Wang

机构 * University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学)

专题命中 安全评测 :alignment(title);分类 cs.AI

AI总结 本文研究了表格基础模型在策略性表格数据上的泛化能力,提出了一种策略感知的先验对齐框架SPN,以提高模型在策略性环境中的鲁棒性和预测性能。

Comments Accepted by ICML2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10234 2026-06-02 cs.HC cs.AI 74%

InFerActive: Interactive Tree-Based Exploration of LLM Sampling for Safety Evaluation

InFerActive: 基于交互式树的安全评估中LLM采样探索

Junhyeong Hwangbo, Soohyun Lee, Hyeon Jeon, Kyochul Jang, Minsoo Cheong, Youngjae Yu, Jinwook Seo

机构 * Seoul National University(首尔国立大学)

专题命中 安全评测 :safety(title);分类 cs.AI

AI总结 提出InFerActive系统,通过广度优先采样构建可导航短语树,提升LLM安全评估中低概率有害输出的覆盖率和效率,相比随机采样减少5倍样本量。

Comments v2: Revised version

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23085 2026-05-29 cs.AI 74%

When Models Learn to Ask Why: Adaptive Causal Reasoning for Trustworthy Medical Vision-Language Models

当模型学会问为什么:面向可信医疗视觉语言模型的自适应因果推理

Jianxin Lin, Chunzheng Zhu, Peter J. Kneuertz, Yunfei Bai, Yuan Xue

机构 * The Ohio State University(俄亥俄州立大学) Hunan University(湖南大学) Amazon(亚马逊)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 提出MedCausalX框架,通过因果推理链、自适应反射架构和轨迹级因果校正,解决医疗VLM中的虚假相关和推理不一致问题,显著提升诊断一致性和减少幻觉。

Comments Accepted by CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28188 2026-05-28 cs.CL 74%

Framing Matters: Addressing Framing Sensitivity in Decision-Making through Behaviorally-Grounded Value Alignment

框架至关重要:通过基于行为的价值对齐解决决策中的框架敏感性

Seojin Hwang, Minju Kim, Junhyuk Choi, JeongHyun Park, Hwanhee Lee

机构 * Chung-Ang University(Chung-Ang 大学)

专题命中 安全评测 :alignment(title);分类 cs.CL

AI总结 本文提出Fragile基准测试框架,系统评估大语言模型在事实等价但不同框架输入下的决策稳定性,并设计Valign方法通过表示级干预有效降低框架引起的决策翻转。

Comments 29 pages, 7 figures, 31 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27345 2026-05-27 cs.CL 74%

MATCHA: Matching Text via Contrastive Semantic Alignment

MATCHA: 通过对比语义对齐进行文本匹配

Siran Li, Ece Sena Etoglu, Carsten Eickhoff, Seyed Ali Bahrainian

机构 * University of Tübingen(图宾根大学)

专题命中 安全评测 :alignment(title);分类 cs.CL

AI总结 针对现有评估指标无法区分语义矛盾的问题,提出MATCHA指标,通过双视角对比学习同时奖励语义一致性和惩罚矛盾,在多个基准上优于ROUGE和BERTScore。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26530 2026-05-27 cs.AI 74%

Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning

哪些变化重要?通过相关性敏感评估和求解器基础推理实现可信赖的法律AI

Chen Linze, Cai Yufan, Hou Zhe, Dong Jin Song

机构 * National University of Singapore(新加坡国立大学) Griffith University(格里菲斯大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 提出法律相关性敏感评估问题,引入统一评估套件,并设计基于形式推理的对抗多智能体框架LexGuard,以提高法律AI对法律相关变化的校准敏感性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11296 2026-05-26 cs.CV cs.LG 74%

$Δ\mathrm{Energy}$: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization

$Δ\mathrm{Energy}$: 优化视觉-语言对齐过程中的能量变化提升OOD检测与OOD泛化

Lin Zhu, Yifeng Yang, Xinbing Wang, Qinying Gu, Nanyang Ye

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 安全评测 :alignment(title);分类 cs.LG

AI总结 本文提出ΔEnergy分数,通过重新对齐视觉-语言模态时的能量变化来同时提升分布外检测和分布外泛化性能,并基于此开发了统一微调框架EBM。

Comments Accepted by NeurIPS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23297 2026-05-25 cs.AI cs.DC 74%

Ontological Knowledge Blocks: Executable Compliance and Profile-Based Validation for Trustworthy AI Systems

本体知识块:可信AI系统的可执行合规与基于配置文件的验证

Aasish Kumar Sharma, Julian M. Kunkel

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 提出本体知识块(OKBs)作为可编程治理基础设施,通过将监管义务编译为机器可检查的约束,实现自动化合规验证,并在AI辅助HPC资源分配场景中验证其有效性。

Comments 6 pages, 3 figures. Accepted at the Security, Trust and Privacy for Software and Applications (STPSA) Workshop, IEEE COMPSAC 2026, Madrid, Spain, July 7-10, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20665 2026-05-22 cs.CV cs.AI 74%

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm

视见之代价:在单体范式内实现可信的多模态推理

Karan Goyal

机构 * IIIT Delhi, India(德里印度理工学院)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 本文提出了一种新的多模态评估方法,通过信息论视角揭示了多模态推理中的视见代价问题,提出了三个新指标并提出了语义充分性准则,挑战了传统多模态评估方法。

Comments Addresses practical viability of Vlabel construction. Writing is grounded. Acknowledgement is duly added

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27389 2026-05-14 cs.CV cs.AI 74%

COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts

COHERENCE:在交错多模态上下文中细粒度图像-文本对齐的基准测试

Bingli Wang, Huanze Tang, Haijun Lv, Zhishan Lin, Lixin Gu, Lei Feng, Qipeng Guo, Kai Chen

机构 * Southeast University Shanghai AI Laboratory(上海大学上海人工智能实验室) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 安全评测 :alignment(title);分类 cs.AI

AI总结 COHERENCE旨在评估MLLM在交错多模态上下文中恢复细粒度图像-文本对应关系的能力,涵盖四个领域,包含6161个高质量问题,并进行六类错误分析以识别当前MLLM的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07422 2026-05-11 cs.SE cs.AI 74%

Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study

基于LLM的软件工程社区心理安全定性编码的提示工程策略:一项受控实证研究

Moaath Alshaikh, Tasneem Alshaher, Ricardo Vieira, Beatriz Santana, Clelio Xavier, Jose Amancio, Glauco Carneiro, Julio Leite, Savio Freire, Manoel Mendonca

机构 * Federal University of Bahia(巴伊亚联邦大学) State University of Feira de Santana(费拉德桑塔纳州立大学) Federal University of Sergipe(塞格皮联邦大学) Federal Institute of Ceara(塞阿拉联邦理工学院)

专题命中 安全评测 :safety(title);分类 cs.AI

AI总结 本文通过对比三种LLM在零样本和多样本提示策略下的表现,验证了多样本提示能提升编码一致性,但对部分模型效果有限,同时发现模型在某些类别上存在系统性偏差。

Comments 9 pages, 5 figures. Accepted at the 1st International Workshop on Prompt Engineering for Software Engineering (PROMPT-SE 2026), co-located with the 30th International Conference on Evaluation and Assessment in Software Engineering (EASE 2026), Glasgow, Scotland, United Kingdom, June 9--12, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.04491 2026-05-11 cs.CY cs.CR 74%

An Evaluation of Chat Safety Moderations in Roblox

在Roblox中聊天安全审核的评估

Priya Kaushik, Sonja Brown, Rakibul Hasan, Sazzadur Rahaman

专题命中 安全评测 :safety(title);分类 cs.CY

AI总结 本文通过分析200万条聊天记录,评估Roblox聊天审核系统的效果,发现大量不安全内容如性侵、欺凌等逃过了审核,并揭示了用户如何通过多种手段规避审核。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15949 2026-05-04 cs.CL 74%

BanglaSocialBench: A Benchmark for Evaluating Sociopragmatic and Cultural Alignment of LLMs in Bangladeshi Social Interaction

BanglaSocialBench: 一个评估大语言模型在孟加拉社会互动中社会语用和文化对齐的基准

Tanvir Ahmed Sijan, S. M Golam Rifat, Pankaj Chowdhury Partha, Md. Tanjeed Islam, Md. Musfique Anwar

机构 * Jahangirnagar University(贾亨吉尔纳加尔大学) Rajshahi University of Engineering & Technology(拉贾沙希工程与技术大学) Bangladesh University of Engineering and Technology(孟加拉工程与技术大学)

专题命中 安全评测 :alignment(title);分类 cs.CL

AI总结 本文提出BanglaSocialBench,通过情境依赖的语言使用评估孟加拉语的社会语用能力,发现当前LLM在文化对齐上存在系统性偏差。

Comments Accepted at ACL SRW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21263 2026-04-24 cs.AI cs.PL cs.SE q-bio.QM 74%

Trustworthy Clinical Decision Support Using Meta-Predicates and Domain-Specific Languages

可信的临床决策支持使用元谓词和领域特定语言

Michael Bouzinier, Sergey Trifonov, Michael Chumack, Eugenia Lvova, Dmitry Etin

机构 * Harvard University(哈佛大学) IDEXX Laboratories(IDEXX实验室) Forome Association(Forome协会) Deggendorf Institute of Technology(德格多夫技术学院)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 本文提出使用元谓词和领域特定语言来确保临床决策规则的证据适当性,通过元谓词验证在部署前约束证据使用,提升决策的可审计性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03686 2026-04-17 cs.AI 74%

AI4S-SDS: A Neuro-Symbolic Solvent Design System via Sparse MCTS and Differentiable Physics Alignment

AI4S-SDS: 一种基于稀疏MCTS和可微物理对齐的神经符号溶剂设计系统

Jiangyu Chen

机构 * State Key Laboratory for Novel Software Technology at Nanjing University(南京大学新型软件技术国家重点实验室) School of Computer Science, Nanjing University(南京大学计算机科学学院) Suzhou Laboratory(苏州实验室)

专题命中 安全评测 :alignment(title);分类 cs.AI

AI总结 本文提出AI4S-SDS系统,通过稀疏MCTS和可微物理对齐,解决化学配方设计中的高维组合空间问题,提升探索多样性并实现物理约束下的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10300 2026-04-14 cs.SE cs.AI 74%

From Helpful to Trustworthy: LLM Agents for Pair Programming

从有益到可信:用于配对编程的LLM代理

Ragib Shahariar Ayon

机构 * Texas State University(德克萨斯州立大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 本文研究多代理LLM配对编程,通过外部化意图和工具迭代验证,提升代码可靠性、可审计性和可维护性。

Comments Accepted in 34th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (FSE Companion 26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07363 2026-04-10 cs.LG 74%

Benchmark Shadows: Data Alignment, Parameter Footprints, and Generalization in Large Language Models

基准阴影:大型语言模型中的数据对齐、参数足迹与泛化

Hongjian Zou, Yidan Wang, Qi Ding, Yixuan Liao, Xiaoxin Chen

机构 * Vivo AI Lab, Shenzhen, China(维沃人工智能实验室,深圳,中国) Hong Kong University of Science and Technology, Hong Kong, China(香港科技大学,香港,中国)

专题命中 安全评测 :alignment(title);分类 cs.LG

AI总结 本文研究了大型语言模型中基准表现与广度能力之间的差异,通过数据干预发现数据分布影响参数适应与泛化能力,提出参数空间诊断方法揭示不同训练模式的结构特征。

Comments 28 pages, 26 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09726 2026-04-08 cs.CL 74%

Forgetting as a Feature: Cognitive Alignment of Large Language Models

作为特征的遗忘:大型语言模型的认知对齐

Alexandros Christoforos

专题命中 安全评测 :alignment(title);分类 cs.CL

AI总结 研究将遗忘视为大型语言模型的认知机制,通过建立基准测试评估时间推理、概念漂移适应和联想回忆,发现模型遗忘率与人类记忆效率的权衡相似,并提出概率记忆提示策略提升长期推理性能。

Comments arXiv admin note: This submission has been withdrawn by arXiv administrators due to incorrect authorship. Author list truncated

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02372 2026-04-06 cs.CR cs.LG 74%

Backdoor Attacks on Decentralised Post-Training

在去中心化后训练中的后门攻击

Oğuzhan Ersoy, Nikolay Blagoev, Jona te Lintelo, Stefanos Koffas, Marina Krček, Stjepan Picek

机构 * University of Neuchâtel(纳沙泰尔大学) Radboud University(拉德堡德大学) Delft University of Technology(代尔夫特理工大学) SecureML University of Zagreb(萨格勒布大学)

专题命中 安全评测 :safety(abstract,comments);alignment(abstract);分类 cs.LG;trustworthy(comments)

AI总结 本文提出首个针对流水线并行的后门攻击,通过控制中间阶段实现模型对齐偏差,实验显示触发词可使对齐率从80%降至6%,且在安全对齐训练下仍成功60%。

Comments Accepted to ICLR 2026 Workshop 'Principled Design for Trustworthy AI - Interpretability, Robustness, and Safety across Modalities'

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01238 2026-04-03 cs.NI cs.AI 74%

Trustworthy AI-Driven Dynamic Hybrid RIS: Joint Optimization and Reward Poisoning-Resilient Control in Cognitive MISO Networks

可信的AI驱动动态混合RIS:在认知MISO网络中的联合优化与抗奖励中毒控制

Deemah H. Tashman, Soumaya Cherkaoui

机构 * LINCS Laboratory, Department of Computer and Software Engineering, Polytechnique Montréal(蒙特利尔理工学院计算机与软件工程系LINCS实验室)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 本文提出一种动态混合RIS,用于解决认知MISO网络中不可靠的直接SU链路和能量约束问题,通过联合优化传输波束成形和RIS相位,采用SAC深度强化学习方法,并提出轻量级实时防御机制以对抗奖励中毒攻击。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20456 2026-04-03 cs.LG 74%

Towards Trustworthy Wi-Fi CSI-based Sensing: Systematic Evaluation of Adversarial Robustness

迈向可信的Wi-Fi CSI基于传感:对抗鲁棒性系统的评估

Shreevanth Krishnaa Gopalakrishnan, Stephen Hailes

机构 * University College London(伦敦大学学院)

专题命中 安全评测 :trustworthy(title);分类 cs.LG

AI总结 本文系统评估了五种不同CSI架构在四个公开数据集上的对抗鲁棒性,发现模型容量并不保证鲁棒性,简单架构比高容量模型更抗攻击,且任务依赖性显著影响脆弱性。

Comments 18 pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24414 2026-03-26 cs.CR cs.AI 74%

ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers

ClawKeeper: 通过技能、插件和观察者为OpenClaw代理提供全面的安全保护

Songyang Liu, Chaozhuo Li, Chenxu Wang, Jinyu Hou, Zejian Chen, Litian Zhang, Zheng Liu, Qiwei Ye, Yiming Hei, Xi Zhang, Zhongyuan Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) China Academy of Information and Communications Technology(信息通信技术研究院)

专题命中 安全评测 :safety(title);分类 cs.AI

AI总结 ClawKeeper通过技能、插件和观察者三层机制,提供实时安全防护,解决OpenClaw生态的安全漏洞问题,经过评估证明其有效性。

Comments 22 pages, 14 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08105 2026-03-09 cs.CL 74%

MERLIN: Multi-Stage Curriculum Alignment for Multilingual Encoder-LLM Integration in Cross-Lingual Reasoning

MERLIN:多阶段课程对齐用于多语言编码器-大语言模型集成的跨语言推理

Kosei Uemura, David Guzmán, Quang Phuoc Nguyen, Jesujoba Oluwadara Alabi, En-shiun Annie Lee, David Ifeoluwa Adelani

机构 * University of Toronto(多伦多大学) Mila-Quebec AI Institute, McGill University(魁北克AI研究所,麦吉尔大学) OntarioTech University(安大略技术大学) Saarland University(萨尔兰州立大学) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 安全评测 :alignment(title);分类 cs.CL

AI总结 MERLIN通过多阶段课程对齐方法提升多语言编码器-大语言模型在跨语言推理中的性能,特别是在低资源语言上表现突出。

Comments Accepted to EACL 2026 (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16235 2026-03-03 cs.CV cs.AI 74%

AI-Powered Dermatological Diagnosis: From Interpretable Models to Clinical Implementation A Comprehensive Framework for Accessible and Trustworthy Skin Disease Detection

AI赋能的皮肤病诊断:从可解释模型到临床实施 一个全面的框架,用于可访问且可信的皮肤疾病检测

Satya Narayana Panda, Vaishnavi Kukkala, Spandana Iyer

机构 * Department of Business Analytics/Data Science Engineering University of New Haven(商业分析/数据科学工程系 罗德岛大学) Department of Healthcare University of New Haven(医疗健康系 罗德岛大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 本研究提出了一种结合家族史数据和临床影像的AI框架,旨在提升皮肤病诊断的准确性与临床应用的可行性。

Comments 9 pages, 5 figures, 1 table. Code available at https://github.com/colabre2020/Enhancing-Skin-Disease-Diagnosis

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08086 2026-02-10 cs.LG 74%

Probability Hacking and the Design of Trustworthy ML for Signal Processing in C-UAS: A Scenario Based Method

概率黑客与用于无人机系统(C-UAS)信号处理的可信机器学习设计:基于场景的方法

Liisa Janssens, Laura Middeldorp

专题命中 安全评测 :trustworthy(title);分类 cs.LG

AI总结 本文提出了一种基于场景的方法,用于增强C-UAS的信号处理能力,通过识别法律机制中的要求来防止概率黑客,从而提升系统的可信度。

Comments 6 pages, Pre-publication. Copyright 2026 IEEE. Peer Reviewed. Accepted at ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), scheduled for 3-8 May 2026 in Barcelona, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20658 2026-02-04 cond-mat.mtrl-sci cs.LG 74%

Trustworthy AI-based crack-tip segmentation using domain-guided explanations

基于领域引导解释的可信AI裂缝尖端分割

Jesco Talies, Eric Breitbarth, David Melching

机构 * Institute for Frontier Materials on Earth and in Space, German Aerospace Center (DLR)(地球和空间前沿材料研究所,德国航空航天中心(DLR))

专题命中 安全评测 :trustworthy(title);分类 cs.LG

AI总结 本文提出一种结合可解释AI和领域先验知识的注意力引导训练框架,用于提升裂缝尖端分割任务的模型泛化能力和解释可信度。

Comments This is the Accepted Manuscript version of an article accepted for publication in Machine Learning: Science and Technology. IOP Publishing Ltd is not responsible for any errors or omissions in this version of the manuscript or any version derived from it. The Version of Record is available online at https://doi.org/10.1088/2632-2153/ae3660

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22653 2026-02-02 cs.HC cs.AI cs.CR 74%

Human-Centered Explainability in AI-Enhanced UI Security Interfaces: Designing Trustworthy Copilots for Cybersecurity Analysts

面向AI增强的用户界面安全性的以人为本的可解释性:为网络安全分析师设计可信的copilot

Mona Rajhans

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 本文提出了一种面向AI增强用户界面安全性的以人为本的可解释性方法,通过设计可信的copilot,提升网络安全分析师对AI输出的信任和决策准确性。

Comments To appear in IEEE ICCA 2025 proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21365 2026-01-30 cs.CY 74%

Small models, big threats: Characterizing safety challenges from low-compute AI models

小模型,大威胁:从低计算量AI模型中characterizing安全挑战

Prateek Puri

专题命中 安全评测 :safety(title);分类 cs.CY

AI总结 本文指出低计算量AI模型因性能提升而变得危险,需加强安全防护以应对新兴威胁。

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18464 2026-01-27 cs.CV cs.AI 74%

Fair-Eye Net: A Fair, Trustworthy, Multimodal Integrated Glaucoma Full Chain AI System

Fair-Eye Net:一种公平、可信的多模态集成青光眼全流程人工智能系统

Wenbin Wei, Suyuan Yao, Cheng Huang, Xiangyu Gao

专题命中 安全评测 :trustworthy(title);分类 cs.AI

AI总结 Fair-Eye Net通过多模态融合和公平性优化,实现青光眼全流程AI系统,提升诊断准确性与公平性,助力全球眼科健康公平。

详情

展开后加载摘要…

URL PDF HTML 收藏