arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-02-27 至 2026-02-27 共收录 17 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 17 篇

2602.22973 2026-02-27 cs.AI 83%

Modeling Expert AI Diagnostic Alignment via Immutable Inference Snapshots

通过不可变推理快照建模专家AI诊断对齐

Dimitrios P. Panagoulias, Evangelia-Aikaterini Tsichrintzi, Georgios Savvidis, Evridiki Tsoureli-Nikita

机构 * organization= Department of Informatics, University of Piraeus , addressline= Karaoli ke Dimitriou 80 , city= Piraeus , postcode= 18534 , country= Greece organization= Department of Research \& Development, Noetiv PC , addressline= Valaoritou 18 , city= Athens , postcode= 10671 , country= Greece organization= Department of Dermatology, Dermacen SA , addressline= Valaoritou 18 , city= Athens , postcode= 10671 , country= Greece

专题命中 安全评测 :alignment(title,abstract);safety(abstract);分类 cs.AI

AI总结 通过不可变推理快照建模专家AI诊断对齐,提升临床决策支持系统评估的准确性与可追溯性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22298 2026-02-27 cs.LG cs.AI 81%

AviaSafe: A Physics-Informed Data-Driven Model for Aviation Safety-Critical Cloud Forecasts

AviaSafe: 一种融合物理约束的数据驱动模型用于航空安全关键的云预报

Zijian Zhu, Qiusheng Huang, Anboyu Guo, Xiaohui Zhong, Hao Li

机构 * Artificial Intelligence Innovation and Incubation Institute, Fudan University(复旦大学人工智能创新与孵化院) Shanghai Innovation Institute(上海创新研究院) Shanghai Academy of Artificial Intelligence for Science(上海人工智能科学研究院) National Marine Environment Forecasting Center(国家海洋环境预报中心) Department of Atmospheric and Oceanic Sciences, Fudan University(复旦大学大气科学与海洋科学系)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

AI总结 AviaSafe通过融合物理约束的神经网络,实现了对航空安全关键的云物种的7天预测,提升了航空路线优化和结冰风险评估能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23075 2026-02-27 cs.CL cs.IR 79%

CiteLLM: An Agentic Platform for Trustworthy Scientific Reference Discovery

CiteLLM:一个用于可信科学引文发现的代理平台

Mengze Hong, Di Jiang, Chen Jason Zhang, Zichang Guo, Yawen Li, Jun Chen, Shaobo Cui, Zhiyang Su

机构 * Hong Kong Polytechnic University(香港理工大学) Beijing University of Posts and Telecommunications(北京邮电大学) Swiss Federal Technology Institute of Lausanne (EPFL)(洛桑联邦理工学院) Hong Kong University of Science and Technology (HKUST)(香港科学大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL

AI总结 CiteLLM通过在LaTeX编辑器中嵌入LLM工具,实现可信的科学引文发现,确保引文的准确性和可靠性。

Comments Accepted by TheWebConf 2026 Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04403 2026-02-27 cs.CV cs.CL cs.CR 79%

Self-adaptive Dataset Construction for Real-World Multimodal Safety Scenarios

自适应数据集构建用于现实世界多模态安全场景

Jingen Qu, Lijun Li, Bo Zhang, Yichen Yan, Jing Shao

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

AI总结 本文提出了一种图像导向的自适应数据集构建方法,用于构建现实世界多模态安全场景的数据集,通过生成35,000个图像-文本对及其指导响应,提升了安全评估的有效性。

Comments Accepted at EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08781 2026-02-27 cs.AI cs.CL 73%

Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions

评估评估者:衡量LLMs对任务评估指令的遵守情况

Bhuvanashree Murugadoss, Christian Poelitz, Ian Drosos, Vu Le, Nick McKenna, Carina Suzana Negreanu, Chris Parnin, Advait Sarkar

专题命中 安全评测 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI

AI总结 本文研究LLMs在任务评估中的表现,发现提示对评估结果影响有限,困惑度有时比提示更符合人类判断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22828 2026-02-27 cs.CL cs.AI 62%

TCM-DiffRAG: Personalized Syndrome Differentiation Reasoning Method for Traditional Chinese Medicine based on Knowledge Graph and Chain of Thought

基于知识图谱和推理链的TCM-DiffRAG:一种用于传统中医个性化辨证的推理方法

Jianmin Li, Ying Chang, Su-Kit Tang, Yujia Liu, Yanwen Wang, Shuyuan Lin, Binkai Ou

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 TCM-DiffRAG结合知识图谱与推理链,提升中医个性化诊断性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23330 2026-02-27 cs.AI q-fin.TR 57%

Toward Expert Investment Teams:A Multi-Agent LLM System with Fine-Grained Trading Tasks

迈向专家投资团队:一个具有细粒度交易任务的多智能体LLM系统

Kunihiro Miyazaki, Takanobu Kawahara, Stephen Roberts, Stefan Zohren

机构 * Japan Digital Design, Inc.(日本数字设计公司)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 本文提出了一种多智能体LLM交易框架,通过细粒度任务分解提升风险调整后的回报,并通过分析输出与决策偏好的一致性提高系统性能。

Comments 14 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23300 2026-02-27 cs.CL eess.AS 57%

A Mixture-of-Experts Model for Multimodal Emotion Recognition in Conversations

一种用于对话中多模态情绪识别的专家混合模型

Soumya Dutta, Smruthi Balaji, Sriram Ganapathy

机构 * LEAP Lab, Department of Electrical Engineering(LEAP实验室,电气工程系) Microsoft(微软)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 MiSTER-E通过专家混合框架提升对话中多模态情绪识别的准确率,实现跨模态一致性与融合。

Comments Accepted to Elsevier Computer Speech and Language. 30 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22865 2026-02-27 cs.CL 57%

Effective QA-driven Annotation of Predicate-Argument Relations Across Languages

跨语言有效的问题回答驱动的谓词-论元关系标注

Jonathan Davidov, Aviv Slobodkin, Shmuel Tomi Klein, Reut Tsarfaty, Ido Dagan, Ayal Klein

机构 * Bar-Ilan University(巴伊兰大学) Ariel University(阿里尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 本文提出一种跨语言问题回答驱动的谓词-论元关系标注方法,通过重用英语QA-SRL解析器生成多语言高质量标注数据,提升语义解析效率。

Comments Accepted to EACL 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22771 2026-02-27 cs.AI cs.DB 57%

ClinDet-Bench: Beyond Abstention, Evaluating Judgment Determinability of LLMs in Clinical Decision-Making

ClinDet-Bench: 超越回避,评估大语言模型在临床决策中的判断可确定性

Yusuke Watanabe, Yohei Kobashi, Takeshi Kojima, Yusuke Iwasawa, Yasushi Okuno, Yutaka Matsuo

机构 * Kyoto University, Department of Biomedical Data Intelligence(京都大学生物医学数据智能系) Kyoto University, Department of Cardiovascular Medicine(京都大学心血管医学系) The University of Tokyo(东京大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 ClinDet-Bench通过评估大语言模型在临床决策中判断可确定性的能力,揭示其在不完整信息下的局限性,并提供评估框架以提高医疗决策的安全性。

Comments 17 pages, 3 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22678 2026-02-27 cs.CV cs.AI 57%

ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport

ViCLIP-OT:首个面向越南语图像-文本检索的视觉-语言基础模型,结合最优传输

Quoc-Khang Tran, Minh-Thien Nguyen, Nguyen-Khang Pham

机构 * Can Tho University(金兰大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 ViCLIP-OT通过整合CLIP对比学习与SIGROT损失,提升越南语图像-文本检索性能,实现域内和零样本设置下的显著改进。

Comments Preprint submitted to Expert Systems with Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22452 2026-02-27 cs.AI cs.RO 57%

CWM: Contrastive World Models for Action Feasibility Learning in Embodied Agent Pipelines

CWM: 用于具身智能体流水线中动作可行性学习的对比世界模型

Chayan Banerjee

机构 * School of Electrical Engineering and Robotics, Queensland University of Technology(电气工程与机器人学学院,昆士兰理工大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 CWM通过对比训练提升动作可行性评分,优于SFT方法,提升精度并增强安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22236 2026-02-27 q-bio.GN cs.CV cs.LG 57%

CrossLLM-Mamba: Multimodal State Space Fusion of LLMs for RNA Interaction Prediction

CrossLLM-Mamba: LLMs多模态状态空间融合用于RNA相互作用预测

Rabeya Tus Sadia, Qiang Ye, Qiang Cheng

机构 * Department of Computer Science, University of Kentucky, Lexington, KY, USA(计算机科学系,肯塔基大学,路易斯维尔,KY,美国) Department of Mathematics, University of Kentucky, Lexington, KY, USA(数学系,肯塔基大学,路易斯维尔,KY,美国) Institute for Biomedical Informatics, University of Kentucky, Lexington, KY, USA(生物医学信息学研究所,肯塔基大学,路易斯维尔,KY,美国)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

AI总结 CrossLLM-Mamba通过多模态状态空间融合实现RNA相互作用预测,采用双向Mamba编码器和动态序列转换模型,达到高精度性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23363 2026-02-27 cs.CV 50%

MediX-R1: Open Ended Medical Reinforcement Learning

MediX-R1:开放端医疗强化学习

Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Omair Mohamed, Mohamed Zidan, Fahad Khan, Salman Khan, Rao Anwer, Hisham Cholakkal

机构 * Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)(迈赫迈德·本·扎耶德人工智能大学) Jubilee Mission Medical College(jubilee mission 医学院) Research Institute(研究院) JJM Medical College(JJM 医学院)

专题命中 安全评测 :alignment(abstract)

AI总结 MediX-R1通过综合奖励信号和LLM评估,实现医疗多模态模型的开放端强化学习,提升临床推理可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22923 2026-02-27 cs.CV cs.RO 50%

WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents

WaterVideoQA: 以ASV为中心的感知与符合规则的推理 via 多模态智能体

Runwei Guan, Shaofeng Liang, Ningwei Ouyang, Weichen Fei, Shanliang Yao, Wei Dai, Chenhao Ge, Penglei Sun, Xiaohui Zhu, Tao Huang, Ryan Wen Liu, Hui Xiong

机构 * Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州)人工智能研究所) Hubei Key Laboratory of Inland Shipping Technology (Wuhan University of Technology)(湖北内河航运技术重点实验室(武汉理工大学)) School of Navigation, Wuhan University of Technology(武汉理工大学航海学院) School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学先进科技学院) School of Artificial Intelligence, Nanjing University(南京大学人工智能学院) School of Information Engineering, Yancheng Institute of Technology(盐城职业技术学院信息工程学院) School of Engineering, Stanford University(斯坦福大学工程学院) Centre for AI and Data Science Innovation and the School of Science and Engineering, James Cook University(詹姆斯库克大学人工智能与数据科学创新中心及科学与工程学院)

专题命中 安全评测 :trustworthy(abstract)

AI总结 WaterVideoQA通过多模态智能体系统,实现ASV在复杂水域环境中的感知与规则合规推理,提升自主航行的安全性和精确性。

Comments 11 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20903 2026-02-27 cs.CV 50%

TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering

TextPecker: 通过奖励结构异常量化提升视觉文本渲染

Hanshen Zhu, Yuliang Liu, Xuecheng Wu, An-Lan Wang, Hao Feng, Dingkang Yang, Chao Feng, Can Huang, Jingqun Tang, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) ByteDance(字节跳动)

专题命中 安全评测 :alignment(abstract)

AI总结 TextPecker通过结构异常感知强化学习策略提升视觉文本渲染的结构忠实度和语义对齐度。

Comments Accepted by CVPR 2026; Code: https://github.com/CIawevy/TextPecker

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22253 2026-02-27 cs.SD 50%

AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs

AR&D: 一种用于音频大语言模型解释的检索与描述框架

Townim Faisal Chowdhury, Ta Duc Huy, Siqi Pan, Jeremy Stoddard, Zhibin Liao

机构 * Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) Dolby Laboratories(杜比实验室) School of Computer and Mathematical Sciences, University of Adelaide, Australia(计算机与数学科学学院,阿德莱德大学,澳大利亚)

专题命中 安全评测 :trustworthy(abstract)

AI总结 AR&D框架通过稀疏自编码器解构音频大语言模型的多义激活,实现对模型内部特征的可解释性增强,为高风险领域应用提供可靠部署基础。

Comments Accepted at International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏