arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9324 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9324 篇

2602.19715 2026-02-24 cs.CV 71%

Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision

像素不会说谎(但你的检测器可能会):通过生成器-评估器过程构建MLLM作为判断者以实现可信的深度伪造检测和推理监督

Kartik Kuckreja, Parul Gupta, Muhammad Haris Khan, Abhinav Dhall

机构 * MBZUAI Monash University(墨尔本大学)

专题命中 安全评测 :trustworthy(title)

AI总结 本文提出DeepfakeJudge框架,通过生成器-评估器过程提升深度伪造检测的推理忠实性,实现高准确率和高一致性的推理监督。

Comments CVPR-2026, Code is available here: https://github.com/KjAeRsTuIsK/DeepfakeJudge

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15676 2026-02-12 cs.MA cs.CR 71%

Trustworthy Decentralized Autonomous Machines: A New Paradigm in Automation Economy

可信去中心化自主机器:自动化经济的新范式

Fernando Castillo, Oscar Castillo, Eduardo Brito, Simon Espinola

专题命中 安全评测 :trustworthy(title)

AI总结 本文提出去中心化自主机器(DAMs)作为自动化经济的新范式,通过整合AI、区块链和物联网技术,实现无信任的资产管理和经济机会民主化。

Comments To be published in IEEE International Workshop on Decentralized Physical Infrastructure Networks 2025, in conjunction with ICBC'25. 7 pages. 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09101 2026-01-30 cs.HC 71%

A Survey of LLM Alignment: Instruction Understanding, Intention Reasoning, and Reliable Generation

大型语言模型对齐综述:指令理解、意图推理和可靠生成

Zongyu Chang, Feihong Lu, Ziqin Zhu, Qian Li, Cheng Ji, Tao Yang, Zhuo Chen, Hao Peng, Yang Liu, Ruifeng Xu, Yangqiu Song, Jianxin Li, Shangguang Wang

专题命中 安全评测 :alignment(title)

AI总结 本文综述了大型语言模型在指令理解、意图推理和可靠生成方面的挑战与解决方案,旨在提升模型在现实应用中的可靠性与适应性。

Comments 37 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18983 2026-01-28 cs.DC 71%

Trustworthy Scheduling for Big Data Applications

大数据应用中的可信调度

Dimitrios Tomaras, Vana Kalogeraki, Dimitrios Gunopulos

专题命中 安全评测 :trustworthy(title)

AI总结 X-Sched通过整合反事实解释与机器学习模型,提供可操作的资源配置指导,提升容器化环境中的任务执行效率和透明度。

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05879 2026-01-19 cs.HC 71%

Human-AI Alignment of Multimodal Large Language Models with Speech-Language Pathologists in Parent-Child Interactions

与语言病理学家在亲子互动中的人机对齐多模态大语言模型

Weiyan Shi, Kenny Tsu Wei Choo

专题命中 安全评测 :alignment(title)

AI总结 本研究通过多模态大语言模型与语言病理学家的对齐,开发了支持亲子互动分析的系统,实现了85%的感知线索提取准确率和75%的判断精确率,并提出了行为观察-判断系统的构建指南。

Comments This is an earlier version of the work released in May 2025. The version accepted at CHI 2026 is available as a separate preprint at arXiv:2511.04366

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08420 2026-01-14 cs.CV 71%

MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP

MMLGNet: 利用CLIP实现遥感数据的跨模态对齐

Aditya Chaudhary, Sneha Barman, Mainak Singha, Ankit Jha, Girish Mishra, Biplab Banerjee

专题命中 安全评测 :alignment(title)

AI总结 MMLGNet通过CLIP实现遥感数据的跨模态对齐,利用多模态语言引导网络有效融合光谱、空间和几何信息,提升语义理解能力。

Comments Accepted at InGARSS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26781 2025-11-04 cs.CV 71%

ChartAB: A Benchmark for Chart Grounding & Dense Alignment

Aniruddh Bansal, Davit Soselia, Dang Nguyen, Tianyi Zhou

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 安全评测 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00192 2025-10-08 cs.CV 71%

Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety

Younggun Kim, Sirnam Swetha, Fazil Kagdi, Mubarak Shah

机构 * Center For Research in Computer Vision, University of Central Florida, USA(计算机视觉研究中心,中央佛罗里达大学) Department of Civil Environmental and Construction Engineering, University of Central Florida, USA(土木环境与建设工程系,中央佛罗里达大学) Department of Computer Science, University of Central Florida, USA(计算机科学系,中央佛罗里达大学)

专题命中 安全评测 :safety(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14396 2025-09-19 econ.TH cs.GT 71%

Friend or Foe: Delegating to an AI Whose Alignment is Unknown

Drew Fudenberg, Annie Liang

专题命中 安全评测 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01209 2025-09-03 cs.CV 71%

Measuring Image-Relation Alignment: Reference-Free Evaluation of VLMs and Synthetic Pre-training for Open-Vocabulary Scene Graph Generation

Maëlic Neau, Zoe Falomir, Cédric Buche, Akihiro Sugimoto

机构 * Computing Science Department, Umeå University(乌梅大学计算科学系) CNRS IRL 2010 CROSSING(法国CNRS IRL 2010 CROSSING) IMT Atlantique(IMT阿蒂昂大学) National Institute of Informatics(日本信息机构)

专题命中 安全评测 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10228 2025-07-15 cs.SE 71%

Towards a Framework for Operationalizing the Specification of Trustworthy AI Requirements

Hugo Villamizar, Daniel Mendez, Marcos Kalinowski

专题命中 安全评测 :trustworthy(title)

Comments This paper has been accepted for presentation at the 2025 IEEE 33rd International Requirements Engineering Conference Workshops (REW-RETRAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17503 2025-06-24 cs.CV 71%

Trustworthy Few-Shot Transfer of Medical VLMs through Split Conformal Prediction

Julio Silva-Rodríguez, Ismail Ben Ayed, Jose Dolz

机构 * ÉTS Montréal(蒙特利尔ÉTS学院)

专题命中 安全评测 :trustworthy(title)

Comments MICCAI 2025. Code: https://github.com/jusiro/SCA-T

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08792 2025-06-12 cs.CR 71%

ProxyGPT: Enabling User Anonymity in LLM Chatbots via (Un)Trustworthy Volunteer Proxies

Dzung Pham, Jade Sheffey, Chau Minh Pham, Amir Houmansadr

专题命中 安全评测 :trustworthy(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05153 2025-06-10 cs.CV 71%

Hummingbird: High Fidelity Image Generation via Multimodal Context Alignment

Minh-Quan Le, Gaurav Mittal, Tianjian Meng, A S M Iftekhar, Vishwas Suryanarayanan, Barun Patra, Dimitris Samaras, Mei Chen

机构 * Microsoft(微软公司) Stony Brook University(史蒂文尼森布鲁克大学)

专题命中 安全评测 :alignment(title)

Comments Accepted to ICLR 2025. Project page with code release: https://roar-ai.github.io/hummingbird

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06912 2025-05-13 cs.CV 71%

Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI

Chao Ding, Mouxiao Bian, Pengcheng Chen, Hongliang Zhang, Tianbin Li, Lihao Liu, Jiayuan Chen, Zhuoran Li, Yabei Zhong, Yongqi Liu, Haiqing Huang, Dongming Shan, Junjun He, Jie Xu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) Shanghai Kupas Technology Limited Company(上海库帕斯科技有限公司)

专题命中 安全评测 :trustworthy(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10635 2025-05-01 cs.CV 71%

MeDSLIP: Medical Dual-Stream Language-Image Pre-training with Pathology-Anatomy Semantic Alignment

Wenrui Fan, Mohammod N. I. Suvon, Shuo Zhou, Xianyuan Liu, Samer Alabed, Venet Osmani, Andrew J. Swift, Chen Chen, Haiping Lu

机构 * Centre for Machine Intelligence and School of Computer Science, University of Sheffield(智能中心和计算机科学学院,谢菲尔德大学) School of Medicine and Population Health, and INSIGNEO, Institute for in Silico Medicine, University of Sheffield(医学与人口健康学院,INSIGNEO,虚拟医学研究所,谢菲尔德大学) Digital Environment Research Institute, Queen Mary University of London(数字环境研究所,伦敦大学玛丽女王学院) School of Computer Science, University of Sheffield(计算机科学学院,谢菲尔德大学) Department of Computing, Imperial College London(计算系,伦敦帝国学院)

专题命中 安全评测 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19338 2025-04-29 physics.flu-dyn 71%

OpenFOAMGPT 2.0: end-to-end, trustworthy automation for computational fluid dynamics

Jingsen Feng, Ran Xu, Xu Chu

专题命中 安全评测 :trustworthy(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07189 2025-04-11 eess.SY cs.SY 71%

Multi-Agent Trustworthy Consensus under Random Dynamic Attacks

Orhan Eren Akgün, Sarper Aydın, Stephanie Gil, Angelia Nedić

专题命中 安全评测 :trustworthy(title)

Comments 16 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19391 2025-03-26 cs.CV cs.MA 71%

TraF-Align: Trajectory-aware Feature Alignment for Asynchronous Multi-agent Perception

Zhiying Song, Lei Yang, Fuxi Wen, Jun Li

专题命中 安全评测 :alignment(title)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07906 2025-03-12 cs.CV 71%

Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning

Qinghao Ye, Xianhan Zeng, Fu Li, Chunyuan Li, Haoqi Fan

专题命中 安全评测 :alignment(title)

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00663 2025-02-20 cs.CV cs.RO 71%

Generalized Robot 3D Vision-Language Model with Fast Rendering and Pre-Training Vision-Language Alignment

Kangcheng Liu, Yong-Jin Liu, Baoquan Chen

专题命中 安全评测 :alignment(title)

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence, Manuscript Info: 17 Pages, 13 Figures, and 6 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09048 2024-12-17 cs.SE 71%

Towards Trustworthy LLMs for Code: A Data-Centric Synergistic Auditing Framework

Chong Wang, Zhenpeng Chen, Tianlin Li, Yilun Zhao, Yang Liu

专题命中 安全评测 :trustworthy(title)

Comments Short Vision Paper, Accepted by ICSE'25-NIER

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10122 2024-10-02 cs.CV 71%

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, Li Yuan

专题命中 安全评测 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16455 2024-09-26 cs.RO 71%

MultiTalk: Introspective and Extrospective Dialogue for Human-Environment-LLM Alignment

Venkata Naren Devarakonda, Ali Umut Kaypak, Shuaihang Yuan, Prashanth Krishnamurthy, Yi Fang, Farshad Khorrami

专题命中 安全评测 :alignment(title)

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10823 2024-08-21 cs.CV eess.IV 71%

Trustworthy Compression? Impact of AI-based Codecs on Biometrics for Law Enforcement

Sandra Bergmann, Denise Moussa, Christian Riess

专题命中 安全评测 :trustworthy(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00098 2024-07-02 eess.IV cs.CV 71%

Scalable, Trustworthy Generative Model for Virtual Multi-Staining from H&E Whole Slide Images

Mehdi Ounissi, Ilias Sarbout, Jean-Pierre Hugot, Christine Martinez-Vinson, Dominique Berrebi, Daniel Racoceanu

专题命中 安全评测 :trustworthy(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01459 2024-01-12 cs.CV 71%

Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization

Jameel Hassan, Hanan Gani, Noor Hussein, Muhammad Uzair Khattak, Muzammal Naseer, Fahad Shahbaz Khan, Salman Khan

专题命中 安全评测 :alignment(title)

Comments Accepted to NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.08533 2024-01-08 cs.CR 71%

Trustchain -- Trustworthy Decentralised Public Key Infrastructure for Digital Credentials

Tim Hobson, Lydia France, Sam Greenbury, Luke Hare, Pamela Wochner

专题命中 安全评测 :trustworthy(title)

Comments 10 pages, 4 figures, presented at the International Conference on AI and the Digital Economy (CADE 2023), Venice, Italy. Replaces the preprint version, with minor changes & additions based on reviewers' comments

Journal ref International Conference on AI and the Digital Economy (CADE 2023), 2023, pp. 31-40

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14880 2023-09-28 cs.CV 71%

SGAligner : 3D Scene Alignment with Scene Graphs

Sayan Deb Sarkar, Ondrej Miksik, Marc Pollefeys, Daniel Barath, Iro Armeni

专题命中 安全评测 :alignment(title)

Comments Accepted at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12132 2023-06-22 cs.SE 71%

ChatGPT as a tool for User Story Quality Evaluation: Trustworthy Out of the Box?

Krishna Ronanki, Beatriz Cabrero-Daniel, Christian Berger

专题命中 安全评测 :trustworthy(title)

Comments 9 Pages, 2 Tables, 1 Figure. Accepted at AI-Assisted Agile Software Development Workshop (Co-located with XP 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏