arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7971 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7971 篇

2412.11871 2025-04-04 cond-mat.soft cond-mat.stat-mech physics.bio-ph 71%

Reentrant phase behavior in binary topological flocks with nonreciprocal alignment

Tian Tang, Yu Duan, Yu-qiang Ma

专题命中 其他安全 :alignment(title)

Comments Supplemental movies are available on request

Journal ref Phys. Rev. Research 7, 023008 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18670 2025-03-25 astro-ph.IM physics.data-an stat.ML 71%

Deep learning-based identification of precipitation clouds from all-sky camera data for observatory safety

Mohammad H. Zhoolideh Haghighi, Alireza Ghasrimanesh, Habib Khosroshahi

专题命中 其他安全 :safety(title)

Comments A version of this work has been published in Machine Learning with Applications (MLWA)

Journal ref Machine Learning with Applications Volume 20, June 2025, 100640

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09416 2025-03-13 cs.CV 71%

OpenVidVRD: Open-Vocabulary Video Visual Relation Detection via Prompt-Driven Semantic Space Alignment

Qi Liu, Weiying Xue, Yuxiao Wang, Zhenao Wei

专题命中 其他安全 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05564 2025-03-10 q-bio.NC 71%

Phase Alignment Enhances Oscillatory Power in Neural Mass Models Optimized for Class Encoding

Alexander Pei

专题命中 其他安全 :alignment(title)

Comments 4 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16410 2025-01-29 cs.CV 71%

DynAlign: Unsupervised Dynamic Taxonomy Alignment for Cross-Domain Segmentation

Han Sun, Rui Gong, Ismail Nejjar, Olga Fink

专题命中 其他安全 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08088 2025-01-15 cs.CV 71%

AgentPose: Progressive Distribution Alignment via Feature Agent for Human Pose Distillation

Feng Zhang, Jinwei Liu, Xiatian Zhu, Lei Chen

专题命中 其他安全 :alignment(title)

Comments 5 pages, 1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08855 2024-12-03 cs.CV 71%

DPA: Dual Prototypes Alignment for Unsupervised Adaptation of Vision-Language Models

Eman Ali, Sathira Silva, Muhammad Haris Khan

专题命中 其他安全 :alignment(title)

Comments Accepted at WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01822 2024-11-05 cs.CV 71%

Distribution alignment based transfer fusion frameworks on quantum devices for seeking quantum advantages

Xi He, Feiyu Du, Xiaohan Yu, Yang Zhao, Tao Lei

专题命中 其他安全 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14238 2024-10-21 cs.CV 71%

Storyboard guided Alignment for Fine-grained Video Action Recognition

Enqi Liu, Liyuan Pan, Yan Yang, Yiran Zhong, Zhijing Wu, Xinxiao Wu, Liu Liu

专题命中 其他安全 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05904 2024-09-05 cs.CV 71%

Enhancing Representation in Radiography-Reports Foundation Model: A Granular Alignment Algorithm Using Masked Contrastive Learning

Weijian Huang, Cheng Li, Hong-Yu Zhou, Hao Yang, Jiarun Liu, Yong Liang, Hairong Zheng, Shaoting Zhang, Shanshan Wang

专题命中 其他安全 :alignment(title)

Journal ref Nature Communications 15, 7620 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00670 2024-06-07 cs.CV 71%

Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation

Yunheng Li, ZhongYu Li, Quansheng Zeng, Qibin Hou, Ming-Ming Cheng

专题命中 其他安全 :alignment(title)

Comments Accepted by ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12678 2024-05-27 cs.CV 71%

Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model

Jihao Dong, Renjie Pan, Hua Yang

专题命中 其他安全 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05103 2024-04-09 cs.HC 71%

Chart What I Say: Exploring Cross-Modality Prompt Alignment in AI-Assisted Chart Authoring

Nazar Ponochevnyi, Anastasia Kuzminykh

专题命中 其他安全 :alignment(title)

Comments Will be published In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10958 2024-01-23 eess.IV 71%

Detection of Thermal Events by Semi-Supervised Learning for Tokamak First Wall Safety

Christian Staron, Hervé Le Borgne, Raphaël Mitteau, Erwan Grelier, Nicolas Allezard

专题命中 其他安全 :safety(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08535 2023-11-16 cs.DB 71%

Taxonomy, Semantic Data Schema, and Schema Alignment for Open Data in Urban Building Energy Modeling

Liang Zhang, Jianli Chen, Jia Zou

专题命中 其他安全 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01317 2023-09-11 cs.CV eess.IV 71%

ELIXR: Towards a general purpose X-ray artificial intelligence system through alignment of large language models and radiology vision encoders

Shawn Xu, Lin Yang, Christopher Kelly, Marcin Sieniek, Timo Kohlberger, Martin Ma, Wei-Hung Weng, Atilla Kiraly, Sahar Kazemzadeh, Zakkai Melamed, Jungyeon Park, Patricia Strachan, Yun Liu, Chuck Lau, Preeti Singh, Christina Chen, Mozziyar Etemadi, Sreenivasa Raju Kalidindi, Yossi Matias, Katherine Chou, Greg S. Corrado, Shravya Shetty, Daniel Tse, Shruthi Prabhakara, Daniel Golden, Rory Pilgrim, Krish Eswaran, Andrew Sellergren

专题命中 其他安全 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07275 2022-11-15 cs.CV 71%

Zero-shot Image Captioning by Anchor-augmented Vision-Language Space Alignment

Junyang Wang, Yi Zhang, Ming Yan, Ji Zhang, Jitao Sang

专题命中 其他安全 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.02830 2021-06-08 eess.AS 71%

Reinforce-Aligner: Reinforcement Alignment Search for Robust End-to-End Text-to-Speech

Hyunseung Chung, Sang-Hoon Lee, Seong-Whan Lee

专题命中 其他安全 :alignment(title)

Comments Accepted in INTERSPEECH 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.12651 2020-05-19 cs.PL 71%

Type Safety with JSON Subschema

Andrew Habib, Avraham Shinnar, Martin Hirzel, Michael Pradel

专题命中 其他安全 :safety(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.13900 2020-04-01 cs.IR cs.SI 71%

A large-scale Twitter dataset for drug safety applications mined from publicly existing resources

Ramya Tekumalla, Juan M. Banda

专题命中 其他安全 :safety(title)

Comments 8 tables, 2 figures, 7 pages, accepted after peer review as a workshop paper in ACM Conference on Health, Inference, and Learning (CHIL) 2020 https://www.chilconference.org/agenda/

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.11861 2020-02-28 cs.MA eess.SP 71%

Simulation of Real-time Routing for UAS traffic Management with Communication and Airspace Safety Considerations

Zhao Jin, Ziyi Zhao, Chen Luo, Franco Basti, Adrian Solomon, M. Cenk Gursoy, Carlos Caicedo, Qinru Qiu

专题命中 其他安全 :safety(title)

Comments The 38th AIAA/IEEE Digital Avionics Systems Conference (DASC)

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.00520 2018-07-25 astro-ph.SR 71%

Tracking the spin axes orbital alignment in selected binary systems - Torun Rossiter-McLaughlin effect survey

P. Sybilski, R. K. Pawłaszek, A. Sybilska, M. Konacki, K. G. Hełminiak, S. K. Kozłowski, M. Ratajczak

专题命中 其他安全 :alignment(title)

Comments 30 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.04766 2017-08-17 cond-mat.mes-hall 71%

Fundamental Band Gap and Alignment of Two-Dimensional Semiconductors Explored by Machine Learning

Zhen Zhu, Baojuan Dong, Teng Yang, Zhi-Dong Zhang

专题命中 其他安全 :alignment(title)

Comments 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.02588 2017-01-11 q-bio.QM 71%

Synthesis of Methotrexate loaded Cerium fluoride nanoparticles with pH sensitive extended release coupled with Hyaluronic acid receptor with plausible theranostic capabilities for preclinical safety studies

Nitish Manu George

专题命中 其他安全 :safety(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11955 2026-08-14 cs.CY cs.HC 版本更新 70%

Philosophical vertigo with artificial intelligence

人工智能引发的哲学眩晕

Thomas A. Pollak, Hamilton Morrin, Murray Shanahan

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CY

AI总结 该研究提出人工智能引发的哲学眩晕概念,分析其产生、传播路径,关联临床妄想案例,指出AI将参与重构人类认知环境,并提出哲学可修正性作为应对对策。

Comments 29 pages, no figures. Source formatting revised to improve arXiv HTML accessibility; article text unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03538 2026-08-11 cs.LG 版本更新 70%

Online Learnability of Chain-of-Thought Verifiers: Soundness and Completeness Trade-offs

链式思维验证器的在线可学习性:正确性与完备性的权衡

Maria-Florina Balcan, Avrim Blum, Kiriaki Fragkia, Zhiyuan Li, Dravyansh Sharma

机构 * Carnegie Mellon University(卡内基梅隆大学) Toyota Technological Institute at Chicago(芝加哥丰田技术研究所) Northwestern University(西北大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG

AI总结 本文提出一种在线学习框架,用于学习链式思维验证器,通过检查解决方案的正确性,解决生成器与验证器之间的反馈循环导致的分布偏移问题,并引入新的Littlestone维度扩展以优化验证器的学习。

Comments The abstract has been abridged due to arXiv length constraints

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02491 2026-08-06 cs.AI 版本更新 70%

Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions

长期测量:迈向对人机交互的纵向理解

Nicole Mitchell, Dhruv Agarwal, Maty Bohacek, Remi Denton, Roma Patel

机构 * Google Research(谷歌研究院) Cornell University(康奈尔大学) Stanford University(斯坦福大学) Google DeepMind(谷歌DeepMind)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 本研究针对语言模型融入生活引发的长期人机交互风险,结合社会科学测量与NLP计算方法,提出通过长期测量建模人类行为变化,实现问题行为在线检测以缓解用户长期风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13904 2026-07-24 cs.AI 版本更新 70%

Diagnosing Pathological Chain-of-Thought in Reasoning Models

诊断推理模型中的病理链式思维

Manqing Liu, David Williams-King, Ida Caspary, Linh Le, Hannes Whittingham, Puria Radmard, Cameron Tice, Edward James Young

机构 * Department of Epidemiology, CAUSALab, Harvard University, Boston, USA(流行病学系、CAUSALab、哈佛大学) Imperial College London, London, UK(伦敦帝国学院) McGill University, Montreal, Canada(麦吉尔大学) Geodesic Research, Cambridge, UK(Geodesic研究)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文提出了一种评估链式思维推理模型中病态的实用工具,通过定义具体度量标准和训练特定模型生物来识别和区分三种不同的病态现象。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31614 2026-07-01 eess.SY cs.AI cs.SY 新提交 70%

Automating Cause-Effect Specification with Knowledge Graphs and Large Language Models

利用知识图谱和大语言模型自动化因果规范生成

Javal Vyas, Milapji Singh Gill, Mehmet Mercangöz

机构 * Autonomous Industrial Systems Lab, Imperial College London(帝国理工学院自主工业系统实验室)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 提出一种语义AI框架,结合知识图谱与约束大语言模型,自动生成因果逻辑和操作安全叙述,减少手动工作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03335 2026-06-30 cs.CL 70%

Compressed Sensing for Capability Localization in Large Language Models

压缩感知在大语言模型能力定位中的应用

Anna Bair, Yixuan Even Xu, Mingjie Sun, J. Zico Kolter

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL

AI总结 研究通过压缩感知方法识别大语言模型中特定能力依赖的稀疏注意力头,发现关闭少量头可显著降低特定能力表现,揭示了模型模块化组织原则。

详情

展开后加载摘要…

URL PDF HTML 收藏