arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7971 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7971 篇

2211.01817 2023-03-20 cs.AI cs.CY cs.LG 67%

Liability regimes in the age of AI: a use-case driven analysis of the burden of proof

David Fernández Llorca, Vicky Charisi, Ronan Hamon, Ignacio Sánchez, Emilia Gómez

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Paper published at the Journal of Artificial Intelligence Research

Journal ref Journal of Artificial Intelligence Research, Vol. 76 (2023), pp. 613-644

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.03430 2023-02-21 cs.LG cs.AI cs.CL cs.CV cs.MM 67%

Foundations and Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions

Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12737 2022-11-24 cs.CV cs.AI cs.CL cs.LG 67%

RoentGen: Vision-Language Foundation Model for Chest X-ray Generation

Pierre Chambon, Christian Bluethgen, Jean-Benoit Delbrouck, Rogier Van der Sluijs, Małgorzata Połacin, Juan Manuel Zambrano Chaves, Tanishq Mathew Abraham, Shivanshu Purohit, Curtis P. Langlotz, Akshay Chaudhari

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.06318 2022-11-14 cs.CY cs.AI cs.LG 67%

Artificial Intelligence and Life in 2030: The One Hundred Year Study on Artificial Intelligence

Peter Stone, Rodney Brooks, Erik Brynjolfsson, Ryan Calo, Oren Etzioni, Greg Hager, Julia Hirschberg, Shivaram Kalyanakrishnan, Ece Kamar, Sarit Kraus, Kevin Leyton-Brown, David Parkes, William Press, AnnaLee Saxenian, Julie Shah, Milind Tambe, Astro Teller

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 52 pages, https://ai100.stanford.edu/2016-report

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.14504 2022-10-11 cs.CL cs.AI cs.LG 67%

GERNERMED++: Transfer Learning in German Medical NLP

Johann Frei, Ludwig Frei-Stuber, Frank Kramer

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.09900 2022-09-21 cs.CL cs.AI cs.LG 67%

LINGUIST: Language Model Instruction Tuning to Generate Annotated Utterances for Intent Classification and Slot Tagging

Andy Rosenbaum, Saleh Soltan, Wael Hamza, Yannick Versley, Markus Boese

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to The 29th International Conference on Computational Linguistics (COLING 2022) October 12-17, 2022, Gyeongju, Republic of Korea https://coling2022.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.10726 2022-09-15 cs.CL cs.AI cs.LG 67%

TWEET-FID: An Annotated Dataset for Multiple Foodborne Illness Detection Tasks

Ruofan Hu, Dongyu Zhang, Dandan Tao, Thomas Hartvigsen, Hao Feng, Elke Rundensteiner

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments LREC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.02222 2022-08-04 cs.LG cs.AI cs.CY eess.SP 67%

Blockchain associated machine learning and IoT based hypoglycemia detection system with auto-injection feature

Rahnuma Mahzabin, Fahim Hossain Sifat, Sadia Anjum, Al-Akhir Nayan, Muhammad Golam Kibria

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Journal ref Indonesian Journal of Electrical Engineering and Computer Science, Vol. 27, No. 1, pp. 447-455, July 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.10086 2022-06-17 cs.AI cs.CY cs.LG 67%

Learning Models of Individual Behavior in Chess

Reid McIlroy-Young, Russell Wang, Siddhartha Sen, Jon Kleinberg, Ashton Anderson

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 12 pages, 11 figures, 5 tables, Published in the Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2022), Code https://github.com/CSSLab/maia-individual

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.08184 2022-05-18 cs.CL cs.AI cs.LG 67%

SKILL: Structured Knowledge Infusion for Large Language Models

Fedor Moiseev, Zhe Dong, Enrique Alfonseca, Martin Jaggi

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.10761 2021-12-28 cs.CL cs.AI cs.LG 67%

Bilingual Lexicon Induction through Unsupervised Machine Translation

Mikel Artetxe, Gorka Labaka, Eneko Agirre

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.07401 2021-09-16 cs.CL cs.AI cs.IR cs.LG 67%

Matching with Transformers in MELT

Sven Hertling, Jan Portisch, Heiko Paulheim

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments accepted at the Ontology Matching Workshop at the International Semantic Web Conference (ISWC 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.06324 2021-09-15 cs.CL cs.AI cs.LG 67%

A Massively Multilingual Analysis of Cross-linguality in Shared Embedding Space

Alex Jones, William Yang Wang, Kyle Mahowald

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 15 pages, 8 figures, EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04319 2021-09-10 cs.CL cs.AI cs.LG 67%

Translate & Fill: Improving Zero-Shot Multilingual Semantic Parsing with Synthetic Data

Massimo Nicosia, Zhongdi Qu, Yasemin Altun

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to EMNLP 2021 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.02164 2020-03-04 cs.CL cs.AI cs.LG 67%

Plug and Play Language Models: A Simple Approach to Controlled Text Generation

Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, Rosanne Liu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICLR 2020 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.04914 2019-05-14 cs.AI cs.CL cs.LG 67%

Learning to Exploit Long-term Relational Dependencies in Knowledge Graphs

Lingbing Guo, Zequn Sun, Wei Hu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by the 36th International Conference on Machine Learning (ICML 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.07178 2019-02-20 eess.AS cs.AI cs.CL cs.LG cs.SD 67%

A spelling correction model for end-to-end speech recognition

Jinxi Guo, Tara N. Sainath, Ron J. Weiss

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to ICASSP 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.02306 2018-09-10 cs.CL cs.AI cs.LG 67%

Unsupervised Cross-lingual Word Embedding by Multilingual Neural Language Models

Takashi Wada, Tomoharu Iwata

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.03938 2017-11-15 cs.CL cs.AI cs.LG 67%

Representation Learning for Grounded Spatial Reasoning

Michael Janner, Karthik Narasimhan, Regina Barzilay

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to TACL 2017, code: https://github.com/jannerm/spatial-reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
1508.04395 2016-03-16 cs.CL cs.AI cs.LG cs.NE 67%

End-to-End Attention-based Large Vocabulary Speech Recognition

Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, Yoshua Bengio

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09984 2025-04-22 cs.LG cs.AI 66%

Symmetry-Breaking Augmentations for Ad Hoc Teamwork

Ravi Hammond, Dustin Craggs, Mingyu Guo, Jakob Foerster, Ian Reid

机构 * Foerster Lab for AI Research, University of Oxford(牛津大学人工智能研究实验室) Australian Institute for Machine Learning, University of Adelaide(阿德莱德大学人工智能研究所) Meta AI Research, UK(英国Meta人工智能研究) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他安全 :alignment(abstract,comments);分类 cs.AI、cs.LG

Comments 21 pages, 12 figures, Bidirectional Human-AI Alignment workshop, ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.08093 2021-12-16 cs.LG cs.AI 66%

Towards Controllable Agent in MOBA Games with Generative Modeling

Shubao Zhang

专题命中 其他安全 :alignment(abstract,comments);分类 cs.AI、cs.LG

Comments Human-Compatible AI; Human-AI Cooperation; AI control; AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10318 2021-10-22 cs.CL cs.LG 66%

Improved Multilingual Language Model Pretraining for Social Media Text via Translation Pair Prediction

Shubhanshu Mishra, Aria Haghighi

专题命中 其他安全 :alignment(abstract,comments);分类 cs.CL、cs.LG

Comments Camera ready version. Accepted to WNUT 2021. Code for reproducing the experiments can be found at: https://github.com/twitter-research/multilingual-alignment-tpp

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14763 2026-08-18 eess.IV cs.AI cs.CV cs.LG 新提交 62%

Cross-Modal Ultrasound-MRI Learning for Fetal Brain Ventricular Volumetry and Abnormality Screening

用于胎儿脑室体积测量与异常筛查的跨模态超声-MRI学习

Yuhao Huang, Yuanji Zhang, Yuhuan Lu, Dong Ni, P. Ellen Grant, Davood Karimi

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本研究提出跨模态学习框架VIFBA,基于超声视频实现胎儿脑室体积预测、VM严重程度分类及非VM异常筛查,性能优于基线与现有模型,为产前脑筛查提供实用方案。

Comments 17 pages, 11 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16514 2026-08-18 cs.CV cs.AI cs.CL cs.HC cs.MM 新提交 62%

Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans

匹配结果,不同注视:中央凹多模态大语言模型(MLLM)的搜索方式与人类的对比

Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Ulas Bagci, Alessandro Bruno

机构 * F-initiatives(F计划) Université Sorbonne Paris Nord(巴黎北索邦大学) Northwestern University(西北大学) IULM university(IULM大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 研究对比三款通用MLLM与人类在目标导向视觉搜索中的表现,发现模型在决策和目标获取上优于人类,但注视过程与人类不同,现有指标无法验证类人视觉,零样本模型不适用于过程层面问题。

Comments Paper accepted at 3rd HCV workshop at ECCV 2026. 12 pages main text, 16 pages supp

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16347 2026-08-18 cs.CL cs.LG 新提交 62%

Architecture-Dependent Causal Transfer of Activation States Across Large Language Models

大语言模型间激活状态的架构依赖型因果迁移

Fernando Cardenas Piepereit

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 本文探究大语言模型间激活状态的因果迁移,发现该迁移依赖模型架构,仅部分仅解码器模型对可实现具统计显著性的因果效应,且迁移的是表征载体而非意义。

Comments 13 pages, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15693 2026-08-18 cs.AI cs.LG 新提交 62%

Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment

面向小型设备的大模型:边缘AI部署的最新进展与实证分析

Subhransu Das, Jiaming Cheng, Arnav Kumar, Sadia Afrose, Mingzhe Han, Michael Silagy, Shreya Palande, Brijesh Soni, Rajiv Ramnath

机构 * The Ohio State University(俄亥俄州立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文调研近期边缘AI部署研究,提炼指南并在多平台测试,发现不同任务适用不同压缩技术,剪枝可能提升分割性能但会增加延迟,相关成果已开源。

Comments Parts of this work were presented at the IEEE Consumer Communications & Networking Conference (CCNC), Las Vegas, NV, USA, January 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12627 2026-08-18 cs.CV cs.AI cs.CL cs.HC 版本更新 62%

EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory

EgoCITE:面向长时程自我中心记忆的上下文增强索引与时序感知检索

Le Zhang, Ke Sun

机构 * University of Michigan(密歇根大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本研究针对长时程自我中心记忆系统的索引不可靠、忽略时序意图的问题,提出EgoCITE框架,经多数据集评估,其准确率优于基线且成本显著低于长上下文LLM智能体。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13596 2026-08-17 cs.LG cs.AI 新提交 62%

Training-Free Knowledge Transfer Across Model Scales through Activation-Guided Pruning

通过激活引导剪枝实现跨模型规模的无训练知识迁移

Jiahe Fan, Si Chen, Yinghao Hou, Aiyuan Zhang, Hong Xie

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出激活剪枝融合框架APM,通过激活引导选择源模型的显著组件并注入目标模型,无需训练和显式语义对齐,在16个基准上将3B目标模型平均准确率从55.5%提升至60.6%。

Comments 9 pages, 3 figures, and 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13341 2026-08-17 cs.LG cs.AI 版本更新 62%

Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples

面向红外光谱化学传感与分析的仿真到真实迁移学习:从分子到复杂样品

Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo, Zhedong Lin, Yuqiang Li, Xiangxiang Zeng, Tong Wang, Jun Xia

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Jilin University(吉林大学) University of Auckland(奥克兰大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Hunan University(湖南大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 研究针对红外光谱化学传感的迁移难题,提出超1亿参数的红外光谱基础模型UltraIR,经6000万仿真光谱预训练后,在多类化学分析任务中性能优于基线,且适配性与数据效率优异。

详情

展开后加载摘要…

URL PDF HTML 收藏