arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7978 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7978 篇

2406.04268 2024-06-07 cs.LG cs.AI 62%

Open-Endedness is Essential for Artificial Superhuman Intelligence

Edward Hughes, Michael Dennis, Jack Parker-Holder, Feryal Behbahani, Aditi Mavalankar, Yuge Shi, Tom Schaul, Tim Rocktaschel

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00160 2024-06-07 cs.CL cs.AI 62%

Self-Specialization: Uncovering Latent Expertise within Large Language Models

Junmo Kang, Hongyin Luo, Yada Zhu, Jacob Hansen, James Glass, David Cox, Alan Ritter, Rogerio Feris, Leonid Karlinsky

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ACL 2024 (Findings; Long Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03030 2024-06-06 cs.CL cs.LG 62%

From Tarzan to Tolkien: Controlling the Language Proficiency Level of LLMs for Content Generation

Ali Malik, Stephen Mayhew, Chris Piech, Klinton Bicknell

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Journal ref In Findings of the Association for Computational Linguistics (ACL 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02600 2024-06-06 cs.LG cs.AI stat.ML 62%

Data Quality in Edge Machine Learning: A State-of-the-Art Survey

Mohammed Djameleddine Belgoumri, Mohamed Reda Bouadjenek, Sunil Aryal, Hakim Hacid

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 31 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02030 2024-06-06 cs.CL cs.AI 62%

Multimodal Reasoning with Multimodal Knowledge Graph

Junlin Lee, Yequan Wang, Jing Li, Min Zhang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by ACL 2024 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18721 2024-06-06 cs.CV cs.AI cs.CL 62%

Correctable Landmark Discovery via Large Models for Vision-Language Navigation

Bingqian Lin, Yunshuang Nie, Ziming Wei, Yi Zhu, Hang Xu, Shikui Ma, Jianzhuang Liu, Xiaodan Liang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by TPAMI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01382 2024-06-04 cs.CL cs.AI 62%

Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function

Keyon Vafa, Ashesh Rambachan, Sendhil Mullainathan

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments To appear in ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.21068 2024-06-03 cs.CL cs.AI 62%

Code Pretraining Improves Entity Tracking Abilities of Language Models

Najoung Kim, Sebastian Schuster, Shubham Toshniwal

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19550 2024-05-31 cs.LG cs.CL 62%

Stress-Testing Capability Elicitation With Password-Locked Models

Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, David Krueger

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03893 2024-05-31 cs.CL cs.AI 62%

From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models

Luiza Pozzobon, Patrick Lewis, Sara Hooker, Beyza Ermis

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15789 2024-05-28 cs.AI cs.IT cs.LG cs.LO math.IT 62%

Semantic Objective Functions: A distribution-aware method for adding logical constraints in deep learning

Miguel Angel Mendez-Lucero, Enrique Bojorquez Gallardo, Vaishak Belle

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 12 pages,4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09647 2024-05-24 cs.SI cs.CL cs.CY 62%

Large Language Models Help Reveal Unhealthy Diet and Body Concerns in Online Eating Disorders Communities

Minh Duc Chu, Zihao He, Rebecca Dorn, Kristina Lerman

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13219 2024-05-24 cs.AI cs.CL 62%

How Reliable AI Chatbots are for Disease Prediction from Patient Complaints?

Ayesha Siddika Nipu, K M Sajjadul Islam, Praveen Madiraju

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments 24th IEEE International Conference on Information Reuse and Integration (IEEE IRI 2024), San Jose, CA, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02752 2024-05-22 cs.LG cs.AI 62%

Offline Reinforcement Learning with Imbalanced Datasets

Li Jiang, Sijie Cheng, Jielin Qiu, Haoran Xu, Wai Kin Chan, Zhao Ding

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Journal ref ICML 2023, workshop on Data-centric Machine Learning Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14381 2024-05-14 cs.CL cs.AI 62%

Editing Knowledge Representation of Language Model via Rephrased Prefix Prompts

Yuchen Cai, Ding Cao, Rongxi Guo, Yaqin Wen, Guiquan Liu, Enhong Chen

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 19pages,3figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14236 2024-05-14 cs.RO cs.AI cs.CV cs.LG 62%

MoDem-V2: Visuo-Motor World Models for Real-World Robot Manipulation

Patrick Lancaster, Nicklas Hansen, Aravind Rajeswaran, Vikash Kumar

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02795 2024-05-07 cs.AI cs.CL 62%

Evaluating and Optimizing Educational Content with Large Language Model Judgments

Joy He-Yueya, Noah D. Goodman, Emma Brunskill

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12420 2024-05-07 cs.CL cs.AI 62%

LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models

Shizhe Diao, Rui Pan, Hanze Dong, Ka Shun Shum, Jipeng Zhang, Wei Xiong, Tong Zhang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Published in NAACL 2024 Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02105 2024-05-06 cs.AI cs.CL cs.IT math.IT 62%

Evaluating Large Language Models for Structured Science Summarization in the Open Research Knowledge Graph

Vladyslav Nechakhin, Jennifer D'Souza, Steffen Eger

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 22 pages, 11 figures. In review at https://www.mdpi.com/journal/information/special_issues/WYS02U2GTD

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01561 2024-05-06 cs.SE cs.AI cs.CY 62%

Rapid Mobile App Development for Generative AI Agents on MIT App Inventor

Jaida Gao, Calab Su, Etai Miller, Kevin Lu, Yu Meng

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY

Journal ref Journal of advances in information science and technology 2(3) 1-8, March 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15597 2024-04-25 cs.NE cs.AI cs.LG cs.MA 62%

GRSN: Gated Recurrent Spiking Neurons for POMDPs and MARL

Lang Qin, Ziming Wang, Runhao Jiang, Rui Yan, Huajin Tang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12399 2024-04-25 cs.LG cs.CL cs.SI 62%

A Survey of Graph Meets Large Language Model: Progress and Future Directions

Yuhan Li, Zhixun Li, Peisong Wang, Jia Li, Xiangguo Sun, Hong Cheng, Jeffrey Xu Yu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments IJCAI 2024 Survey Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14906 2024-04-24 cs.CV cs.AI cs.LG 62%

Driver Activity Classification Using Generalizable Representations from Vision-Language Models

Ross Greer, Mathias Viborg Andersen, Andreas Møgelmose, Mohan Trivedi

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14864 2024-04-23 cs.AI cs.CV cs.LG 62%

Evaluating the Stability of Semantic Concept Representations in CNNs for Robust Explainability

Georgii Mikriukov, Gesina Schwalbe, Christian Hellert, Korinna Bade

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11797 2024-04-19 cs.CV cs.AI cs.LG 62%

When are Foundation Models Effective? Understanding the Suitability for Pixel-Level Classification Using Multispectral Imagery

Yiqun Xie, Zhihao Wang, Weiye Chen, Zhili Li, Xiaowei Jia, Yanhua Li, Ruichen Wang, Kangyang Chai, Ruohan Li, Sergii Skakun

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11589 2024-04-18 cs.CV cs.AI cs.LG 62%

Prompt Optimizer of Text-to-Image Diffusion Models for Abstract Concept Understanding

Zezhong Fan, Xiaohan Li, Chenhao Fang, Topojoy Biswas, Kaushiki Nag, Jianpeng Xu, Kannan Achan

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments WWW 2024 Companion

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08295 2024-04-17 cs.CL cs.AI 62%

Gemma: Open Models Based on Gemini Research and Technology

Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amélie Héliou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari, Charline Le Lan, Christopher A. Choquette-Choo, Clément Crepy, Daniel Cer, Daphne Ippolito, David Reid, Elena Buchatskaya, Eric Ni, Eric Noland, Geng Yan, George Tucker, George-Christian Muraru, Grigory Rozhdestvenskiy, Henryk Michalewski, Ian Tenney, Ivan Grishchenko, Jacob Austin, James Keeling, Jane Labanowski, Jean-Baptiste Lespiau, Jeff Stanway, Jenny Brennan, Jeremy Chen, Johan Ferret, Justin Chiu, Justin Mao-Jones, Katherine Lee, Kathy Yu, Katie Millican, Lars Lowe Sjoesund, Lisa Lee, Lucas Dixon, Machel Reid, Maciej Mikuła, Mateo Wirth, Michael Sharman, Nikolai Chinaev, Nithum Thain, Olivier Bachem, Oscar Chang, Oscar Wahltinez, Paige Bailey, Paul Michel, Petko Yotov, Rahma Chaabouni, Ramona Comanescu, Reena Jana, Rohan Anil, Ross McIlroy, Ruibo Liu, Ryan Mullins, Samuel L Smith, Sebastian Borgeaud, Sertan Girgin, Sholto Douglas, Shree Pandya, Siamak Shakeri, Soham De, Ted Klimenko, Tom Hennigan, Vlad Feinberg, Wojciech Stokowiec, Yu-hui Chen, Zafarali Ahmed, Zhitao Gong, Tris Warkentin, Ludovic Peran, Minh Giang, Clément Farabet, Oriol Vinyals, Jeff Dean, Koray Kavukcuoglu, Demis Hassabis, Zoubin Ghahramani, Douglas Eck, Joelle Barral, Fernando Pereira, Eli Collins, Armand Joulin, Noah Fiedel, Evan Senter, Alek Andreev, Kathleen Kenealy

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09574 2024-04-16 cs.LG cs.AI 62%

Predicting and Analyzing Pedestrian Crossing Behavior at Unsignalized Crossings

Chi Zhang, Janis Sprenger, Zhongjun Ni, Christian Berger

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 10 figures, 4 tables. Accepted in 2024 IEEE Intelligent Vehicles Symposium (IV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07484 2024-04-16 cs.CL cs.AI 62%

Psychometric Predictive Power of Large Language Models

Tatsuki Kuribayashi, Yohei Oseki, Timothy Baldwin

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 23 pages; Findings of NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06971 2024-04-11 cs.CV cs.AI cs.LG 62%

TrajPRed: Trajectory Prediction with Region-based Relation Learning

Chen Zhou, Ghassan AlRegib, Armin Parchami, Kunjan Singh

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏