arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7978 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7978 篇

2406.17804 2024-11-25 physics.med-ph cs.AI cs.CV cs.LG eess.IV 62%

A Review of Electromagnetic Elimination Methods for low-field portable MRI scanner

Wanyu Bian, Panfeng Li, Mengyao Zheng, Chihang Wang, Anying Li, Ying Li, Haowei Ni, Zixuan Zeng

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by 2024 5th International Conference on Machine Learning and Computer Application

Journal ref Proceedings of the 2024 5th International Conference on Machine Learning and Computer Application (ICMLCA), 2024, pp. 614-618

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09955 2024-11-22 cs.CV cs.AI cs.HC cs.LG cs.MM 62%

Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era

Thanh Tam Nguyen, Zhao Ren, Trinh Pham, Thanh Trung Huynh, Phi Le Nguyen, Hongzhi Yin, Quoc Viet Hung Nguyen

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Fixed a serious error in author information

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10861 2024-11-19 cs.LG cs.AI cs.SE 62%

See-Saw Generative Mechanism for Scalable Recursive Code Generation with Generative AI

Ruslan Idelfonso Magaña Vsevolodovna

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 18 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11595 2024-11-15 cs.CL cs.AI 62%

Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate

Kai Xiong, Xiao Ding, Yixin Cao, Ting Liu, Bing Qin

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2023 Findings Camera Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08900 2024-11-15 q-bio.GN cs.AI cs.CE cs.LG q-bio.BM 62%

RNA-GPT: Multimodal Generative System for RNA Sequence Understanding

Yijia Xiao, Edward Sun, Yiqiao Jin, Wei Wang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Machine Learning for Structural Biology Workshop, NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08814 2024-11-14 cs.AI cs.LG 62%

Process-aware Human Activity Recognition

Jiawei Zheng, Petros Papapanagiotou, Jacques D. Fleuriot, Jane Hillston

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07462 2024-11-12 cs.CV cs.AI cs.LG 62%

MAN TruckScenes: A multimodal dataset for autonomous trucking in diverse conditions

Felix Fent, Fabian Kuttenreich, Florian Ruch, Farija Rizwin, Stefan Juergens, Lorenz Lechermann, Christian Nissler, Andrea Perl, Ulrich Voll, Min Yan, Markus Lienkamp

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to NeurIPS 2024 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17372 2024-11-12 cs.AI cs.LG cs.RO 62%

BehaviorGPT: Smart Agent Simulation for Autonomous Driving with Next-Patch Prediction

Zikang Zhou, Haibo Hu, Xinhong Chen, Jianping Wang, Nan Guan, Kui Wu, Yung-Hui Li, Yu-Kai Huang, Chun Jason Xue

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11549 2024-11-05 cs.CL cs.AI 62%

How Personality Traits Influence Negotiation Outcomes? A Simulation based on Large Language Models

Yin Jou Huang, Rafik Hadfi

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Findings of EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05954 2024-11-05 cs.AI cs.LG cs.SY eess.SY 62%

Aligning Large Language Models with Representation Editing: A Control Perspective

Lingkai Kong, Haorui Wang, Wenhao Mu, Yuanqi Du, Yuchen Zhuang, Yifei Zhou, Yue Song, Rongzhi Zhang, Kai Wang, Chao Zhang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.04559 2024-11-04 cs.AI cs.CL cs.HC 62%

Can Large Language Model Agents Simulate Human Trust Behavior?

Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye, Shiyang Lai, Kai Shu, Jindong Gu, Adel Bibi, Ziniu Hu, David Jurgens, James Evans, Philip Torr, Bernard Ghanem, Guohao Li

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to Proceedings of NeurIPS 2024. The first two authors contributed equally. 10 pages for main paper, 56 pages including appendix. Project website: https://agent-trust.camel-ai.org

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01475 2024-11-04 cs.LG cond-mat.mtrl-sci cs.AI physics.chem-ph 62%

Are large language models superhuman chemists?

Adrian Mirza, Nawaf Alampara, Sreekanth Kunchapu, Martiño Ríos-García, Benedict Emoekabu, Aswanth Krishnan, Tanya Gupta, Mara Schilling-Wilhelmi, Macjonathan Okereke, Anagha Aneesh, Amir Mohammad Elahi, Mehrdad Asgari, Juliane Eberhardt, Hani M. Elbeheiry, María Victoria Gil, Maximilian Greiner, Caroline T. Holick, Christina Glaubitz, Tim Hoffmann, Abdelrahman Ibrahim, Lea C. Klepsch, Yannik Köster, Fabian Alexander Kreth, Jakob Meyer, Santiago Miret, Jan Matthias Peschel, Michael Ringleb, Nicole Roesner, Johanna Schreiber, Ulrich S. Schubert, Leanne M. Stafast, Dinga Wonanke, Michael Pieler, Philippe Schwaller, Kevin Maik Jablonka

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02518 2024-10-30 cs.SE cs.AI cs.CL cs.CR cs.MA cs.PL 62%

INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness

Hung Le, Yingbo Zhou, Caiming Xiong, Silvio Savarese, Doyen Sahoo

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to The Thirty-Eighth Annual Conference on Neural Information Processing Systems (NeurIPS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14186 2024-10-30 cs.CL cs.AI 62%

Empowering Cross-lingual Abilities of Instruction-tuned Large Language Models by Translation-following demonstrations

Leonardo Ranaldi, Giulia Pucci, Andre Freitas

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Journal ref 2024.findings-acl.473

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19974 2024-10-29 cs.LG cs.CL cs.IR 62%

Evaluating Cost-Accuracy Trade-offs in Multimodal Search Relevance Judgements

Silvia Terragni, Hoang Cuong, Joachim Daiber, Pallavi Gudipati, Pablo N. Mendes

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Journal ref CIKM MMSR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17438 2024-10-24 cs.LG cs.AI 62%

Interpreting Affine Recurrence Learning in GPT-style Transformers

Samarth Bhargav, Alexander Gu

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 21 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14710 2024-10-23 cs.CL cs.AI 62%

ERABAL: Enhancing Role-Playing Agents through Boundary-Aware Learning

Yihong Tang, Jiao Ou, Che Liu, Fuzheng Zhang, Di Zhang, Kun Gai

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2402.10618

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14347 2024-10-21 cs.LG cs.AI 62%

A Scientific Machine Learning Approach for Predicting and Forecasting Battery Degradation in Electric Vehicles

Sharv Murgai, Hrishikesh Bhagwat, Raj Abhijit Dandekar, Rajat Dandekar, Sreedath Panat

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14273 2024-10-21 cs.CL cs.AI cs.CR 62%

REEF: Representation Encoding Fingerprints for Large Language Models

Jie Zhang, Dongrui Liu, Chen Qian, Linfeng Zhang, Yong Liu, Yu Qiao, Jing Shao

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12187 2024-10-18 cs.LG cs.AI 62%

DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs

Yingsong Luo, Ling Chen

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12858 2024-10-18 cs.CL cs.AI 62%

Large Language Models for Medical OSCE Assessment: A Novel Approach to Transcript Analysis

Ameer Hamza Shakur, Michael J. Holcomb, David Hein, Shinyoung Kang, Thomas O. Dalton, Krystle K. Campbell, Daniel J. Scott, Andrew R. Jamieson

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12104 2024-10-17 cs.CY cs.LG cs.SE 62%

To Err is AI : A Case Study Informing LLM Flaw Reporting Practices

Sean McGregor, Allyson Ettinger, Nick Judd, Paul Albee, Liwei Jiang, Kavel Rao, Will Smith, Shayne Longpre, Avijit Ghosh, Christopher Fiorelli, Michelle Hoang, Sven Cattell, Nouha Dziri

专题命中 其他安全 :safety(abstract);分类 cs.CY、cs.LG

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12508 2024-10-17 cs.CL cs.AI cs.CV 62%

MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline

Donghoon Han, Eunhwan Park, Gisang Lee, Adam Lee, Nojun Kwak

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2024 Industry Track Accepted (Camera-Ready Version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10860 2024-10-16 cs.CL cs.AI 62%

A Recipe For Building a Compliant Real Estate Chatbot

Navid Madani, Anusha Bagalkotkar, Supriya Anand, Gabriel Arnson, Rohini Srihari, Kenneth Joseph

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19918 2024-10-16 cs.RO cs.AI cs.LG 62%

CtRL-Sim: Reactive and Controllable Driving Agents with Offline Reinforcement Learning

Luke Rowe, Roger Girgis, Anthony Gosselin, Bruno Carrez, Florian Golemo, Felix Heide, Liam Paull, Christopher Pal

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments CoRL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09649 2024-10-15 cs.CV cs.CL cs.LG 62%

Learning the Bitter Lesson: Empirical Evidence from 20 Years of CVPR Proceedings

Mojtaba Yousefi, Jack Collins

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments NLP4Sceince Workshop, EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03102 2024-10-15 cs.CL cs.AI 62%

"In Dialogues We Learn": Towards Personalized Dialogue Without Pre-defined Profiles through In-Dialogue Learning

Chuanqi Cheng, Quan Tu, Shuo Shang, Cunli Mao, Zhengtao Yu, Wei Wu, Rui Yan

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08698 2024-10-14 cs.CL cs.CY 62%

SocialGaze: Improving the Integration of Human Social Norms in Large Language Models

Anvesh Rao Vijjini, Rakesh R. Menon, Jiayi Fu, Shashank Srivastava, Snigdha Chaturvedi

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08431 2024-10-14 cs.CL cs.AI 62%

oRetrieval Augmented Generation for 10 Large Language Models and its Generalizability in Assessing Medical Fitness

Yu He Ke, Liyuan Jin, Kabilan Elangovan, Hairil Rizal Abdullah, Nan Liu, Alex Tiong Heng Sia, Chai Rick Soh, Joshua Yi Min Tung, Jasmine Chiat Ling Ong, Chang-Fu Kuo, Shao-Chun Wu, Vesela P. Kovacheva, Daniel Shu Wei Ting

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2402.01733

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07638 2024-10-11 cs.LG cs.AI cs.IT math.IT stat.ML 62%

Almost Minimax Optimal Best Arm Identification in Piecewise Stationary Linear Bandits

Yunlong Hou, Vincent Y. F. Tan, Zixin Zhong

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 69 pages. Accepted to NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏