arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7978 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7978 篇

2511.18319 2025-11-25 cs.AI cs.LG cs.SY eess.SY 62%

Weakly-supervised Latent Models for Task-specific Visual-Language Control

弱监督潜在模型用于任务特定的视觉语言控制

Xian Yeow Lee, Lasitha Vidyaratne, Gregory Sin, Ahmed Farahat, Chetan Gupta

机构 * Industrial AI Lab, Hitachi America, Ltd.(日立美国有限公司工业人工智能实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种任务特定的潜在动态模型,利用目标状态监督学习动作诱导位移,以提高空间定位任务中的视觉语言控制性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13814 2025-11-25 cs.HC cs.AI cs.LG 62%

Reversing the Lens: Using Explainable AI to Understand Human Expertise

反转镜头:利用可解释人工智能理解人类专家能力

Roussel Rahman, Aashwin Ananda Mishra, Wan-Lin Hu

机构 * Department of Photon Science(光子科学系) SLAC National Accelerator Laboratory(SLAC国家加速器实验室) Department of Machine Learning(机器学习系)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本研究利用可解释人工智能方法分析人类在调节粒子加速器任务中的学习过程,揭示专家能力发展中的问题分解与策略演变机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06843 2025-11-25 cs.CL cs.AI 62%

Integrating Cognitive Processing Signals into Language Models: A Review of Advances, Applications and Future Directions

将认知处理信号整合进语言模型:关于进展、应用与未来方向的综述

Angela Lopez-Cardona, Sebastian Idesis, Ioannis Arapakis

机构 * Telefónica Scientific Research(Telefónica科学研究院) Universitat Politècnica de Catalunya(加泰罗尼亚理工大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文综述了将认知信号,特别是眼动数据整合到语言模型中的最新进展、应用及未来方向,探讨了其在提升模型性能和解决环境成本方面的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08016 2025-11-24 q-bio.NC cs.AI cs.CL 62%

Emergence of psychopathological computations in large language models

大语言模型中精神病理计算的出现

Soo Yong Lee, Hyunjin Hwang, Taekwan Kim, Yuyeong Kim, Kyuri Park, Jaemin Yoo, Denny Borsboom, Kijung Shin

机构 * KAIST, Kim Jaechul Graudate School of AI(KAIST人工智能研究生院) KAIST, School of Electrical Engineering(KAIST电子工程学院) UCL, Mental Health Neuroscience Department(伦敦大学学院心理健康神经科学系) UvA, Informatics Institute(乌得勒支大学信息学院) UvA, Department of Psychology(乌得勒支大学心理学系)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过建立计算理论框架,证明大语言模型中已出现精神病理学的网络计算结构,并揭示其可能带来的安全风险。

Comments pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15750 2025-11-21 cs.CY cs.AI cs.ET cs.HC 62%

Writing With Machines and Peers: Designing for Critical Engagement with Generative AI

用机器和同伴写作:设计促进生成式AI批判性参与的方法

Xinran Zhu, Cong Wang, Duane Searsmith

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本研究提出一种整合AI与同伴反馈的教学设计,通过八周的写作活动,探讨学生如何批判性地使用生成式AI并建立与AI和人类的协作关系。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13761 2025-11-19 cs.DC cs.AI cs.LG 62%

What happens when nanochat meets DiLoCo?

Alexander Acker, Soeren Becker, Sasho Nedelkoski, Dominik Scheinert, Odej Kao, Philipp Wiesner

机构 * Team exalsius(exalsius团队) Technische Universität Berlin(柏林技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 8pages, 3 figures, technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12630 2025-11-18 cs.CL cs.AI 62%

Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing

Maoqi Liu, Quan Fang, Yang Yang, Can Zhao, Kaiquan Cai

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beihang University(北航) State Key Laboratory of CNS/ATM(国家空管流量管理技术实验室) Aviation Data Communication Corporation(航空数据通信公司)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to Advanced Engineering Informatics

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04998 2025-11-18 cs.CL cs.AI 62%

ProFuser: Progressive Fusion of Large Language Models

Tianyuan Shi, Fanqi Wan, Canbin Huang, Xiaojun Quan, Chenliang Li, Ming Yan, Ji Zhang, Minhua Huang, Wu Kai

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20856 2025-11-17 cs.LG cs.AI 62%

Strada-LLM: Graph LLM for traffic prediction

Seyed Mohamad Moghadas, Bruno Cornelis, Alexandre Alahi, Adrian Munteanu

机构 * Department of Electronics and Informatics, Vrije Universiteit Brussel(布鲁塞尔自由大学电子与信息系) VITA Lab, EPFL(EPFL VITA实验室)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10675 2025-11-17 cs.CL cs.AI cs.IR 62%

Learn to Select: Exploring Label Distribution Divergence for In-Context Demonstration Selection in Text Classification

Ye Jiang, Taihang Wang, Youzheng Liu, Yimin Wang, Yuhan Xia, Yunfei Long

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16657 2025-11-17 cs.CV cs.AI cs.CL 62%

DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation

Zun Wang, Jialu Li, Han Lin, Jaehong Yoon, Mohit Bansal

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments AAAI 2026, Project website: https://zunwang1.github.io/DreamRunner

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10627 2025-11-14 cs.AI cs.CV cs.FL cs.LG 62%

Querying Labeled Time Series Data with Scenario Programs

Edward Kim, Devan Shanker, Varun Bharadwaj, Hongbeen Park, Jinkyu Kim, Hazem Torfah, Daniel J Fremont, Sanjit A Seshia

机构 * University of California, Berkeley(加州大学伯克利分校) Korea University(韩国大学) Chalmers University of Technology(查尔姆斯理工大学) University of Gothenburg(哥德堡大学) University of California, Santa Cruz(加州大学圣克ruz分校)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Journal ref NASA Formal Methods Conference 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10583 2025-11-14 cs.CL cs.AI 62%

Evaluating Prompting Strategies with MedGemma for Medical Order Extraction

Abhinand Balachandran, Bavana Durgapraveen, Gowsikkan Sikkan Sudhagar, Vidhya Varshany J S, Sriram Rajkumar

机构 * EXL Health AI Lab at MEDIQA-OE 2025(EXL健康AI实验室)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments 2 figures 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09796 2025-11-14 cs.CL cs.AI 62%

Predicate-Argument Structure Divergences in Chinese and English Parallel Sentences and their Impact on Language Transfer

Rocco Tripodi, Xiaoyu Liu

机构 * Department of Environmental Sciences, Informatics and Statistics(环境科学、信息学与统计学系) Department of Linguistic Sciences And Foreign Literatures(语言科学与外国文学系)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08402 2025-11-12 cs.CV cs.AI cs.LG 62%

Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation

Difei Gu, Yunhe Gao, Mu Zhou, Dimitris Metaxas

机构 * Rutgers University(新泽西罗格斯大学) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08052 2025-11-12 cs.AI cs.CL cs.SE 62%

Dual-Process Scaffold Reasoning for Enhancing LLM Code Debugging

Po-Chung Hsieh, Chin-Po Chen, Jeng-Lin Li, Ming-Ching Chang

机构 * National Taiwan University(国立台湾大学) AI Research Center, Inventec Corporation(Inventec公司人工智能研究中心) University at Albany, SUNY(纽约州立大学阿尔巴尼分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07166 2025-11-11 cs.CL cs.AI cs.CE 62%

AdaRec: Adaptive Recommendation with LLMs via Narrative Profiling and Dual-Channel Reasoning

Meiyun Wang, Charin Polpanumas

机构 * The University of Tokyo(东京大学) Amazon.com, Inc.(亚马逊公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07080 2025-11-11 cs.CL cs.AI 62%

Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora

Khalil Hennara, Ahmad Bastati, Muhammad Hreden, Mohamed Motasim Hamed, Zeina Aldallal, Sara Chrouf, Safwan AlModhayan

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06497 2025-11-11 cs.CL cs.AI 62%

Rethinking what Matters: Effective and Robust Multilingual Realignment for Low-Resource Languages

Quang Phuoc Nguyen, David Anugraha, Felix Gaschi, Jun Bin Cheng, En-Shiun Annie Lee

机构 * Ontario Tech University(安大略技术大学) Stanford University(斯坦福大学) SAS Posos University of Toronto(多伦多大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06215 2025-11-11 cs.CL cs.AI 62%

Explicit Knowledge-Guided In-Context Learning for Early Detection of Alzheimer's Disease

Puzhen Su, Yongzhu Miao, Chunxi Guo, Jintao Tang, Shasha Li, Ting Wang

机构 * College of Computer Science and Technology, National University of Defense Technology(计算机科学与技术学院,国防科技大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments This paper was accepted by IEEE BIBM 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05570 2025-11-11 cs.CV cs.CY cs.LG 62%

Do Street View Imagery and Public Participation GIS align: Comparative Analysis of Urban Attractiveness

Milad Malekzadeh, Elias Willberg, Jussi Torkko, Silviya Korpilo, Kamyar Hasanzadeh, Olle Järv, Tuuli Toivonen

专题命中 其他安全 :alignment(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02009 2025-11-11 cs.CY cs.CL 62%

Urban Computing in the Era of Large Language Models

Zhonghang Li, Lianghao Xia, Xubin Ren, Jiabin Tang, Tianyi Chen, Yong Xu, Chao Huang

机构 * South China University of Technology(华南理工大学) The University of Hong Kong(香港大学) City University of Hong Kong(城市大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.CY

Comments https://github.com/HKUDS/Awesome-LLM4Urban-Papers

Journal ref ACM TIST, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12547 2025-11-11 cs.CL cs.AI 62%

Revealing emergent human-like conceptual representations from language prediction

Ningyu Xu, Qi Zhang, Chao Du, Qiang Luo, Xipeng Qiu, Xuanjing Huang, Menghan Zhang

机构 * College of Computer Science and Artificial Intelligence(计算机科学与人工智能学院) Fudan University(复旦大学) Institute of Modern Languages and Linguistics(现代语言与语言学研究所) Shanghai Key Laboratory of Intelligent Information Processing(上海智能信息处理重点实验室) Research Institute of Intelligent Complex Systems(智能复杂系统研究所) Institute of Trustworthy Embodied Artificial Intelligence(可信具身人工智能研究所) State Key Laboratory of Genetics and Development of Complex Phenotypes(复杂表型遗传与发育国家重点实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 66 pages. Accepted manuscript. Final version published in Proceedings of the National Academy of Sciences (PNAS): https://www.pnas.org/doi/10.1073/pnas.2512514122

Journal ref Proceedings of the National Academy of Sciences, U.S.A., 122 (44) e2512514122 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04995 2025-11-10 cs.HC cs.AI cs.CL 62%

Enhancing Public Speaking Skills in Engineering Students Through AI

Amol Harsh, Brainerd Prince, Siddharth Siddharth, Deepan Raj Prabakar Muthirayan, Kabir S Bhalla, Esraaj Sarkar Gupta, Siddharth Sahu

机构 * Center for Thinking, Language and Communication(思考、语言与交流中心) Plaksha University(普拉克斯大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22780 2025-11-10 cs.AI cs.CL cs.HC 62%

How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations

Zora Zhiruo Wang, Yijia Shao, Omar Shaikh, Daniel Fried, Graham Neubig, Diyi Yang

机构 * Carnegie Mellon University(卡内基梅隆大学) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03684 2025-11-07 cs.CE cs.AI cs.LG cs.SY eess.SY 62%

Simulation-Based Validation of an Integrated 4D/5D Digital-Twin Framework for Predictive Construction Control

Atena Khoshkonesh, Mohsen Mohammadagha, Navid Ebrahimi

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03559 2025-11-06 cs.CL cs.AI 62%

AILA--First Experiments with Localist Language Models

Joachim Diederich

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03159 2025-11-06 cs.LG cs.AI 62%

CoTox: Chain-of-Thought-Based Molecular Toxicity Reasoning and Prediction

Jueon Park, Yein Park, Minju Song, Soyon Park, Donghyeon Lee, Seungheun Baek, Jaewoo Kang

机构 * Department of Computer Science and Engineering(计算机科学与工程系) AIGEN Sciences(AIGEN公司)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to IEEE BIBM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05588 2025-11-06 cs.RO cs.AI cs.LG 62%

Deep Learning Warm Starts for Trajectory Optimization on the International Space Station

Somrita Banerjee, Abhishek Cauligi, Marco Pavone

机构 * Apple(苹果公司) Johns Hopkins University(约翰霍普金斯大学) Stanford University(斯坦福大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to 2025 International Conference on Space Robotics (iSpaRo). Presented at RSS 2025 Workshop on Space Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01894 2025-11-05 cs.GR cs.AI cs.LG 62%

LGCC: Enhancing Flow Matching Based Text-Guided Image Editing with Local Gaussian Coupling and Context Consistency

Fangbing Liu, Pengfei Duan, Wen Li, Yi He

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏