arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9311 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9311 篇

2504.15211 2026-01-23 cs.AI stat.AP 79%

Embracing Ambiguity: Bayesian Nonparametrics and Stakeholder Participation for Ambiguity-Aware Safety Evaluation

拥抱模糊性:基于贝叶斯非参数方法和利益相关者参与的模糊性感知安全评估

Yanan Long

专题命中 安全评测 :safety(title);trustworthy(abstract);分类 cs.AI

AI总结 本文提出了一种基于贝叶斯非参数方法和利益相关者参与的框架,用于模糊性感知的安全评估,通过量化尾部风险和整合利益相关者偏好提升生成模型的可信度。

Comments AAAI 2026 workshop MURE

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14127 2026-01-21 cs.CV cs.CL 79%

The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning

智能的副作用:MLLMs多图像推理中的安全风险

Renmiao Chen, Yida Lu, Shiyao Cui, Xuan Ouyang, Victor Shea-Jay Huang, Shumin Zhang, Chengwei Pan, Han Qiu, Minlie Huang

机构 * CoAI group, DCST, Tsinghua University(清华大学) Beihang University(北航) Tsinghua University(清华大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

AI总结 研究发现,多图像推理能力越强的模型在安全测试中越容易产生不安全响应,揭示了模型在任务解决中可能忽视安全约束的风险。

Comments *15 pages, 5 figures. Introduces MIR-SafetyBench (2,676 instances; 9 multi-image relations). Equal contribution; †Corresponding author. Code/data: https://github.com/thu-coai/MIR-SafetyBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13669 2026-01-21 cs.CL 79%

CommunityBench: Benchmarking Community-Level Alignment across Diverse Groups and Tasks

CommunityBench: 跨多样化群体和任务的社区级对齐基准测试

Jiayu Lin, Zhongyu Wei

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

AI总结 CommunityBench通过四个基于共同身份和共同纽带理论的任务,评估了社区级对齐方法,揭示了现有LLMs在建模社区特定偏好上的局限性,并探索了社区级对齐在促进个体建模中的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12775 2026-01-21 cs.CL 79%

Persistent Personas? Role-Playing, Instruction Following, and Safety in Extended Interactions

持久的人格?角色扮演、指令遵循与在扩展互动中的安全性

Pedro Henrique Luz de Araujo, Michael A. Hedderich, Ali Modarressi, Hinrich Schuetze, Benjamin Roth

机构 * University of Vienna, Faculty of Computer Science(维也纳大学计算机科学系) Doctoral School Computer Science, University of Vienna(维也纳大学计算机科学博士学院) Center for Information and Language Processing, LMU Munich(慕尼黑大学信息与语言处理中心) Munich Center for Machine Learning(慕尼黑机器学习中心) University of Vienna, Faculty of Philological and Cultural Studies(维也纳大学文学与文化研究系)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

AI总结 研究探讨了在长对话中人格保持的稳定性,发现随着对话延长,人格保真度下降,且非人格基线模型在初期表现更优。

Comments 31 pages, 35 figures, accepted to EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12960 2026-01-21 cs.CL 79%

Trustworthy Data-driven Chronological Age Estimation from Panoramic Dental Images

可信的数据驱动全景牙科图像Chronological年龄估计

Ainhoa Vivel-Couso, Nicolás Vila-Blanco, María J. Carreira, Alberto Bugarín-Diz, Inmaculada Tomás, Jose M. Alonso-Moral

机构 * Centro Singular de Investigación en Tecnoloxías Intelixentes (CiTIUS), Universidade de Santiago de Compostela(智能技术研究中心(CiTIUS)、圣地亚哥-德孔波斯特拉大学) Department of Electronics and Computing, Universidade de Santiago de Compostela(电子与计算系、圣地亚哥-德孔波斯特拉大学) Oral Sciences Research Group, Universidade de Santiago de Compostela(口腔科学研究组、圣地亚哥-德孔波斯特拉大学) Instituto de Investigación Sanitaria de Santiago de Compostela (IDIS), Universidade de Santiago de Compostela(圣地亚哥-德孔波斯特拉卫生研究院(IDIS)、圣地亚哥-德孔波斯特拉大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL

AI总结 本文提出了一种结合不透明和透明方法的系统,通过生成临床医生友好的文本解释来提高全景牙科图像中Chronological年龄估计的可信度。

Comments This paper is a preliminary version of an accepted article in Information Systems Frontiers, Springer, Special Issue "Explainability in Human-Centric AI". Please cite the final published version of the paper, not this preprint. The final published version can be found at https://link.springer.com/article/10.1007/s10796-025-10682-3

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12392 2026-01-21 cs.AI 79%

PsychēChat: An Empathic Framework Focused on Emotion Shift Tracking and Safety Risk Analysis in Psychological Counseling

PsychēChat: 一种专注于心理辅导中情感转变跟踪和安全风险分析的共情框架

Zhentao Xia, Yongqi Fan, Yuxiang Chu, Yichao Yin, Liangliang Chen, Tong Ruan, Weiyan Zhang

机构 * East China University of Science and Technology(东华大学) Shanghai Changning Mental Health Center, Affiliated Mental Health Center of East China Normal University(上海昌宁心理健康中心,华东师范大学附属心理健康中心)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

AI总结 PsychēChat通过情感管理与风险控制模块,提升心理辅导中情感转变跟踪和安全风险分析的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11286 2026-01-19 cs.AI 79%

XChoice: Explainable Evaluation of AI-Human Alignment in LLM-based Constrained Choice Decision Making

XChoice: 用于基于LLM的受限选择决策中的可解释性AI-人类对齐评估

Weihong Qi, Fan Huang, Rasika Muralidharan, Jisun An, Haewoon Kwak

机构 * Indiana University Bloomington(印第安纳大学布卢明顿分校)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 XChoice是一种用于评估基于LLM的受限选择决策中AI-人类对齐的可解释框架,通过机制模型恢复可解释参数以诊断不一致并支持改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10460 2026-01-16 cs.CL cs.AI cs.CY cs.LG 79%

Contextual StereoSet: Stress-Testing Bias Alignment Robustness in Large Language Models

上下文立体集:在大型语言模型中压力测试偏见对齐的鲁棒性

Abhinaba Basu, Pavan Chakraborty

机构 * Indian Institute of Information Technology, Allahabad (IIITA)(印度信息与技术研究所(Allahabad)) National Institute of Electronics and Information Technology (NIELIT)(国家电子与信息技术研究所)

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI、cs.CY

AI总结 上下文立体集通过压力测试大型语言模型在不同上下文中的偏见对齐鲁棒性,揭示固定条件测试的偏见分数可能无法推广。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02602 2026-01-15 stat.ML astro-ph.IM cs.LG stat.AP stat.ME 79%

Trustworthy scientific inference with generative models

可信的生成模型科学推断

James Carzon, Luca Masserano, Joshua D. Ingram, Alex Shen, Antonio Carlos Herling Ribeiro Junior, Tommaso Dorigo, Michele Doro, Joshua S. Speagle, Rafael Izbicki, Ann B. Lee

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

AI总结 FreB通过提供可解释的诊断和有效性保证,实现可信的科学推断,尤其在直接似然评估不可行的领域。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04211 2026-01-09 cs.CL 79%

Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays

Qwerty AI:可解释的自动年龄评级和俄语剧本内容安全评估系统

Nikita Zmanovskii

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

AI总结 Qwerty AI通过可解释的自动化系统实现俄语剧本的年龄评级和内容安全评估,利用微调模型和量化技术,在严格限制下实现高准确率和高效处理。

Comments 15 pages, 7 tables, 1 figure, 4 appendices. System paper describing automated age-rating for Russian screenplays using fine-tuned Phi-3-mini. Includes baseline comparisons, human evaluation, and production deployment. Code and model weights available at https://github.com/nikita-zmanovskiy/qwertyAI. Developed during Wink Hackathon, November 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03281 2026-01-08 eess.SY cs.AI cs.SY 79%

$α^3$-Bench: A Unified Benchmark of Safety, Robustness, and Efficiency for LLM-Based UAV Agents over 6G Networks

$α^3$-Bench:一种统一的安全性、鲁棒性和效率评估基准,用于基于大语言模型的无人机代理在6G网络中的自主性

Mohamed Amine Ferrag, Abderrahmane Lakas, Merouane Debbah

机构 * Department of Computer and Network Engineering, College of Information Technology, United Arab Emirates University(计算机与网络工程系,信息科技学院,阿联酋大学) Khalifa University of Science and Technology(科技大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

AI总结 本文提出$α^3$-Bench,用于评估基于LLM的无人机代理在6G网络中的安全性、鲁棒性和效率,通过多轮对话推理和控制问题测试其性能。

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02215 2026-01-06 cs.SE cs.AI 79%

LLM-Empowered Functional Safety and Security by Design in Automotive Systems

基于大语言模型的汽车系统功能安全与安全设计

Nenad Petrovic, Vahid Zolfaghari, Fengjunjie Pan, Alois Knoll

机构 * European Chips Joint Undertaking(欧洲芯片联合计划) Federal Ministry of Research, Technology and Space of Germany(德国联邦科研、技术和空间 ministry)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

AI总结 本文提出利用大语言模型提升汽车系统功能安全与安全设计的方法,通过事件链模型和MDE方法实现安全拓扑设计和代码分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01024 2026-01-06 cs.CV cs.AI cs.IR 79%

ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval

ITSELF: 基于注意力引导的细粒度对齐用于视觉-语言检索

Tien-Huy Nguyen, Huu-Loc Tran, Thanh Duc Ngo

机构 * University of Information Technology(信息技术大学) Vietnam National University(越南国家大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 ITSELF通过基于注意力引导的细粒度对齐框架,在视觉-语言检索任务中实现最先进的性能和跨数据集泛化能力。

Comments Accepted at WACV Main Track 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00926 2026-01-06 cs.IR cs.AI 79%

MACA: A Framework for Distilling Trustworthy LLMs into Efficient Retrievers

MACA:一种将可信大语言模型提炼为高效检索器的框架

Satya Swaroop Gudipudi, Sahil Girhepuje, Ponnurangam Kumaraguru, Kristine Ma

机构 * JP Morgan Chase(摩根大通公司) IIIT Hyderabad(海得拉巴印度理工学院)

专题命中 安全评测 :trustworthy(title);alignment(abstract);分类 cs.AI

AI总结 MACA通过提炼可信LLM为高效检索器,提升检索准确率并减少LLM调用成本

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00588 2026-01-06 cs.CL 79%

CSSBench: Evaluating the Safety of Lightweight LLMs against Chinese-Specific Adversarial Patterns

CSSBench: 评估轻量级大语言模型对中文特定对抗模式的安全性

Zhenhong Zhou, Shilinlu Yan, Chuanpu Liu, Qiankun Li, Kun Wang, Zhigang Zeng

机构 * Nanyang Technological University(南洋理工大学) Beijing University of Posts and Telecommunications(北京邮电大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

AI总结 CSSBench通过评估中文特定对抗模式,揭示轻量级大语言模型在中文环境下的安全挑战,为实际应用提供安全评估框架。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19991 2026-01-06 cs.SD cs.LG 79%

SAMUeL: Efficient Vocal-Conditioned Music Generation via Soft Alignment Attention and Latent Diffusion

SAMUeL:通过软对齐注意力和潜在扩散实现高效的语音条件音乐生成

Hei Shing Cheung, Boya Zhang, Jonathan H. Chan

机构 * Division of Engineering Science University of Toronto(工程科学系 马尼托瓦大学) Innovative Cognitive Computing Center King Mongkut's University of Technology Thonburi(创新认知计算中心 王后大学技术吞武里)

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

AI总结 SAMUeL通过软对齐注意力和潜在扩散技术,实现高效的语音条件音乐生成,参数减少220倍,推理速度提升52倍,性能优于OpenAI Jukebox。

Comments 7 pages, 3 figures, accepted to IEEE/WIC WI-IAT

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00513 2026-01-05 cs.LG 79%

When Small Models Are Right for Wrong Reasons: Process Verification for Trustworthy Agents

当小模型适用于错误的原因:信任代理的过程验证

Laksh Advani

机构 * Laksh Advani

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

AI总结 本文揭示小型语言模型在推理过程中的可靠性问题,提出推理完整性评分,并展示RAG提升推理质量而元认知可能损害性能,强调过程验证对可信代理的重要性。

Comments Accepted to Trustagent workshop AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21842 2026-01-05 cs.CL 79%

AlignAR: Generative Sentence Alignment for Arabic-English Parallel Corpora of Legal and Literary Texts

AlignAR: 生成式句子对齐用于法律和文学文本的阿拉伯-英语平行语料库

Baorong Huang, Ali Asiri

机构 * School of Foreign Languages, Huaihua University(怀化大学外国语言学院) Al-lith University College, Umm al-Qura University(乌姆·盖拉大学阿尔利思学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

AI总结 AlignAR通过生成式句子对齐方法,提升阿拉伯-英语法律和文学文本的平行语料库质量,验证了大语言模型在对齐任务中的优越性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20144 2026-01-05 cs.CL 79%

Multi-hop Reasoning via Early Knowledge Alignment

通过早期知识对齐实现多跳推理

Yuxin Wang, Shicheng Fang, Bo Wang, Qi Luo, Xuanjing Huang, Yining Zheng, Xipeng Qiu

机构 * Computer Science, Fudan University(复旦大学计算机科学系) Institute of Modern Languages and Linguistics, Fudan University(复旦大学现代语言与语言学研究所) Shanghai SII(上海SII)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

AI总结 EKA通过早期知识对齐提升多跳推理效率,增强检索精度和系统性能。

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00089 2025-12-30 cs.LG stat.ME 79%

A new machine learning framework for occupational accidents forecasting with safety inspections integration

一种整合安全检查的新型机器学习框架用于职业事故预测

Aho Yapi, Pierre Latouche, Arnaud Guillin, Yan Bailly

机构 * Laboratoire de Mathématique Blaise Pascal UMR 6620 CNRS, Université Clermont Auvergne(布列塔尼数学实验室(皮萨实验室))

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

AI总结 本文提出一种整合安全检查的新型机器学习框架,用于预测职业事故,通过将安全检查数据转化为二元时间序列,实现短期风险信号的生成与评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21220 2025-12-29 cs.AI cs.CV cs.RO 79%

RoboSafe: Safeguarding Embodied Agents via Executable Safety Logic

RoboSafe: 通过可执行的安全逻辑保障具身智能体

Le Wang, Zonghao Ying, Xiao Yang, Quanchen Zou, Zhenfei Yin, Tianlin Li, Jian Yang, Yaodong Yang, Aishan Liu, Xianglong Liu

机构 * Beihang University(北航大学) AI Security Lab(360AI安全实验室) The University of Sydney(悉尼大学) Nanyang Technological University(南洋理工大学) Peking University(北京大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

AI总结 RoboSafe通过可执行的安全逻辑保障具身智能体,减少危险行为并保持任务性能。

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21127 2025-12-25 cs.AI 79%

A Real-World Evaluation of LLM Medication Safety Reviews in NHS Primary Care

对NHS初级医疗中基于LLM的药物安全审查系统的真实世界评估

Oliver Normand, Esther Borsi, Mitch Fruin, Lauren E Walker, Jamie Heagerty, Chris C. Holmes, Anthony J Avery, Iain E Buchan, Harry Coppock

机构 * i.AI, Department for Science, Innovation, and Technology(i.AI,科学、创新与技术部门) Centre for Experimental Therapeutics, University of Liverpool(实验治疗中心,利物浦大学) Civic Health Innovation Labs, University of Liverpool(公民健康创新实验室,利物浦大学) Downing Street(唐宁街10号) Department of Statistics, University of Oxford(统计系,牛津大学) Ellison Institute of Technology(埃利森技术研究所) Centre for Academic Primary Care, University of Nottingham(学术初级护理中心,诺丁汉大学) The UK AI Security Institute(英国人工智能安全研究所) Department of Computing, Imperial College London(计算系,伦敦帝国学院)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

AI总结 本文评估了基于LLM的药物安全审查系统在真实临床数据中的表现,揭示了其在不同复杂性下的失败模式,强调了上下文推理的重要性及改进方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09016 2025-12-25 cs.SD cs.AI eess.AS 79%

DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment

DiTSinger: 通过扩散变换器和隐式对齐扩展歌唱语音合成

Zongcai Du, Guilin Deng, Xiaofeng Guo, Xin Gao, Linke Li, Kaichang Cheng, Fubo Han, Siyu Yang, Peng Liu, Pan Zhong, Qiang Fu

机构 * Migu Music, China Mobile Communications Corporation, China(咪咕音乐,中国移动通信集团,中国)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 DiTSinger通过扩散变换器和隐式对齐机制实现高效且高保真的歌唱语音合成,解决了数据稀缺和模型扩展性问题。

Comments ICASSP26 under review. Demo page: https://nju-jet.github.io/DiTSinger

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19663 2025-12-23 cs.CV cs.AI 79%

Beyond CLIP: Knowledge-Enhanced Multimodal Transformers for Cross-Modal Alignment in Diabetic Retinopathy Diagnosis

超越CLIP:基于知识的多模态Transformer用于糖尿病视网膜病变诊断中的跨模态对齐

Argha Kamal Samanta, Harshika Goyal, Vasudha Joshi, Tushar Mungle, Pabitra Mitra

机构 * Department of Medicine Stanford University Stanford, USA(医学系 斯坦福大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出一种基于知识的多模态Transformer框架,通过整合视网膜图像、临床文本和结构化数据,提升糖尿病视网膜病变诊断中的跨模态对齐与检索性能。

Comments 14 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15210 2025-12-23 cs.CL cs.IR 79%

Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphs

对先验的探讨:大型语言模型在知识图谱上的可信推理

Jie Ma, Ning Qu, Zhitao Gao, Rui Xing, Jun Liu, Hongbin Pei, Jiang Xie, Linyun Song, Pinghui Wang, Jing Tao, Zhou Su

机构 * MOE KLINNS Lab, Xi’an Jiaotong University(MOE KLINNS实验室,西安交通大学) School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) Shaanxi Province Key Laboratory of Big Data Knowledge Engineering(陕西省大数据知识工程重点实验室) School of Artificial Intelligence, Chongqing University of Post and Telecommunications(人工智能学院,重庆邮电大学) School of Computer Science, Northwestern Polytechnical University(计算机学院,西北工业大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL

AI总结 本研究提出DP框架,通过整合知识图谱的结构和约束先验,提升大型语言模型在知识图谱上的推理准确性和响应可靠性。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17795 2025-12-22 cs.DL cs.AI cs.IR 79%

Intelligent Knowledge Mining Framework: Bridging AI Analysis and Trustworthy Preservation

智能知识挖掘框架:弥合人工智能分析与可信保存之间的鸿沟

Binh Vu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

AI总结 本文提出智能知识挖掘框架,旨在弥合人工智能分析与可信保存之间的鸿沟,通过双流架构实现数据的高效挖掘与可信存档。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17172 2025-12-22 cs.HC cs.AI 79%

PILAR: Personalizing Augmented Reality Interactions with LLM-based Human-Centric and Trustworthy Explanations for Daily Use Cases

PILAR: 基于LLM的人本化和可信的增强现实交互个性化

Ripan Kumar Kundu, Istiak Ahmed, Khaza Anuarul Hoque

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

AI总结 PILAR通过基于LLM的个性化解释提升AR交互的可信度和用户体验

Comments Published in the 2025 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16250 2025-12-19 cs.AI cs.MA 79%

AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding

AMUSE:面向代理多说话者理解的音频-视觉基准与对齐框架

Sanjoy Chowdhury, Karren D. Yang, Xudong Liu, Fartash Faghri, Pavan Kumar Anasosalu Vasu, Oncel Tuzel, Dinesh Manocha, Chun-Liang Li, Raviteja Vemulapalli

机构 * University of Maryland, College Park(马里兰大学学院公园分校) Apple(苹果公司)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 AMUSE提出一个面向多说话人理解的音频-视觉基准与对齐框架RAFT,通过代理推理提升多模态模型能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15894 2025-12-19 cs.AI 79%

PediatricAnxietyBench: Evaluating Large Language Model Safety Under Parental Anxiety and Pressure in Pediatric Consultations

儿童焦虑基准:在儿科咨询中评估大语言模型在父母焦虑和压力下的安全性

Vahideh Zolfaghari

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

AI总结 PediatricAnxietyBench评估大语言模型在父母焦虑和压力下的安全性,发现70B模型在诊断和紧急情况识别方面表现更优,但所有模型均存在现实压力下的漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15538 2025-12-18 cs.AI cs.MA 79%

TrafficGamer: Reliable and Flexible Traffic Simulation for Safety-Critical Scenarios with Game-Theoretic Oracles

TrafficGamer: 用于安全关键场景的可靠且灵活的交通仿真方法,基于博弈论 oracle

Guanren Qiao, Guorui Quan, Jiawei Yu, Shujun Jia, Guiliang Liu

机构 * School of Data Science, the Chinese University of Hong Kong, Shenzhen(数据科学学院,香港中文大学(深圳)) University of Manchester(曼彻斯特大学) Shenyang MXNavi Co.,Ltd.(沈阳MXNavi有限公司)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

AI总结 TrafficGamer通过将道路驾驶视为多智能体游戏,实现安全关键场景的可靠且灵活的交通仿真,利用博弈论方法提高仿真保真度和均衡捕捉能力。

详情

展开后加载摘要…

URL PDF HTML 收藏