arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 2611 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 2611 篇

2607.11594 2026-07-14 cs.AI cs.GR 新提交 89%

MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models

MAGIC:利用大语言模型生成具有可导航多场景的游戏世界并感知过渡

Tsz Hei Fan, Choi Wing Fung, Yuxuan Wan, Shuqing Li, Michael R. Lyu

机构 * Department of Computer Science and Engineering(计算机科学与工程系)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 研究利用大语言模型生成多场景游戏世界时面临的问题,提出MAGIC系统,通过四阶段管道将自然语言提示转化为可运行项目,解决跨场景一致性等问题,在新基准测试中表现出色,生成可执行项目且过渡识别指标优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11459 2026-07-14 eess.SY cs.AI cs.SY 新提交 89%

A Multimodal Dataset for Large Language Model Applications in the Energy Domain

用于能源领域大语言模型应用的多模态数据集

Costas Mylonas, Magda Foti

机构 * UBITECH

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 介绍mAIEnergy多模态数据集,整合文本、图像、时间序列及地理空间等数据,统一为结构化格式并伴有元数据与工作流程,可作能源知识库,遵循FAIR原则,助力能源领域大语言模型应用的研究、建模和决策。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31614 2026-07-01 eess.SY cs.AI cs.SY 新提交 89%

Automating Cause-Effect Specification with Knowledge Graphs and Large Language Models

利用知识图谱和大语言模型自动化因果规范生成

Javal Vyas, Milapji Singh Gill, Mehmet Mercangöz

机构 * Autonomous Industrial Systems Lab, Imperial College London(帝国理工学院自主工业系统实验室)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 提出一种语义AI框架,结合知识图谱与约束大语言模型,自动生成因果逻辑和操作安全叙述,减少手动工作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18108 2026-06-17 astro-ph.IM cs.AI 新提交 89%

Querying an astronomical database using large language models: the ALeRCE text-to-SQL system

使用大语言模型查询天文数据库:ALeRCE文本到SQL系统

P. A. Estevez, J. Espejo-Moreira, S. Sanfeliu-Alvarez, F. Forster, A. M. Munoz Arancibia, G. Cabrera-Vives, F. E. Bauer, A. Bayo, M. Catelan, R. Dastidar, L. Hernandez-Garcia, J. A. Intriago, G. Pignata

机构 * Department of Electrical Engineering, University of Chile, Av. Tupper 2007, Santiago, Chile Millennium Institute of Astrophysics (MAS), Nuncio Monseñor Sótero Sanz 100, Providencia, Santiago, Chile Data Artificial Intelligence Initiative (ID\&IA), Universidad de Chile Center for Mathematical Modeling, Universidad de Chile, Beauchef 851, North building, 7th floor, Santiago 8320000, Chile Departamento de Astronom\'ia, Universidad de Chile, Casilla 36D, Santiago, Chile Department of Computer Science, Universidad de Concepción, Edmundo Larenas 219, Concepción, Chile Center for Data Artificial Intelligence, Universidad de Concepción, Edmundo Larenas 310, Concepción, Chile Heidelberg Institute for Theoretical Studies, Heidelberg, Baden-Württemberg, Germany Instituto de Alta Investigación, Universidad de Tarapacá, Casilla 7D, Arica, 1010000, Chile European Southern Observatory, Karl-Schwarzschild-Strasse 2, 85748 Garching bei München, Germany Instituto de Astrofísica, Facultad de Física, Pontificia Universidad Católica de Chile, Casilla 306, Santiago 22, Chile Centro de Astroingeniería, Pontificia Universidad Católica de Chile, Av. Vicuña Mackenna 4860, 7820436 Macul, Santiago, Chile Instituto de Estudios Astrof\'isicos, Facultad de Ingenier\'ia y Ciencias, Universidad Diego Portales, Av. Ej\'ercito Libertador 441, Santiago, Chile Centro Interdisciplinario de Data Science, Facultad de Ingenier\'ia y Ciencias, Universidad Diego Portales, Av. Ej\'ercito Libertador 441, Santiago, Chile

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract_cn);分类 cs.AI

AI总结 提出基于大语言模型的文本到SQL系统,通过上下文学习和逐步生成框架(模式链接、查询分类、提示分解、自纠正)实现自然语言查询天文数据库,在ALeRCE数据集上评估13个模型,Claude Opus 4.6等表现最佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15762 2026-06-16 cs.CR cs.AI cs.SE 新提交 89%

Snyk VulnBench JS 1.0: Can LLMs Find the Same Bugs Twice?

Snyk VulnBench JS 1.0: LLM 能否两次发现相同的漏洞?

Liran Tal, Johannes Kloos, Arsenii Rudich, Stephen Thoemmes, Manoj Nair

机构 * Snyk

专题命中 评测与基准 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 通过300次重复漏洞扫描实验,评估LLM在相同JavaScript代码上安全审查的可重复性,发现引用匹配结果稳定但额外报告波动大,建议结合确定性SAST使用。

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12821 2026-06-12 cs.AI cs.ET 新提交 89%

GeoNatureAgent Benchmark: Benchmarking LLM Agents for Environmental Geospatial Analysis Across Frontier and Open-Weight Foundation Models

GeoNatureAgent Benchmark:面向前沿与开源基础模型的环境地理空间分析LLM智能体基准测试

Gabriel Diaz-Ireland, Diego Prieto-Herráez, Mario García Peces, Javier Velázquez, Devika Jain

机构 * Universidad Católica de Ávila (UCAV)(阿维拉天主教大学) Johns Hopkins University(约翰霍普金斯大学) Independent Researcher(独立研究者) Center for Geographic Analysis, Harvard University(哈佛大学地理分析中心)

专题命中 评测与基准 :LLM(title,title_cn);foundation model(title);分类 cs.AI

AI总结 提出首个通过结构化工具调用真实API评估环境分析智能体的基准,包含93个任务,发现Claude Sonnet 4领先,但开源模型在成本效益上占优,且比较任务普遍未解决。

Comments Preprint. 10 pages, 8 figures. Submitted to ACM SIGSPATIAL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09351 2026-06-09 cs.CL stat.ME 新提交 89%

In-Context Learning for the Imputation of Public Opinion Data with Large Language Models

基于上下文学习的民意数据插补方法

Tobias Holtdirk, Georg Ahnert, Joseph W Sakshaug, Anna-Carolina Haensch

机构 * LMU Munich(慕尼黑大学) Munich Center for Machine Learning(慕尼黑机器学习中心) University of Mannheim(曼海姆大学) Institute for Employment Research (IAB)(就业研究所(IAB)) University of Maryland, College Park(马里兰大学帕克分校)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract_cn);分类 cs.CL

AI总结 提出通过上下文学习(ICL)插补调查缺失数据,在150个意见变量上评估,相比MICE PMM方法,在所有缺失机制下绝对误差更低,尤其非随机缺失时优势显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07541 2026-06-09 cs.HC cs.AI cs.CV cs.CY cs.MM 新提交 89%

Multimodal Large Language Models as Synthetic Participants in Video-Based Studies: An Evaluation

多模态大语言模型作为视频研究中的合成参与者:一项评估

Prabal Shrestha, Bohan Jiang, Haoning Xue, Huan Liu, Xinyi Zhou

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.AI

AI总结 本研究评估多模态大语言模型在视频感知任务中模拟人类主观评分的表现,发现模型存在偏差且与人类一致性有限。

Comments Accepted to SocialLLM @ ICWSM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11816 2026-08-13 cs.CR cs.AI cs.CL 新提交 89%

How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment

中国起源的多模态视觉语言模型如何在状态对齐中从拒绝转向重构

Guang Yang, Fengchen Liu, Alex Wang, Homa Hosseinmardi, Amir Ghasemian

专题命中 评测与基准 :language model(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);prompting(abstract)

AI总结 该研究构建基准测试9个视觉语言模型,发现中国起源模型更易进行状态对齐重构,且审查从可见的拒绝转向不可见的重构,这对人机交互存在影响。

Comments 41 pages, 31 figures, 9 tables. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08160 2026-08-11 cs.CL cs.AI 新提交 89%

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

大语言模型智能体能坚持剧本吗?交互式叙事中长期一致性的基准测试

Yingpeng Ma, Jianhao Yan, Bei Shi, Ka Hou Kam, Runnan Wang, Xuebo Liu, Yulong Chen, Yue Zhang, Derek F. Wong

专题命中 评测与基准 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 该研究针对LLM智能体在交互式叙事中难以维持长期一致性的问题,构建了NCP-Bench基准测试,发现现有顶尖LLM的承诺保持率较低,存在显著的逻辑冲突问题。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18110 2026-07-21 cs.LG cs.CL 新提交 89%

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

大语言模型作为教练:不可验证任务的体验式学习

Tianzhu Ye, Li Dong, Guanheng Chen, He Zhu, Xun Wu, Shaohan Huang, Furu Wei

机构 * Microsoft Research(微软研究院) Tsinghua University(清华大学) Peking University(北京大学)

专题命中 评测与基准 :LLM(title,summary_cn);post-training(abstract);分类 cs.CL、cs.LG

AI总结 研究不可验证任务,提出体验式学习(EL),将LLM反馈模型从评判转为教练,通过提炼体验知识提供密集监督,在开放式任务上表现优于基于规则的RL,泛化性好且减轻奖励作弊。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21937 2026-06-23 cs.CY cs.AI cs.CL 新提交 89%

Latent Confidence Alignment for LLM Self-Assessment

潜在置信对齐用于大语言模型自我评估

Ting-Yu Chen, Tingting Yu, Pei-Cing Huang, Chan Hsu, Ming-Yen Lin, Yihuang Kang

机构 * Department of Information Management(信息管理系) National Sun Yat-Sen University(国立中山大学) Kaohsiung Medical University Hospital(高雄医学大学附设医院) Kaohsiung Medical University(高雄医学大学)

专题命中 评测与基准 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出基于Rasch模型的潜在置信对齐误差(LCAE)来评估LLM自我评估与潜在错误概率的一致性,并引入项目难度作为外部信号,实验表明该方法在不影响模型能力的情况下提升自我评估质量。

Comments 2026 IEEE 27th International Conference on Information Reuse and Integration for Data Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06379 2026-08-10 cs.HC 新提交 89%

Preventive Care Recommendations by Large Language Models

基于大语言模型的预防性护理建议

Eden Avnat, Elia Yanko, Ori Yoran, Raja-Elie E. Abdulnour

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 本研究对比7种LLMs与医生的预防性护理优先级排序,发现LLMs与医生高度一致但对生活方式干预重视不足,部分模型表现更优,强化效果需价值对齐训练等。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06287 2026-08-07 cs.SE 新提交 89%

Automatic Translation of Unstructured Requirements into Linear Temporal Logic through Large Language Models

基于大语言模型将非结构化需求自动翻译为线性时序逻辑

Alexandra Newcomb, Omar Ochoa

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(abstract)

AI总结 本文评估6种大语言模型,通过少样本提示策略在含15个需求的基准上生成线性时序逻辑公式,验证通用LLMs无需微调即可完成非结构化自然语言转LTL任务,可作为半自动化形式化工作流的前端助手。

Comments Copyright 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28991 2026-08-03 cs.CV 新提交 89%

CAER: Conflict-Aware Evidence Routing with Dual Prefix Experts for Multimodal Large Language Models

CAER:面向多模态大语言模型的冲突感知证据路由,采用双前缀专家机制

Zixuan Liu, Juntao Cai, Xiaoxu Cai, Haishuai Wang, Jiajun Bu

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(abstract)

AI总结 本研究提出与主干无关的CAER框架,通过双前缀专家路由机制实现视觉-语言冲突检测与冲突感知生成,在MMMC及AgriConflict数据集上验证可提升开源多模态大语言模型的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19670 2026-07-23 cs.MA 新提交 89%

Same Game, Different Story: A Minimal Conservative Strategic Robustness Benchmark for Large Language Model Agents

相同博弈,不同故事:大语言模型智能体的最小保守战略稳健性基准

Seyed Pouyan Mousavi Davoudi, Alireza Amiri-Margavi, Amin Gholami Davodi, Hamidreza Hasani Balyani, Arshia Gharagozlou

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 研究大语言模型智能体在战略环境中的可靠性,通过“相同博弈,不同故事”基准,以收益不变框架变化下行动分布不变性定义稳健性,经二次分析已发表数据得出社会关系框架会改变模型行为,应分别评估稳健性与能力。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13820 2026-07-16 cs.SE 新提交 89%

PROBE: Benchmarking Code Generation in Large Language Models

PROBE:大语言模型中代码生成的基准测试

Rodrigo Pato Nogueira, Marco Vieira, João R. Campos

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(abstract)

AI总结 针对大语言模型代码生成评估不足,介绍PROBE基准框架,基于多维度评估代码,用于评估多种模型和语言,发现LLMs虽有成果,但处理难题及小模型在特定语言上有困难,且常因易避免错误失败。

Comments Accepted for publication in Empirical Software Engineering (Springer)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03734 2026-07-07 cs.SE 新提交 89%

Fault Detection and Explainable Classification in Automotive HIL Validation via Denoising Autoencoders and In-Context Large Language Models

通过去噪自编码器和上下文大语言模型进行汽车硬件在环验证中的故障检测与可解释分类

Mohammad Abboush, Hamza Ouarrad, Andreas Rausch

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(abstract)

AI总结 提出用于汽车实时验证中故障检测与分类的通用且可解释的两阶段框架,先以去噪自编码器标记异常,再用大语言模型分类并给出解释,评估显示该方法有效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26960 2026-06-26 cs.NI 新提交 89%

Toward Agentic SysAdmin: Rethinking System Administration with AI Agents

迈向智能系统管理:用AI智能体重新思考系统管理

Gianmaria Frigo, Davide Saladino, Alberto Castagnaro, Francesco Marchiori, Denis Donadel, Luca Pajola, Mauro Conti

专题命中 评测与基准 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);foundation model(abstract)

AI总结 提出NetLLMeval框架,通过实时网络仿真自动评估LLM在系统管理任务中的表现,发现求解器设计显著影响准确性,本地模型在合适配置下可媲美前沿大模型。

Comments Under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26196 2026-06-26 cs.CL cs.AI cs.CV cs.LG cs.MM 新提交 89%

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

从结构到协同:多模态大语言模型中视觉-语言感知范式演进综述

Haoxiang Sun, Tao Wang, Li Yuan, Jian Zhao, Jiancheng Lv

机构 * School of Computer Science, Sichuan University(四川大学计算机学院) School of Electronic and Computer Engineering, Peking University Shenzhen Graduate School(北京大学深圳研究生院电子与计算机工程学院) Institute of Artificial Intelligence (TeleAI), China Telecom and Northwestern Polytechnical University(中国电信与西北工业大学人工智能研究院(TeleAI))

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文系统综述多模态大语言模型中统一视觉-语言感知的范式演进,提出五阶段分类法,梳理各阶段代表性方法,并指出开放挑战与未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26109 2026-06-26 cs.CY cs.MA 新提交 89%

Simulating Eating Disorder Patients with LLMs: Evaluating Psychological Persona Stability in Multi-Turn Conversations

使用LLM模拟饮食障碍患者:评估多轮对话中的心理角色稳定性

Jennifer Haase, Jana Gonnermann-Müller, See Heng Yim, Nicolas Leins, Jan Mendling, Sebastian Pokutta

专题命中 评测与基准 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract)

AI总结 本研究通过饮食障碍案例和双评估框架,发现LLM在模拟患者时过度稳定且不准确,系统性地高估严重程度12-30%,并存在“缺失中间态”问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24655 2026-06-24 cs.CL cs.AI cs.LG cs.PF 新提交 89%

AI-PAVE-Br: Leveraging Large Language Models for Enhanced Product Attribute Value Extraction through a Golden Set Approach

AI-PAVE-Br:利用大语言模型通过黄金集方法增强产品属性值提取

Murilo Gazzola, Hugo Gobato Souto, Samuel Silva, Júlia Schubert Peixoto, Felipe Siqueira, André Luis Pedroso de Morais, Caio Gomes

机构 * LuizaLabs – Center of Excellence in Artificial Intelligence(LuizaLabs人工智能卓越中心) Department of Computing and Informatics – Mackenzie Presbyterian University(计算与信息学系-麦金西 Presbyterian大学) Institute of Mathematics and Computer Sciences - University of São Paulo(数学与计算机科学研究所-圣保罗大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 针对巴西电商产品数据爆炸和复杂性,提出AI-PAVE-Br系统,利用大语言模型和精心标注的黄金集,显著提升葡萄牙语产品属性值提取的准确性。

Journal ref Proceedings of the 15th Symposium in Information and Human Language Technology (STIL 2025), Brazilian Computer Society (SBC), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22269 2026-06-23 cs.CL cs.AI cs.LG 新提交 89%

Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability

评估大语言模型在豪萨语和丰贝语机器翻译中的表现:基准、失败与指标可靠性

Mahounan Pericles Adjovi, Roald Eiselen, Prasenjit Mitra

机构 * Carnegie Mellon University Africa(卡内基梅隆大学非洲校区) North-West University(西北大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究评估了四种大语言模型在英语到豪萨语和丰贝语(两种西非低资源语言)的翻译质量,发现质量因语言而异,模型排名不一致,且自动指标与人类判断的相关性波动大,建议采用多指标评估并注意神经指标的局限性。

Comments 19 pages, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13812 2026-06-15 q-fin.CP q-fin.GN 新提交 89%

CFOs Meet LLMs

CFO 与 LLM 相遇

John R. Graham, Campbell R. Harvey, Manish Jha

专题命中 评测与基准 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract)

AI总结 本研究利用大语言模型扮演特定公司CFO,基于杜克-美联储CFO调查数据预测经济乐观情绪,发现LLM能有效复现个体人类反应,为金融研究和政策提供可扩展的高频预期数据。

Comments 21 pages, 4 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11712 2026-06-11 cs.CL cs.AI cs.LG 新提交 89%

Substrate Asymmetry in User-Side Memory: A Diagnostic Framework

用户侧记忆中的子模块不对称性:一个诊断框架

Youwang Deng

机构 * EpistemicaLab — Independent Research(EpistemicaLab — 独立研究)

专题命中 评测与基准 :RLHF(summary_cn,abstract);LLM(summary_cn,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 提出一个诊断框架,将LLM用户侧记忆分解为行为一致性、事实存在和事实缺失三个正交子模块,发现参数记忆与检索记忆在不同子模块上存在不对称性,且RLHF调优加剧了这种不对称性。

Comments Preprint. Code: https://github.com/EpistemicaLab/substrate-asymmetry-memory

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07341 2026-06-08 cs.CR 新提交 89%

Empirical Evaluation of Large Language Models for Migration of Code Fragments to Post-Quantum Cryptography

大型语言模型在代码片段向后量子密码迁移中的实证评估

Javier Pallarés de Bonrostro, Ana I. González-Tablas, María Isabel González Vasco

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract_cn)

AI总结 评估大型语言模型在将经典密码代码片段迁移至后量子密码中的能力,通过微调GPT-4.1-mini实现92.5%的功能正确率,优于零样本基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11584 2026-08-13 cs.AI 新提交 89%

EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval

EnterpriseRAG:非理想企业检索场景下LLM指令遵循与鲁棒性的基准测试

Huiqi Miao, Xinbao Sun, Bo Wang, Fanyu Meng, Lijun Mei, Na Wu, Di Jin, Chao Deng, Junlan Feng

机构 * Jiutian Research, China Mobile(中国移动九天研究院)

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.AI

AI总结 研究针对企业RAG部署的可靠性问题,推出含983个样本的EnterpriseRAG基准,评估13个LLM发现其指令遵循崩溃,为生产级RAG提供可复现的评估基础

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08801 2026-08-11 cs.CL cs.AR cs.ET 新提交 89%

IDRAAK: From Multi-Agent NLP to Few-Shot Prompting for Semantic Drift Detection in Technical Requirements

IDRAAK:从多智能体自然语言处理到用于技术需求语义漂移检测的少样本提示

Shiva Ahir

专题命中 评测与基准 :LLM(summary_cn,abstract);prompting(title,abstract);分类 cs.CL

AI总结 IDRAAK是基于语义需求表示的可解释框架,通过含6个少样本示例的LLM单次调用检测技术需求语义漂移,性能优于多种替代方案,证明少样本提示是高效替代方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05199 2026-08-07 cs.CR cs.AI 新提交 89%

Post-Hoc Trajectory-Risk Certification for Modular LLM-Based Security Agents

基于模块化大语言模型(LLM)的安全智能体的事后轨迹风险验证

Zhenpeng Li

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.AI

AI总结 本文针对模块化LLM安全智能体的轨迹风险验证问题,提出生成树替代方案,在两阶段入侵检测实验中实现92.7%±2.4%的平均轨迹覆盖率,量化了联合访问缺失的成本。

Comments 18 pages, 11 tables; preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00718 2026-08-04 cs.CR cs.AI cs.MA 新提交 89%

Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling Structural Vulnerabilities in Agentic AI Architectures

多智能体大语言模型流水线中的对抗攻击:揭示智能体AI架构的结构性漏洞

Faisal Haque Bappy, Tahrim Hossain, Tarannum Shaila Zaman, Raiful Hasan, Kamrul Hasan, Tariqul Islam

专题命中 评测与基准 :LLM(title,summary_cn);language model(abstract);分类 cs.AI

AI总结 该研究针对多智能体LLM流水线的结构性漏洞,通过在GPT-5-mini等模型上的实验,发现对抗性漏洞源于架构而非模型能力,推动流水线级防御的发展。

Comments This paper has been accepted at the 2026 IEEE Global Communications Conference (GLOBECOM)

详情

展开后加载摘要…

URL PDF HTML 收藏