arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 32006 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 32006 篇

2604.08948 2026-04-23 cs.CL 89%

TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice

TaxPraBen:一种可扩展的基准,用于评估LLM在中文实际税收实践中的结构化表现

Gang Hu, Yating Chen, Haiyan Ding, Wang Gao, Jiajia Huang, Min Peng, Qianqian Xie, Kun Yue

机构 * Yunnan University(云南大学) Wuhan University(武汉大学) Jianghan University(江汉大学) Nanjing Audit University(南京审计大学)

专题命中 评测与基准 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出TaxPraBen,首个专注于中文税收实践的基准,结合10项传统任务和3项创新场景,评估19种LLM在税收实践中的表现,揭示封闭源大参数LLM和中文LLM的优势。

Journal ref ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09373 2026-04-23 cs.CL 89%

The Imperfective Paradox in Large Language Models

大型语言模型中的不完美悖论

Bolei Ma, Yusuke Miyao

机构 * LMU Munich(慕尼黑大学) The University of Tokyo(东京大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL

AI总结 研究揭示大型语言模型在事件复合语义理解上的局限性,发现其依赖表面概率启发式而非逻辑推理,提出ImperfectiveNLI数据集揭示目标导向事件的预测偏差。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19533 2026-04-23 cs.CR cs.AI 89%

Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps

网络防御基准:针对LLM在SecOps中的代理威胁狩猎评估

Alankrit Chona, Igor Kozlov, Ambuj Kumar

机构 * Simbian AI

专题命中 评测与基准 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出网络防御基准,评估LLM代理在威胁狩猎任务中的表现,发现现有模型在无引导情况下无法有效识别恶意事件。

Comments 13 pages, 3 figures, 5 tables. Complete benchmark and hunt traces available on request

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20652 2026-04-23 cs.AI cs.HC econ.GN q-fin.EC 89%

Large Language Models Outperform Humans in Fraud Detection and Resistance to Motivated Investor Pressure

大语言模型在欺诈检测中优于人类且对有动机投资者压力具有抵抗力

Nattavudh Powdthavee

机构 * Nanyang Technological University(南洋理工大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract_cn);分类 cs.AI

AI总结 研究发现大语言模型在欺诈检测中表现优于人类,且在面对有动机投资者压力时未抑制欺诈预警,而人类顾问在压力下更倾向于支持欺诈投资。

Comments 36 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06837 2026-04-22 cs.CL 89%

Persuasion with Large Language Models: A Survey of Empirical Evidence, Study Methodologies, and Ethical Implications

利用大语言模型的说服:对实证证据、研究方法和伦理影响的综述

Sander Noels, Alexander Rogiers, Maarten Buyl, Tijl De Bie

机构 * AIDA-IDLab, Electronics and Information Systems, Ghent University(AIDA-ID实验室,电子与信息系统,根特大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL

AI总结 本文综述了基于大语言模型的说服领域,分析了其在政治、营销等领域的应用,指出其说服力接近或超越人类,并强调了伦理和社会风险,呼吁制定伦理准则和监管框架。

Comments Main changes: - Slightly altered title & author ordering - New section detailing survey methodology - Expanded literature coverage and improved discussion of all references for clarity, precision & conciseness - Removed the "appealing to authority" subsection & integrated its content elsewhere - Overhauled the experimental design section - Significantly expanded success metrics discussion

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19016 2026-04-22 cs.CL 89%

AlignCultura: Towards Culturally Aligned Large Language Models?

AlignCultura: 向文化对齐的大语言模型迈进?

Gautam Siddharth Kashyap, Mark Dras, Usman Naseem

机构 * School of Computing, Macquarie University, Australia(麦考瑞大学计算机学院,澳大利亚)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract_cn);分类 cs.CL

AI总结 本文提出AlignCultura框架,通过构建HHH-English数据集和评估不同模型的文化对齐能力,提升输出的尊重性和文化多样性。

Comments Accepted at ACL Mains 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10687 2026-04-21 cs.AI cs.CY 89%

Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users

为谁安全?重新思考如何为真实用户评估LLM的安全性

Manon Kempermann, Sai Suresh Macharla Vasu, Mahalakshmi Raveenthiran, Theo Farrell, Ingmar Weber

机构 * Interdisciplinary Institute for Societal Computing(社会计算跨学科研究所)

专题命中 评测与基准 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文探讨了评估LLM对真实用户安全性的挑战,指出需考虑用户情境差异,通过测试不同LLM在金融和健康建议上的表现,发现用户背景影响评估结果,且仅依赖用户披露信息不足以提升安全性评估。

Comments Paper accepted at IASEAI'26; please cite that peer-reviewed version instead

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16729 2026-04-21 cs.CV cs.AI 89%

Agentic Large Language Models for Training-Free Neuro-Radiological Image Analysis

基于代理的大型语言模型用于无训练神经放射学图像分析

Ayhan Can Erdur, Daniel Scholz, Jiazhen Pan, Benedikt Wiestler, Daniel Rueckert, Jan C. Peeken

机构 * Department of Radiation Oncology, TUM University Hospital, Munich, Germany(放射肿瘤科,慕尼黑技术大学医院,德国) Chair for AI in Healthcare and Medicine, Technical University of Munich (TUM)(人工智能在医疗和健康领域的主任,慕尼黑技术大学) Chair for AI for Image-Guided Diagnosis and Therapy, Technical University of Munich (TUM)(人工智能在影像引导诊断和治疗领域的主任,慕尼黑技术大学) Munich Center for Machine Learning (MCML), Munich, Germany(慕尼黑机器学习中心(MCML),德国慕尼黑) Department of Computing, Imperial College London, London, UK(计算系,伦敦帝国学院,英国伦敦) Deutsches Konsortium für Translationale Krebsforschung (DKTK), Partner Site Munich, Munich, Germany(德国转化癌症研究联盟(DKTK),慕尼黑分部,德国慕尼黑)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract_cn);分类 cs.AI

AI总结 本文提出一种无训练代理框架,利用外部工具实现脑MRI分析,涵盖预处理、病灶分割和体积分析,并通过公开BraTS数据集评估代理AI在神经放射学任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16387 2026-04-21 cs.IR cs.AI cs.DL 89%

Large language models for post-publication research evaluation: Evidence from expert recommendations and citation indicators

基于大语言模型的发表后研究评估:专家推荐与引用指标的证据

Mengjia Wu, Yi Zhang, Robin Haunschild, Lutz Bornmann

机构 * Australian Artificial Intelligence Institute, Faculty of Engineering and Information Technology, University of Technology Sydney(澳大利亚人工智能研究所,工程与信息科技学院,新南威尔士大学) Max Planck Institute for Solid State Research(马克斯·普朗克固体物理研究所) Science Policy and Strategy Department, Administrative Headquarters of the Max Planck Society(马克斯·普朗克学会科学政策与战略部门,行政总部)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.AI

AI总结 本文探讨大语言模型在发表后同行评审中的应用,通过专家判断和引用指标对比,发现其在粗粒度评估中表现良好,但在细粒度评分上效果下降,监督微调效果最佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05523 2026-04-20 cs.SE cs.AI 89%

Capture the Flags: Family-Based Evaluation of Agentic LLMs via Semantics-Preserving Transformations

捕获旗帜:通过语义保持变换进行基于家庭的代理LLM评估

Shahin Honarvar, Amber Gorzynski, James Lee-Jones, Harry Coppock, Marek Rei, Joseph Ryan, Alastair F. Donaldson

机构 * Department of Computing, Imperial College London(帝国理工学院伦敦计算机系) AI Security Institute(人工智能安全研究所) Royal Grammar School Guildford(格里菲斯皇家文法学校)

专题命中 评测与基准 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出CTF挑战家族,通过语义保持的程序变换生成等效挑战,评估代理鲁棒性和泛化能力,展示了Evolve-CTF工具及13种代理LLM配置的评估结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15460 2026-04-20 cs.HC cs.AI 89%

The Crutch or the Ceiling? How Different Generations of LLMs Shape EFL Student Writings

支点还是天花板?不同世代的LLM如何塑造EFL学生的写作

Hengky Susanto, David James Woo, Chingyi Yeung, Stephanie Wing Yan Lo-Philip, Chi Ho Yeung

机构 * Education University of Hong Kong(教育大学(香港)) Everwrite Limited(Everwrite有限公司) International Christian School(国际基督教学校)

专题命中 评测与基准 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究探讨LLM在辅助EFL学生写作中的作用及局限,发现高级LLM虽提升评分和词汇多样性,但可能掩盖真实能力,建议教学应关注学习过程而非仅输出质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09643 2026-04-17 cs.ET cs.AI 89%

MM-tau-p$^2$: Persona-Adaptive Prompting for Robust Multi-Modal Agent Evaluation in Dual-Control Settings

MM-tau-p$^2$: 人格自适应提示用于双控制设置中多模态代理的鲁棒性评估

Anupam Purwar, Aditya Choudhary

机构 * Sprinklr AI

专题命中 评测与基准 :LLM(summary_cn,abstract);prompting(title,comments);language model(abstract);分类 cs.AI

AI总结 本文提出MM-tau-p$^2$基准,通过12个新指标评估多模态代理在双控制环境中的鲁棒性,结合用户输入解决查询,并展示即使使用前沿LLM,多模态鲁棒性等指标仍需考虑。

Comments A benchmark for evaluating multimodal both voice and text LLM agents in dualcontrol settings. We introduce persona adaptive prompting and 12 new metrics to assess robustness safety efficiency and recovery in customer support scenarios

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02135 2026-04-17 cs.CL 89%

Graph-Based Alternatives to LLMs for Human Simulation

基于图的LLM替代方案:人类模拟

Joseph Suh, Suhong Moon, Serina Chang

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 评测与基准 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文研究了多种封闭式模拟任务,提出图神经网络可匹配或超越LLM方法,GEMS模型在三个数据集上表现优异,参数更少且透明。

Comments Conference: ACL 2026 Long Main Code: https://github.com/schang-lab/gems

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09819 2026-04-16 cs.LG cs.SY eess.SY 89%

FDM-Bench: A Comprehensive Benchmark for Evaluating Large Language Models in Additive Manufacturing Tasks

FDM-Bench:用于评估大型语言模型在增材制造任务中的综合基准

Ahmadreza Eslaminia, Adrian Jackson, Beitong Tian, Avi Stern, Hallie Gordon, Rajiv Malhotra, Klara Nahrstedt, Chenhui Shao

机构 * Department of Mechanical Science and Engineering, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校机械科学与工程系) Coordinated Science Laboratory, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校协调科学实验室) Department of Mechanical and Aerospace Engineering, Rutgers University(罗格斯大学机械与航空航天工程系) Department of Mechanical Engineering, University of Michigan(密歇根大学机械工程系)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.LG

AI总结 本文提出FDM-Bench基准,用于评估大型语言模型在增材制造任务中的能力,通过用户查询和G代码样本评估不同模型的性能,发现闭源模型在检测G代码异常方面表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11287 2026-04-14 cs.AI q-bio.OT 89%

Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Using a Large Language Model

人工智能生成锻炼处方的一致性:使用大型语言模型的重复生成研究

Kihyuk Lee

机构 * Seoul National University Bundang Hospital(首尔大学盆唐医院)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本研究评估了大型语言模型生成锻炼处方在相同条件下的内在一致性,发现语义一致性高,但关键定量成分存在差异,需进一步结构约束和专家验证。

Comments 15 pages, 5 tables, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09606 2026-04-14 cs.AI cs.SE 89%

Evaluating Reliability Gaps in Large Language Model Safety via Repeated Prompt Sampling

通过重复提示采样评估大语言模型安全性的可靠性差距

Keita Broadwater

机构 * Independent Researcher(独立研究员)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出APST框架,通过重复采样相同提示在受控条件下揭示LLM潜在故障模式,量化不同模型和配置下的操作风险,展示浅层基准分数可能掩盖持续使用下的可靠性差异。

Comments 9 pages, 4 figures; accepted at the CCAI 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13987 2026-04-14 cs.CL 89%

RiTeK: A Dataset for Large Language Models Complex Reasoning over Textual Knowledge Graphs in Medicine

RiTeK:一个用于大型语言模型在医学领域文本知识图谱上复杂推理的数据集

Jiatan Huang, Mingchen Li, Zonghai Yao, Dawei Li, Yuxin Zhang, Zhichao Yang, Yongkang Xiao, Feiyun Ouyang, Xiaohan Li, Shuo Han, Hong Yu

机构 * University of Connecticut(康涅狄格大学) University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校) School of Computing, and Augmented Intelligence, Arizona State University(亚利桑那州立大学计算与增强智能学院) UMass Chan Medical School(马萨诸塞大学陈医学院) University of Minnesota(明尼苏达大学) Rollins School of Public Health, Emory University(埃默里大学罗林斯公共卫生学院) Optum AI University of Massachusetts, Lowell(马萨诸塞大学洛厄尔分校)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL

AI总结 本文提出RiTeK数据集,用于评估大型语言模型在医学领域文本知识图谱上的复杂推理能力,通过合成高质量查询和专家评估,揭示现有检索方法的不足。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03107 2026-04-13 cs.CL 89%

The Mask of Civility: Benchmarking Chinese Mock Politeness Comprehension in Large Language Models

礼貌的面具:在大型语言模型中基准测试中文虚伪礼貌理解

Yitong Zhang, Yuhan Xiang, Mingxuan Liu

机构 * Tsinghua University(清华大学) National University of Singapore(新加坡国立大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL

AI总结 本文从语用学角度评估大型语言模型在识别中文礼貌、不礼貌和虚伪礼貌现象中的表现差异,构建了结合真实和模拟中文话语的三类数据集,并通过六种代表性模型在四种提示条件下进行测试。

Comments Preprint

Journal ref International Conference in Applied Language Sciences (ALS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08293 2026-04-10 cs.SE cs.AI 89%

CIAO - Code In Architecture Out - Automated Software Architecture Documentation with Large Language Models

CIAO - 代码输入,架构输出 - 利用大语言模型实现自动化软件架构文档生成

Marco De Luca, Tiziano Santilli, Domenico Amalfitano, Anna Rita Fasolino, Patrizio Pelliccione

机构 * University of Naples Federico II(那不勒斯腓特烈二世大学) University of Southern Denmark(南丹麦大学) Gran Sasso Science Institute(格兰萨索科学研究所)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出CIAO方法,利用大语言模型从GitHub仓库自动生成系统级架构文档,基于ISO/IEC/IEEE标准模板,评估显示文档被开发者认为有价值且准确,但存在图表质量和高层次上下文建模的局限。

Comments Manuscript accepted for the 23rd International Conference on Software Architecture (ICSA 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22416 2026-04-10 cs.CL cs.IR 89%

Hallucination Detection and Evaluation of Large Language Model

大语言模型的幻觉检测与评估

Chenggong Zhang, Haopeng Wang, Hexi Meng

机构 * Department of Electrical and Computer Engineering, University of California, Los Angeles(加州大学洛杉矶分校电气与计算机工程系)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL

AI总结 本文提出HHEM模型,通过轻量级分类框架有效检测大语言模型幻觉,提升效率并保持高准确率,通过对比分析不同模型的检测效果,发现HHEM在时间与准确性上表现优异,但总结任务中存在局部幻觉问题,引入分段检索并分析模型规模与幻觉的关系。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06562 2026-04-09 cs.AI 89%

On Emotion-Sensitive Decision Making of Small Language Model Agents

小语言模型代理的情绪敏感决策研究

Jiaju Lin, Xingjian Du, Qingyun Wu, Ellen Wenting Zou, Jindong Wang

机构 * Pennsylvania State University(宾夕法尼亚州立大学) University of Rochester(罗切斯特大学) William & Mary(威廉与玛丽学院)

专题命中 评测与基准 :language model(title,abstract);small language model(title,abstract);SLM(abstract);分类 cs.AI

AI总结 本文研究小语言模型在情绪影响下的决策机制,通过结合情绪诱导与结构化博弈评估,发现情绪扰动系统影响战略选择,但行为不稳定且不完全符合人类预期。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05755 2026-04-08 cs.SE cs.AI 89%

CAKE: Cloud Architecture Knowledge Evaluation of Large Language Models

CAKE:大型语言模型云架构知识评估

Tim Lukas Adam, Phongsakon Mark Konrad, Riccardo Terrenzi, Florian Girardo Lukas, Rahime Yilmaz, Krzysztof Sierszecki, Serkan Ayvaz

机构 * Centre for Industrial Software(工业软件中心) University of Southern Denmark(南丹麦大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出CAKE基准测试,评估大型语言模型对云原生软件架构的理解,涵盖四个认知层次和五个主题,通过多选和自由回答形式评估22种模型配置,发现多选准确率随参数增加而稳定,自由回答持续区分模型,推理增强提升表现,工具增强则降低小型模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04168 2026-04-08 cs.CL cs.IR 89%

A Semi-Automated Annotation Workflow for Paediatric Histopathology Reports Using Small Language Models

一种用于儿童病理科报告的半自动化标注工作流程使用小型语言模型

Avish Vijayaraghavan, Jaskaran Singh Kawatra, Sebin Sabu, Jonny Sheldon, Will Poulett, Alex Eze, Daniel Key, John Booth, Shiren Patel, Jonny Pearson, Dan Schofield, Jonathan Hope, Pavithra Rajendran, Neil Sebire

机构 * Imperial College London(帝国理工学院) NHS England(英国国家医疗服务体系) Great Ormond Street Hospital(大奥蒙德街儿童医院) University College London(伦敦大学学院)

专题命中 评测与基准 :language model(title,abstract);small language model(title,abstract);large language model(abstract);分类 cs.CL

AI总结 本文提出一种基于小型语言模型的半自动化标注流程,用于从非结构化电子病历数据中提取结构化信息,尤其针对儿童病理科报告,通过临床监督和少量示例提升提取准确性。

Comments 36 pages, includes supplementary information

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03742 2026-04-07 cs.AI 89%

Structured Multi-Criteria Evaluation of Large Language Models with Fuzzy Analytic Hierarchy Process and DualJudge

结构化多标准评估大型语言模型:基于模糊层次分析法与DualJudge

Yulong He, Ivan Smirnov, Dmitry Fedrushkov, Sergey Kovalchuk, Ilya Revin

机构 * St. Petersburg State University(圣彼得堡国立大学) ITMO University(ITMO大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出基于模糊层次分析法的结构化评估方法,通过引入置信度感知的模糊AHP扩展,提升对大型语言模型评估的准确性与稳定性,提出DualJudge框架实现直观与 deliberative评估的互补。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03361 2026-04-07 cs.LG q-bio.QM 89%

The limits of bio-molecular modeling with large language models : a cross-scale evaluation

大语言模型在生物分子建模中的局限性:跨尺度评估

Yaxin Xu, Yue Zhou, Tianyu Zhao, Fengwei An, Zhixiang Ren

机构 * Southern University of Science and Technology(南方科技大学) Pengcheng Laboratory(鹏城实验室) Institute of Mechanics, Chinese Academy of Sciences(中国科学院力学研究所)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.LG

AI总结 本文通过跨尺度生物分子基准测试,揭示了大语言模型在生物分子建模中的性能与机制理解之间的差距,指出链式思维数据对生物任务的有限帮助,混合mamba-注意力架构在长序列中的有效性,以及监督微调对专业化与泛化能力的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03216 2026-04-06 cs.CL 89%

BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence

BAS:一种基于决策理论的评估大语言模型置信度的方法

Sean Wu, Fredrik K. Gustafsson, Edward Phillips, Boyan Gao, Anshul Thakur, David A. Clifton

机构 * Department of Engineering Science, University of Oxford(牛津大学工程科学系) Oxford Suzhou Centre for Advanced Research(牛津大学苏州高等研究院)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL

AI总结 本文提出BAS,一种基于决策理论的评估方法,用于评估大语言模型置信度在不同风险偏好下的决策支持能力,揭示置信度与决策可靠性之间的关系,并展示如何通过干预提升置信度可靠性。

Comments 24 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01366 2026-04-03 cs.AI 89%

CogBias: Measuring and Mitigating Cognitive Bias in Large Language Models

CogBias: 测量和缓解大型语言模型中的认知偏差

Fan Huang, Songheng Zhang, Haewoon Kwak, Jisun An

机构 * Indiana University Bloomington(印第安纳大学伯明顿分校) Singapore Management University(新加坡管理大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文研究大型语言模型中的认知偏差,通过定义四种偏差类型,评估三种模型并发现偏差在不同家族中表现不同,通过激活引导技术减少偏差并保持模型能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07645 2026-04-03 cs.SE cs.AI cs.CR 89%

A Self-Improving Architecture for Dynamic Safety in Large Language Models

为大语言模型动态安全设计的自改进架构

Tyler Slater

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出自改进安全框架SISF,通过反馈循环实现LLM系统在运行时自动检测安全故障并生成防御策略,实验显示其能有效降低攻击成功率,提升系统鲁棒性。

Comments Under review at the journal Information and Software Technology (Special Issue on Software Architecture for AI-Driven Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01108 2026-04-02 cs.AI 89%

Adversarial Moral Stress Testing of Large Language Models

对抗性道德压力测试:大语言模型的伦理鲁棒性评估

Saeid Jamshidi, Foutse Khomh, Arghavan Moradi Dakhel, Amin Nikanjam, Mohammad Hamdaqa, Kawser Wazed Nafi

机构 * SWAT Laboratory, Polytechnique Montréal(蒙特利尔理工学院SWAT实验室) Huawei Distributed Scheduling and Data Engine Lab(华为分布式调度与数据引擎实验室) SæT Laboratory, Polytechnique Montréal(蒙特利尔理工学院SæT实验室)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出AMST框架,通过对抗性多轮交互评估大语言模型的伦理鲁棒性,揭示传统单轮评估无法发现的降级模式,强调分布稳定性与尾部行为的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08206 2026-04-02 cs.AI 89%

EHRStruct: A Comprehensive Benchmark Framework for Evaluating Large Language Models on Structured Electronic Health Record Tasks

EHRStruct:一个全面的评估框架,用于评估大型语言模型在结构化电子健康记录任务上的表现

Xiao Yang, Xuejiao Zhao, Zhiqi Shen

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出EHRStruct框架,用于评估大型语言模型在结构化电子健康记录任务上的性能,包含11种临床任务和2200个样本,分析了影响模型性能的关键因素,并提出EHRMaster方法以提升性能。

Comments 28pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏