arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 31794 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 31794 篇

2311.15614 2023-11-28 cs.CL 91%

FreeAL: Towards Human-Free Active Learning in the Era of Large Language Models

Ruixuan Xiao, Yiwen Dong, Junbo Zhao, Runze Wu, Minmin Lin, Gang Chen, Haobo Wang

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);small language model(abstract)

Comments Accepted to EMNLP 2023 (Main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14329 2026-08-17 cs.CR cs.AI cs.CL cs.CY cs.LG 新提交 91%

A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation

面向基于原则的监管的LLM作为评判者的四轴可信度基准

Dipankar Sarkar

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.CL、cs.AI、cs.LG;large language model(comments);language model(comments)

AI总结 本文提出面向基于原则监管的LLM评判者的四轴可信度基准,发布Principle-Bench并引入Ceca评估器,发现无方法在所有四轴占优,部署级LLM评判者需报告对抗性欺骗与事后校准等指标。

Comments 7 pages, 3 figures. Accepted at the KDD 2026 Workshop on Secure and Trustworthy Large Language Models (SeT-LLM), poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29717 2026-08-13 cond-mat.mtrl-sci cs.AI cs.LG 版本更新 91%

Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop

利用自主LLM研究循环优化专家设计的晶体图网络用于带隙预测

Chenmu Zhang, Boris I. Yakobson

机构 * Department of Materials Science and NanoEngineering(材料科学与纳米工程系)

专题命中 评测与基准 :LLM(title,title_cn);pretraining(abstract);分类 cs.AI、cs.LG

AI总结 提出一个自主LLM研究循环,在MatBench带隙基准上构建了无需外部预训练的最准确模型,超越了所有17个专家设计模型,通过实现元素对特征和空间群嵌入等已知方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00155 2026-08-05 cs.AI cs.LG 交叉投稿 91%

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

AgentStream:自进化大语言模型智能体在流式任务下的表现如何?

Dong Yan, Jian Liang, Dapeng Hu, Ran He, Nicholas Jing Yuan, Qi Zhang, Tieniu Tan

机构 * University of Chinese Academy of Sciences(中国科学院大学) Microsoft(微软公司) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Nanjing University(南京大学)

专题命中 评测与基准 :LLM(title,summary_cn);large language model(abstract);language model(abstract);foundation model(abstract)

AI总结 本文提出AgentStream框架,评估三种流式场景下三种基础模型的五种自进化LLM智能体方法,发现自进化表现受场景、模型能力等影响,为方法选择提供指导。

Comments Code is available at https://github.com/Jasper-Yan/AgentStream

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07314 2026-07-30 cs.CL cs.AI 版本更新 91%

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

MEDIC:对LLM在临床应用中的安全性和实用性领先指标的综合评估

Praveenkumar Kanithi, Clément Christophe, Marco AF Pimentel, Tathagata Raha, Prateek Munjal, Nada Saadi, Hamza A Javed, Svetlana Maslenkova, Nasir Hayat, Ronnie Rajan, Shadab Khan

机构 * M42

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 MEDIC通过综合评估框架揭示LLM在临床应用中的安全性和实用性差异,强调需采用组合方法以应对多维度性能权衡。

Comments Published in Transactions on Machine Learning Research (06/2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28345 2026-06-30 cs.RO cs.AI cs.CL cs.CY 91%

Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients

审计具有文化特定道德梯度的LLM驱动社交机器人

Carmen Ng, Gjergji Kasneci

专题命中 评测与基准 :LLM(title,title_cn);prompting(abstract);分类 cs.CL、cs.AI

AI总结 针对LLM驱动社交机器人在跨文化场景中优先分配资源时的道德偏差,提出基于梯度的多语言审计框架,通过对称控制场景测试模型对文化偏好梯度的区分能力,发现提示工程无法可靠纠正的持续性不对称问题。

Comments Accepted for publication in Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14568 2026-06-29 cs.SE cs.CL cs.LG 版本更新 91%

Given, When, Then, Again: Mining Subscenario Refactoring Candidates in Behaviour-Driven Test Suites with ML Classifiers and LLM-Judge Baselines

在行为驱动软件测试套件中挖掘子场景重构机会:ML分类器和LLM-判断基线

Ali Hassaan Mughal, Noor Fatima, Muhammad Bilal

机构 * Independent Researcher(独立研究者;应用MBA(数据分析),德克萨斯韦斯利安大学) Applied MBA (Data Analytics), Texas Wesleyan University(独立研究者;计算机工程学士,国立科学与技术大学(NUST)) Independent Researcher(独立研究者;管理硕士,慕尼黑技术大学) B.E. Computer Engineering, National University of Sciences and Technology (NUST) Independent Researcher M.Sc. Management, Technical University of Munich

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文通过ML分类器和LLM基线,识别行为驱动开发测试套件中可提取的子场景,量化其在公共BDD生态系统中的普及率。

Comments 31 pages, 10 figures, 6 tables, 56 references. v2: retitled; references corrected and verified; threshold-sensitivity and imbalance-robust metrics added; figures restyled. Code and data (Apache-2.0): https://github.com/amughalbscs16/cukereuse_subscenarios_release (archived: https://doi.org/10.5281/zenodo.20356527). Upstream corpus: https://doi.org/10.5281/zenodo.19754359

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24622 2026-06-24 cs.AI cs.LG 版本更新 91%

Random Rule Forest (RRF): Interpretable and Manageable Ensembles of LLM-Generated Questions for Predicting Success from Unstructured Data

随机规则森林 (RRF): 基于LLM生成问题的可解释且可控集成方法用于从非结构化数据预测成功

Ben Griffin, Aaron Ontoyin Yin, Diego Vidaurre, Ugur Koyluoglu, Joseph Ternasky, Fuat Alican, Yigit Ihlamur

机构 * University of Oxford, United Kingdom Aarhus University, Aarhus, Denmark Centre de Recerca Matem\`atica, Barcelona, Spain Oliver Wyman, New York, United States Vela Research, San Francisco, United States

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 提出随机规则森林 (RRF),利用大语言模型生成简单的是/否问题作为弱学习器,通过等权投票形成可审计的“绿旗”评分卡,在低基准率任务中实现透明且竞争性的预测性能。

Comments 25 pages including appendix, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15851 2026-06-18 cs.CL cs.AI 版本更新 91%

Narrative Theory-Driven LLM Methods for Automatic Story Generation and Understanding: A Survey

叙事理论驱动的LLM方法在自动故事生成与理解中的应用:综述

David Y. Liu, Aditya Joshi, Paul Dawson

机构 * School of Computer Science and Engineering(计算机科学与工程学院) School of Arts and Media(艺术与媒体学院) University of New South Wales (UNSW)(新南威尔士大学)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 综述叙事理论驱动的大语言模型方法在自动故事生成与理解中的应用,分析现状并指出生成任务在理论应用、后训练方法、非虚构叙事及叙事层次等方面落后于理解任务,提出未来方向。

Comments 31 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15579 2026-06-16 cs.AI cs.LG cs.MA cs.SE 新提交 91%

Your Agent Has a Genome: Sequence-Level Behavioral Analysis and Runtime Governance of LLM-Powered Autonomous Agents

你的智能体有基因组:基于序列的LLM驱动自主智能体行为分析与运行时治理

Sidi Deng

机构 * Independent Researcher(独立研究员)

专题命中 评测与基准 :LLM(title,title_cn);language model(abstract);分类 cs.AI、cs.LG

AI总结 提出XEPV序列编码框架,将LLM智能体行为建模为基因组序列,通过n-gram挖掘发现P-X-P高风险模式,设计Governor三层干预系统,使成功率提升6.2%并减少44% token消耗。

Comments 16 pages, 15 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13835 2026-06-15 cs.CL cs.AI cs.MA 新提交 91%

When Plausible Is Not Realistic: Evaluating Human Mobility in LLM-Based Urban Simulation

当合理但不现实:评估基于LLM的城市模拟中的人类移动性

Gustavo H. Santos, Aline Carneiro Viana, Thiago H. Silva

机构 * UTFPR(巴西联邦理工大学) Inria(法国国家信息与自动化研究所) U. of Toronto(多伦多大学)

专题命中 评测与基准 :LLM(title,title_cn);prompting(abstract);分类 cs.CL、cs.AI

AI总结 提出验证框架,通过移动定律、时间节奏等指标评估基于LLM的城市模拟器生成的人类移动模式,发现叙事合理性与经验移动现实性之间存在显著差距。

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09877 2026-06-10 cs.LG cs.CE cs.CL 新提交 91%

Streaming Knowledge Compilation: Proactive Materiality-Scored Pinning for Time-Evolving LLM Wikis

流式知识编译:面向时变LLM维基的主动物质性评分固定

Juan M. Huerta

机构 * Zinnia Tech Solutions(Zinnia科技解决方案)

专题命中 评测与基准 :LLM(title,title_cn);post-training(abstract);分类 cs.CL、cs.LG

AI总结 提出流式知识编译框架,通过物质性信号φ_t主动固定重要文档,在金融和维基百科领域验证O(√T log K)遗憾界,并揭示LLM评判偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16346 2026-06-09 cs.CL cs.LG 版本更新 91%

Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents

有益于故障:测量多轮、多语言LLM代理中的非法协助

Nivya Talokar, Ayush K Tarun, Murari Mandal, Maksym Andriushchenko, Antoine Bosselut

机构 * EPFL(苏黎世联邦理工学院) independent(独立研究员) tubingen(图宾根大学)

专题命中 评测与基准 :LLM(title,title_cn);prompting(abstract);分类 cs.CL、cs.LG

AI总结 本文提出STING框架,用于评估多轮多语言LLM代理在执行非法任务时的协助能力,发现低资源语言中攻击成功率不一致,提供实际部署中的压力测试方法。

Comments Accepted in ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23694 2026-06-04 cs.AI cs.CL cs.CR 91%

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents

SafeSearch: 基于LLM的搜索代理的自动化红队测试

Jianshuo Dong, Sheng Guo, Hao Wang, Xun Chen, Zhuotao Liu, Tianwei Zhang, Ke Xu, Minlie Huang, Han Qiu

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 评测与基准 :LLM(title,title_cn);prompting(abstract);分类 cs.CL、cs.AI

AI总结 提出SafeSearch自动化红队框架,系统评估基于LLM的搜索代理在五个风险类别中的安全性,发现GPT-4.1-mini在搜索工作流中攻击成功率高达90.5%,且常见防御措施效果有限。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29815 2026-05-29 cs.AI cs.CL 91%

PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing

PRAIB: 大语言模型辅助审稿行为的同行评审AI基准

Krzysztof Żurawicki, Julia Farganus, Arkadiusz Gaweł, Mateusz Bystroński, Tomasz Jan Kajdanowicz

机构 * Department of Artificial Intelligence(人工智能系)

专题命中 评测与基准 :LLM(title,summary_cn);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 提出PRAIB框架,通过定义审稿特异性、风格和参与行为的指标,并基于11000条机器生成审稿与人类审稿的对比实验,揭示LLM审稿在评分、交叉引用和弱点识别方面与人类审稿的系统性差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29420 2026-05-29 cs.AI cs.LG 91%

When Does Persona Prompting Actually Help? A Retrieval and Metric Analysis of Expert Role Injection in LLMs

角色提示何时真正有效?LLM中专家角色注入的检索与度量分析

Shuai Xiao, Su Liu, Weikai Zhou, Jialun Wu, Xinjie He, Zhiyuan Lin, Qiyang Xie

机构 * Independent Researchers(独立研究者)

专题命中 评测与基准 :LLM(title_cn,abstract);prompting(title,abstract);large language model(abstract);language model(abstract)

AI总结 通过对比四种提示条件在1140个开放式问题上的表现,发现角色提示系统性地增加专家深度但降低清晰度,其效果高度依赖于问题类型和领域,且混合检索优于纯嵌入检索。

Comments 6 pages, 2 figures. Submitted for peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25440 2026-05-26 cs.CL cs.AI cs.MA 91%

A Multi-Agent LLM Framework for Rating the Quality of Surgical Feedback

用于评估手术反馈质量的多智能体LLM框架

Rafal Kocielnik, J. Everett Knudsen, Steven Y. Cen, Jasmine Lin, Cherine H. Yang, Atharva Deo, Ujjwal Pasupulety, Peter Wager, Anima Anandkumar, Andrew J. Hung

机构 * Computing + Mathematical Sciences, California Institute of Technology(加州理工学院计算与数学科学系) Department of Urology, Cedars-Sinai(塞斯医疗中心泌尿科) Keck School of Medicine, University of Southern California(美国南加州大学凯克医学院)

专题命中 评测与基准 :LLM(title,title_cn);prompting(abstract);分类 cs.CL、cs.AI

AI总结 提出一个两阶段LLM框架,通过多智能体提示和手术领域知识注入发现可解释的反馈质量标准,并利用LLM作为评判者自动评分,在预测反馈有效性上优于先前方法。

Comments 25 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24755 2026-05-26 cs.AI cs.CL 91%

Automated Detection and Classification of Delusion-related Content in Naturalistic Audio Diaries Using Multi-Agent Language Models

使用多智能体语言模型自动检测和分类自然音频日记中的妄想相关内容

Feng Chen, Justin Tauscher, Changye Li, Meliha Yetisgen, Alex Cohen, Adam Kuczynski, Angelina Pei-Tzu Tsai, Benjamin Buck, Dror Ben-Zeev, Trevor Cohen

机构 * Department of Biomedical Informatics and Medical Education, University of Washington, Seattle, WA, USA(生物医学信息学与医学教育系,华盛顿大学,西雅图,华盛顿州,美国) Department of Psychiatry and Behavioral Sciences, University of Washington, Seattle, WA, USA(精神病学与行为科学系,华盛顿大学,西雅图,华盛顿州,美国) Department of Psychology, Louisiana State University, Baton Rouge, LA, USA(心理学系,路易斯安那州立大学,巴吞鲁日,路易斯安那州,美国) Department of Psychiatry, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA(精神病学系,北卡罗来纳大学教堂山分校,教堂山,北卡罗来纳州,美国)

专题命中 评测与基准 :LLM(summary_cn,abstract);language model(title,abstract);large language model(abstract);foundation model(abstract)

AI总结 提出一种多智能体LLM流水线,从自然音频日记中自动检测和分类妄想信念、情感和行为反应,通过多数投票实现稳健性能。

Comments Accepted by CLPych 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19999 2026-05-20 cs.LG cs.AI cs.CR 91%

LLM Benchmark Datasets Should Be Contamination-Resistant

LLM基准数据集应具备抗污染性

Ali Al-Lawati, Jason Lucas, Dongwon Lee, Suhang Wang

机构 * The Pennsylvania State University, University Park, PA, USA(宾夕法尼亚州立大学)

专题命中 评测与基准 :LLM(title,title_cn);pretraining(abstract);分类 cs.AI、cs.LG

AI总结 本文探讨了LLM基准数据集应具备抗污染性,提出通过改进数据集设计和架构来提高其可靠性和通用性。

Comments Accepted to ICML 2026 Position Paper Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08590 2026-05-12 cs.HC cs.AI cs.CL cs.CY 91%

Causal Stories from Sensor Traces: Auditing Epistemic Overreach in LLM-Generated Personal Sensing Explanations

从传感器轨迹中生成因果故事:审计LLM生成的个人传感解释中的知识过度延伸

Shanshan Zhu, Han Zhang, J. Doris Chi, Subigya Nepal, Koustuv Saha

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Chicago(芝加哥大学) Yale University(耶鲁大学) University of Virginia(弗吉尼亚大学)

专题命中 评测与基准 :LLM(title,title_cn);prompting(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了LLM生成个人传感解释时的知识过度延伸问题,通过分析不同数据集中的异常日场景,发现LLM在缺乏足够证据时仍会做出因果归因,且丰富的上下文无法有效减少这种过度延伸。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23701 2026-04-28 cs.CL cs.AI cs.CV 91%

Agri-CPJ: A Training-Free Explainable Framework for Agricultural Pest Diagnosis Using Caption-Prompt-Judge and LLM-as-a-Judge

Agri-CPJ:一种无需训练的可解释框架,用于利用Caption-Prompt-Judge和LLM-as-a-Judge进行农业害虫诊断

Wentao Zhang, Qi Zhang, Mingkun Xu, Mu You, Henghua Shen, Zhongzhi He, Keyan Jin, Derek F. Wong, Tao Fang

机构 * Business School, Shandong University of Technology(山东理工大学商学院) Faculty of Data Science, City University of Macau(澳门城市大学数据科学学院) Guangdong Institute of Intelligent Science and Technology(广东智能科学与技术研究院) Macau Millennium College(澳门 millennium 学院) Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系)

专题命中 评测与基准 :LLM(title,title_cn);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出Agri-CPJ框架,通过生成结构化形态描述并利用LLM进行判断,解决农业病害诊断中模型易产生错误物种名称和推理不可用的问题,实验显示其在病害分类和问答评分上有显著提升。

Comments This work is an expanded version of our prior paper published in the IEEE ICASSP 2026 conference arXiv:2512.24947, from 4 to 20+ pages, presenting a well-structured and principled framework, extensive experiments, and deeper insights. Tao Fang is the corresponding author

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12826 2026-04-28 cs.CL cs.AI cs.MA 91%

Scheming Ability in LLM-to-LLM Strategic Interactions

大语言模型在LLM到LLM战略互动中的策略能力

Thao Pham

机构 * Berea College(伯克利学院)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 研究通过博弈论框架评估大语言模型的策略欺骗能力,发现模型在无提示下倾向于欺骗,尤其在Peer Evaluation中全部选择欺骗,揭示了多智能体环境下需加强高风险博弈场景的评估。

Comments 20 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20273 2026-04-23 cs.AI cs.CL 91%

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks

ActuBench: 一个用于生成和评估精算推理任务的多智能体LLM流水线

Jan-Philipp Schmidt

机构 * TH Köln, Institut für Versicherungswesen (ivwKöln)(TH Köln保险学系(ivwKöln))

专题命中 评测与基准 :LLM(title,title_cn);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出ActuBench,一个多智能体LLM流水线,用于自动生成和评估符合国际精算协会教育大纲的高级精算评估题目。通过四个LLM角色的分工,结合成本优化的辅助代理,评估了50个模型并在两个基准上报告了三个主要发现。

Comments 19 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11199 2026-04-22 cs.CL cs.LG 91%

When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification

何时以及为何提问:AskBench与基于规则的强化学习与验证器(RLVR)用于LLM澄清

Jiale Zhao, Ke Fang, Lu Cheng

机构 * University of Illinois Chicago(伊利诺伊大学香槟分校)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文研究如何评估和改进LLM在缺乏关键信息或包含误导信息时的澄清能力,提出AskBench交互基准和基于规则的强化学习与验证器(RLVR)方法,提升准确性和交互效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03174 2026-04-06 cs.CL cs.AI 91%

Beyond the Parameters: A Technical Survey of Contextual Enrichment in Large Language Models: From In-Context Prompting to Causal Retrieval-Augmented Generation

超越参数:大型语言模型中上下文丰富技术的综述:从上下文提示到因果检索增强生成

Prakhar Bansal, Shivangi Agarwal

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(title);分类 cs.CL、cs.AI

AI总结 本文综述了大型语言模型中通过增加结构化上下文来提升性能的技术,涵盖上下文学习、提示工程、检索增强生成等方法,并提出透明的文献筛选协议和可信检索增强NLP的研究优先级。

Comments 7 pages, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13302 2025-07-18 cs.AI cs.CL 91%

The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations

Carlos Arriaga, Gonzalo Martínez, Eneko Sendin, Javier Conde, Pedro Reviriego

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17204 2024-12-02 cs.CL cs.AI 91%

Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks

Ratnesh Kumar Joshi, Priyanshu Priya, Vishesh Desai, Saurav Dudhate, Siddhant Senapati, Asif Ekbal, Roshni Ramnani, Anutosh Maitra, Shubhashis Sengupta

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(title);分类 cs.CL、cs.AI

Comments 39 pages, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13217 2023-04-03 cs.CL cs.AI 91%

Fairness-guided Few-shot Prompting for Large Language Models

Huan Ma, Changqing Zhang, Yatao Bian, Lemao Liu, Zhirui Zhang, Peilin Zhao, Shu Zhang, Huazhu Fu, Qinghua Hu, Bingzhe Wu

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04186 2022-10-12 cs.CL cs.AI 91%

Analogy Generation by Prompting Large Language Models: A Case Study of InstructGPT

Bhavya Bhavya, Jinjun Xiong, Chengxiang Zhai

专题命中 评测与基准 :language model(title,abstract);prompting(title,abstract);large language model(title);分类 cs.CL、cs.AI

Comments Accepted to 15th International Conference on Natural Language Generation (INLG 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23781 2026-07-15 cs.CR 版本更新 91%

Leveraging Large Language Models for Trustworthiness Assessment of Web Applications

利用大语言模型评估网络应用的可信度

Oleksandr Yarotskyi, José D'Abruzzo Pereira, João R. Campos

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);prompting(abstract)

AI总结 本文提出利用大语言模型自动评估网络应用的可信度,通过比较不同提示工程技术,提出基于逻辑偏好分数的分层质量模型,实验表明规则提示能提高评估可靠性。

Comments Accepted for publication in the 19th IEEE International Conference on Software Testing, Verification and Validation (ICST) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏