arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12565 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 12565 篇

2604.23600 2026-06-05 cs.CL 85%

Personality Shapes Gender Bias in Persona-Conditioned LLM Narratives Across English and Hindi: An Empirical Investigation

性格在英语和印地语中影响人物条件化大语言模型叙事中的性别偏见:一项实证研究

Tanay Kumar, Shreya Gautam, Aman Chadha, Vinija Jain, Francesco Pierri

机构 * Politecnico di Milano(米兰理工学院) Apple(苹果公司) Meta

专题命中 领域大模型 :LLM(title,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本研究探讨了在英语和印地语中,人物条件化大语言模型叙事中的性别偏见如何受到性格特征的影响,发现性格特质与性别偏见的幅度和方向显著相关,特别是黑暗三联体性格特质与性别刻板印象的表示更相关,但这些关联在不同模型和语言中有所变化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03693 2026-06-03 cs.CL cs.CV 85%

Does Language Shift Break Medical Vision-Language Models? Indonesian Radiology Visual Question Answering Case Study

语言转换会破坏医学视觉语言模型吗?印度尼西亚放射学视觉问答案例研究

Pieter Christy Yan Yudhistira, Dzaki Rafif Malik, Novanto Yudistira

机构 * Intelligent System Laboratory, Faculty of Computer Science Brawijaya University(智能系统实验室,计算机科学学院布拉维亚大学)

专题命中 领域大模型 :language model(title,abstract);foundation model(abstract);prompting(abstract);分类 cs.CL

AI总结 本研究通过构建印尼语放射学VQA数据集IndoRad-VQA,评估医学视觉语言模型在非英语临床语言下的鲁棒性,发现英语与印尼语设置间存在8-25%的性能差距,表明需要更包容的多语言评估。

Comments accepted to MMFM-BIOMED Workshop @ CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22672 2026-05-25 cs.AI 85%

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most

能力是负担吗?更强大的语言模型在关键时刻做出更差的预测

Nick Merrill, Jaeho Lee, Ezra Karger

机构 * Forecasting Research Institute(预测研究 institute)

专题命中 领域大模型 :language model(title);LLM(abstract,abstract_cn);post-training(abstract);分类 cs.AI

AI总结 本文发现,在时间序列呈现超线性增长和制度转换尾部风险的预测问题上,更强大的语言模型反而产生更差的分位数预测,并分析了这一逆缩放现象的机制、影响因素及评估指标问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14215 2026-05-19 cs.IR cs.AI 85%

PriHA: A RAG-Enhanced LLM Framework for Primary Healthcare Assistant in Hong Kong

PriHA:一种增强型大语言模型框架,用于香港初级医疗服务助手

Richard Wai Cheung Chan, Shanru Lin, Ya-nan Ma, Hao Chen, Liangjun Jiang, Wenqi Fan

机构 * The Hong Kong Polytechnic University(香港理工大学) Haikou Affiliated Hospital of Central South University Xiangya School of Medicine(中南大学湘雅医学院海口附属医院) Sun Yat-Sen University Cancer Center(中山大学肿瘤中心)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出PriHA框架,通过检索增强生成技术解决香港初级医疗服务中指南碎片化问题,提升信息访问准确性与清晰度。

Comments Accepted to PAKDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14202 2026-05-15 cs.SE cs.AI 85%

LLM-Based Robustness Testing of Microservice Applications: An Empirical Study

基于大型语言模型的微服务应用鲁棒性测试:实证研究

Hrushitha Goud Tigulla, Marco Vieira

机构 * College of Computing(计算学院) Informatics University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校信息学院) Charlotte, USA(美国夏洛特)

专题命中 领域大模型 :LLM(title,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文通过实证研究探讨大型语言模型在生成微服务应用鲁棒性测试用例中的有效性,发现提示策略对测试用例多样性影响更大,引入两种策略提升测试覆盖率并保持低跨模型相似性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12265 2026-05-13 cs.AI 85%

How Useful Is Cross-Domain Generalization for Training LLM Monitors?

在训练LLM监控器时,跨领域泛化有多有用?

Sam Martin, Fabien Roger

机构 * Anthropic Fellows Program(Anthropic 后备计划)

专题命中 领域大模型 :LLM(title,title_cn);language model(abstract);分类 cs.AI

AI总结 研究通过多任务训练提升跨领域分类性能,发现部分泛化效果,但存在特定场景下的失败,混合训练能保持优势并缓解泛化问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15141 2026-04-23 cs.IR cs.AI 85%

ItemRAG: Item-Based Retrieval-Augmented Generation for LLM-Based Recommendation

ItemRAG: 基于物品的检索增强生成用于基于大语言模型的推荐系统

Sunwoo Kim, Geon Lee, Kyungho Kim, Jaemin Yoo, Kijung Shin

机构 * Seoul National University(首尔国立大学)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 ItemRAG通过细粒度物品级检索提升推荐效果,结合共购信息与语义信息,有效解决冷启动问题,实验表明其优于现有RAG方法。

Comments Published as a conference paper at SIGIR 2026 (short)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09952 2026-04-14 cs.LG 85%

SLM Finetuning for Natural Language to Domain Specific Code Generation in Production

为生产环境中的自然语言到领域特定代码生成进行SLM微调

Renjini R. Nair, Damian K. Kowalczyk, Marco Gaudesi, Chhaya Methani

专题命中 领域大模型 :SLM(title);large language model(abstract);language model(abstract);small language model(abstract)

AI总结 本文研究了通过微调小型语言模型提升自然语言到领域代码生成的性能与效率,展示了在生产环境中优于大模型的延迟和成本效益。

Comments 11 pages (including appendix), 5 tables, 1 figure. Submitted to arXiv as a preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09584 2026-04-14 cs.AI cs.CV 85%

Agentic Exploration of PDE Spaces using Latent Foundation Models for Parameterized Simulations

基于潜在基础模型的PDE空间代理探索用于参数化模拟

Abhijeet Vishwasrao, Francisco Giral, Mahmoud Golestanian, Federica Tonti, Andrea Arroyo Ramo, Adrian Lozano-Duran, Steven L. Brunton, Sergio Hoyas, Soledad Le Clainche, Hector Gomez, Ricardo Vinuesa

机构 * University of Michigan(密歇根大学) Universidad Politécnica de Madrid(马德里理工大学) Purdue University(普渡大学) Universitat Politècnica de València(瓦伦西亚理工大学) Caltech(加州理工学院) University of Washington(华盛顿大学)

专题命中 领域大模型 :foundation model(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出利用多代理LLM与潜在基础模型耦合,实现对PDE参数空间的连续探索,发现流体中不同尺度的标度律,展示了代理推理在PDE系统中的自动化科学发现能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09551 2026-04-14 cs.IR cs.AI 85%

SemaCDR: LLM-Powered Transferable Semantics for Cross-Domain Sequential Recommendation

SemaCDR: 基于大语言模型的跨域序列推荐语义

Chunxu Zhang, Shanqiang Huang, Zijian Zhang, Jiahong Liu, Linsong Yu, Ruiqi Wan, Bo Yang, Irwin King

机构 * Jilin University(吉林大学) Hong Kong Polytechnic University(香港理工大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 SemaCDR通过大语言模型构建统一语义空间,整合领域无关和领域特定语义,利用对比正则化对齐多视角物品特征,通过自适应融合生成统一偏好表示,提升跨域序列推荐效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08691 2026-04-13 cs.SI cs.AI 85%

AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society

AgentSociety: 基于大语言模型的生成代理大规模模拟推动对人类行为与社会的理解

Jinghua Piao, Yuwei Yan, Jun Zhang, Nian Li, Junbo Yan, Xiaochong Lan, Zhihong Lu, Zhiheng Zheng, Jing Yi Wang, Di Zhou, Chen Gao, Fengli Xu, Fang Zhang, Ke Rong, Jun Su, Yong Li

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出AgentSociety平台,通过大规模模拟人类行为和社会动态,探索社会问题如极化、虚假信息传播等,验证其在社会科学研究中的应用价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08956 2026-04-13 cs.CV cs.LG 85%

Low-Data Supervised Adaptation Outperforms Prompting for Cloud Segmentation Under Domain Shift

低数据监督适应在云分割领域优于提示

Harshith Kethavath, Weiming Hu

机构 * University of Georgia(佐治亚大学)

专题命中 领域大模型 :prompting(title,abstract);language model(abstract);pretraining(abstract);分类 cs.LG

AI总结 本文研究了在遥感影像云分割任务中,低数据监督微调优于提示方法,发现使用少量标注数据可显著提升性能,而提示方法效果有限。

Comments 10 pages, 6 figures, to be published in EarthVision @ CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08126 2026-04-10 cs.CL 85%

LLM-Based Data Generation and Clinical Skills Evaluation for Low-Resource French OSCEs

基于大语言模型的数据生成与低资源法语OSCEs临床技能评估

Tian Huang, Tom Bourgeade, Irina Illina

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出利用大语言模型生成和评估法语OSCE对话,通过合成数据缓解资源限制,展示中等规模模型在医疗教育中的可行性。

Comments 11 pages, 2 figures, to be published in LREC 2026 proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07584 2026-04-10 cs.AI 85%

From Papers to Property Tables: A Priority-Based LLM Workflow for Materials Data Extraction

从论文到属性表:一种基于优先级的LLM工作流用于材料数据提取

Koushik Rameshbabu, Jing Luo, Ali Shargh, Khalid A. El-Awady, Jaafar A. El-Awady

机构 * Department of Applied Mathematics and Statistics, Johns Hopkins University(约翰霍普金斯大学应用数学与统计系) Department of Mechanical Engineering, Johns Hopkins University(约翰霍普金斯大学机械工程系)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出一种基于优先级的LLM工作流,用于从科研论文中自动提取和重建结构化材料实验数据,通过整合文本、表格、图表和物理推导信息,以合金劈裂强度为案例,实现了高精度的数据提取与验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06487 2026-04-09 cs.CL 85%

Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR

通过有限音频缩小语音-文本差距以实现基于LLM的ASR的有效领域适应

Thibault Bañeras-Roux, Sergio Burdisso, Esaú Villatoro-Tello, Dairazalia Sánchez-Cortés, Shiran Liu, Severin Baroudi, Shashi Kumar, Hasindri Watawana, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke

机构 * Idiap Research Institute(Idiap 研究所) Laboratoire d'Informatique et des Systèmes(信息系统实验室) Uniphore Brno University of Technology(布尔诺理工大学)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文研究了有限音频在基于LLM的ASR领域适应中的作用,比较了纯文本适应、配对语音-文本适应和混合批处理策略,发现少量语音能有效提升性能。

Comments Submitted to Interspeech

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04174 2026-04-07 cs.AI 85%

CoALFake: Collaborative Active Learning with Human-LLM Co-Annotation for Cross-Domain Fake News Detection

CoALFake:协作主动学习与人类-大语言模型共注释用于跨领域假新闻检测

Esma Aïmeur, Gilles Brassard, Dorsaf Sallami

机构 * University of Montreal(蒙特利尔大学)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出CoALFake,通过人类与大语言模型共注释和领域感知主动学习,解决跨领域假新闻检测中数据获取困难和领域特征丢失的问题,实验显示其在多个数据集上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03898 2026-04-07 cs.AI stat.CO 85%

LLM-Agent-based Social Simulation for Attitude Diffusion

基于大语言模型的社交模拟用于态度扩散

Deepak John Reji

机构 * University of Limerick(利默里克大学)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出discourse_sim框架,结合LLM与基于代理的建模,模拟公众对移民态度随时间变化的动态,通过生成社交媒体帖子和解读观点来研究态度扩散和信念演变。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02660 2026-04-06 cs.CL 85%

SocioEval: A Template-Based Framework for Evaluating Socioeconomic Status Bias in Foundation Models

SocioEval: 一种基于模板的评估框架,用于评估基础模型中的社会经济地位偏见

Divyanshu Kumar, Ishita Gupta, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi

专题命中 领域大模型 :foundation model(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出SocioEval框架,通过决策任务系统评估基础模型中的社会经济偏见,发现不同主题的偏见程度差异显著,部署防护措施能有效防止显性歧视但易受领域特定刻板印象影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01461 2026-04-03 cs.AI 85%

Reducing Hallucinations in LLM-based Scientific Literature Analysis Using Peer Context Outlier Detection

通过同侪上下文异常检测减少基于LLM的科学文献分析中的幻觉

Daniel Xie, Maxwell J. Jacobson, Adil Wazeer, Haiyan Wang, Xinghang Zhang, Yexiang Xue

专题命中 领域大模型 :LLM(title);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出Peer Context Outlier Detection方法,通过文档间关系提升数据提取准确性,实验显示在6个科学领域达到98%的异常检测精度,减少幻觉并提高自动化系统可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04921 2026-04-02 cs.CL 85%

AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent

AgentExpt: 基于LLM的资源检索代理自动化AI实验设计

Yu Li, Lehui Li, Lin Chen, Qingmin Liao, Fengli Xu, Yong Li

机构 * Tsinghua University(清华大学) Shandong University(山东大学) Northeastern University(东北大学)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出基于LLM的资源检索代理框架,通过自动化数据收集和集体感知增强检索,提升实验设计的自动化和可解释性,实验结果显示在Recall@20和HitRate@5上分别提升5.85%和8.30%。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00281 2026-04-02 cs.AI 85%

Human-in-the-Loop Control of Objective Drift in LLM-Assisted Computer Science Education

人机协同控制LLM辅助计算机科学教育中的目标漂移

Mark Dranias, Adam Whitley

机构 * Asheville Institute for Memory and Longevity(阿什维尔记忆与长寿研究所) University of North Carolina Asheville(北卡罗来纳大学阿什维尔分校)

专题命中 领域大模型 :LLM(title);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出一种人机协同的教学方法,通过分离规划与执行,训练学生在代码生成前指定接受标准和架构约束,以稳定AI辅助学习过程。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11349 2026-03-31 cs.LG nlin.CD physics.comp-ph 85%

Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learning

上下文鹦鹉:一种简单但难以击败的基础模型基准,用于科学机器学习中的基础模型

Yuanzhao Zhang, William Gilpin

机构 * Santa Fe Institute(圣塔菲研究所) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 领域大模型 :foundation model(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文指出基础模型常通过简单鹦鹉策略预测,而非鹦鹉时易出现收敛到均值等失败模式。简单鹦鹉模型在低维混沌、湍流等动态系统预测中表现优于主流模型,且计算成本低。

Comments International Conference on Learning Representations (ICLR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13082 2026-03-30 cs.CR cs.LG 85%

Adversarial News and Lost Profits: Manipulating Headlines in LLM-Driven Algorithmic Trading

对抗性新闻与损失利润:在基于大语言模型的算法交易中操纵头条

Advije Rizvani, Giovanni Apruzzese, Pavel Laskov

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 研究探讨了在基于大语言模型的算法交易中,通过操纵新闻头条导致模型误判的风险,分析了Unicode同音替代和隐藏文本条款对模型影响,并量化了其对收益的负面影响。

Comments This work has been accepted for publication at the IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). The final version will be available on IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25422 2026-03-27 cs.CL cs.CY 85%

Navigating the Prompt Space: Improving LLM Classification of Social Science Texts Through Prompt Engineering

在提示空间中导航:通过提示工程提高LLM对社会科学文本的分类

Erkan Gunes, Christoffer Florczak, Tevfik Murat Yildirim

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文通过系统变化提示工程的三个方面,探讨如何通过增加提示上下文提高LLM对社会科学文本的分类准确性,发现最小的上下文增加能显著提升性能,但过度增加反而可能降低准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12138 2026-03-25 cs.AI 85%

DriveSafe: A Hierarchical Risk Taxonomy for Safety-Critical LLM-Based Driving Assistants

DriveSafe:面向安全关键的LLM驾驶助手风险分类

Abhishek Kumar, Riya Tapwal, Carsten Maple

机构 * The Alan Turing Institute(艾伦·图灵研究所) University of Warwick(沃里克大学)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出DriveSafe,一种四层风险分类体系,用于系统表征LLM驾驶助手的安全关键失效模式,涵盖技术、法律、社会和伦理维度,通过专家评审和模型评估验证其现实性与安全性。

Comments This is the revised version of the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21530 2026-03-24 cs.SE cs.AI 85%

LLM-Based Test Case Generation in DBMS through Monte Carlo Tree Search

通过蒙特卡洛树搜索的基于大语言模型的数据库管理系统测试用例生成

Yujia Chen, Yingli Zhou, Fangyuan Zhang, Cuiyun Gao

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Huawei Hong Kong Research Center(华为香港研究中心)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出MIST框架,通过蒙特卡洛树搜索生成不同数据库管理系统方言的语法正确且语义多样的测试用例,提升代码覆盖率。

Comments Accepted to ICSE 2026 Industry Challenge Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14783 2026-03-24 cs.CL cs.CY 85%

Human or LLM as Standardized Patients? A Comparative Study for Medical Education

人类或大语言模型作为标准化患者?医学教育中的比较研究

Bingquan Zhang, Xiaoxiao Liu, Yuchi Wang, Lei Zhou, Qianqian Xie, Benyou Wang

机构 * School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Freedom AI

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出EasyMED框架和SPBench基准,通过对比实验显示其在医学教育中更接近人类标准化患者行为,尤其在案例一致性与可控披露方面表现更优,且在学习效果和成本效率上具有优势。

Comments 24 pages, 13 figures, 10 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12632 2026-03-24 cs.CL 85%

Prompt-Induced Linguistic Fingerprints for LLM-Generated Fake News Detection

基于提示的语言指纹用于LLM生成虚假新闻检测

Chi Wang, Min Gao, Zongwei Wang, Junwei Yin, Kai Shu, Chenghua Lin

机构 * Chongqing University(重庆大学) Emory University(埃默里大学) University of Manchester(曼彻斯特大学) Key Laboratory of Dependable Service Computing in Cyber Physical Society (Chongqing University), Ministry of Education of China(可信服务计算网络物理社会关键实验室(重庆大学),中国教育部)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出LIFE方法,通过重建词级概率分布发现语言指纹,提升LLM生成虚假新闻检测性能,实验显示其在LLM和人工生成虚假新闻中均表现优异。

Comments published in WWW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01503 2026-03-24 cs.CL 85%

A Theory of Adaptive Scaffolding for LLM-Based Pedagogical Agents

基于大语言模型的教育代理自适应支架理论

Clayton Cohn, Surya Rayala, Namrata Srivastava, Joyce Horn Fonteles, Shruti Jain, Xinying Luo, Divya Mereddy, Naveeduddin Mohammed, Gautam Biswas

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出结合证据中心设计与社会认知理论的自适应支架框架,开发出Inquizzitor评估代理,通过人机混合智能提供基于认知科学的反馈,验证了理论驱动的大语言模型在教育中的应用潜力。

Comments Published in the proceedings of AAAI 2026 (main technical track)

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(3), 1757-1765. 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17017 2026-03-24 cs.CL 85%

SafeSearch: Do Not Trade Safety for Utility in LLM Search Agents

SafeSearch: 不应为效用而牺牲安全性的LLM搜索代理

Qiusi Zhan, Angeline Budiman-Chan, Abdelrahman Zayed, Xingzhi Guo, Daniel Kang, Joo-Kyung Kim

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文研究了基于大语言模型的搜索代理在安全与效用之间的平衡问题,提出SafeSearch方法通过多目标强化学习提升安全性和效用,实验表明其显著降低有害输出并保持问答性能。

Comments EACL 2026 Findings. Code available at https://github.com/amazon-science/SafeSearch

详情

展开后加载摘要…

URL PDF HTML 收藏