arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-06-30 至 2026-06-30 共收录 143 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 143 篇

2606.29876 2026-06-30 cs.CL cs.AI q-bio.QM 92%

Clinical Reasoning Graphs: Structured Evaluation of LLM Diagnostic Reasoning Reveals Competence Without Consistency

临床推理图:LLM诊断推理的结构化评估揭示能力与一致性无关

Nisarg A. Patel

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 通过从LLM诊断痕迹中提取结构化图(5种节点和7种边),分析50例临床病例,发现图相似性在相似病例与不相似病例间无显著差异,表明LLM具备诊断能力但缺乏推理一致性。

Comments Spotlight Paper, Proceedings of the Workshop on Structured Data for Health at the 43rd International Conference on Machine Learning (ICML), Seoul, South Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28963 2026-06-30 cs.CL cs.CY cs.LG 92%

Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data

超越均值:从小型试点数据对齐基于LLM的调查模拟器的三轴保真度

Eun Cheol Choi, Youngrae Kim, Prabhu Pugalenthi, Hong-En Chen, Bo-Ruei Huang

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 针对LLM模拟社会调查响应时的系统偏差,提出结构、边际和个体三轴保真度框架,通过提示、修正和微调三种方法对比,发现微调小型试点数据可实现平衡保真度,但保真度水平因子样本而异。

Comments 11 pages, 8 tables, 3 figures; Pluralistic Alignment @ ICML 2026 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10109 2026-06-30 cs.AI cs.HC cs.LG 92%

LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals

基于自我报告的LLM代理能够实现通用个体模拟

Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, Michael S. Bernstein

机构 * Computer Science Department, Stanford University(斯坦福大学计算机科学系) Department of Communication Studies, Northwestern University(西北大学传播学系) Department of Communication, University of Washington(华盛顿大学传播学系) Google DeepMind(谷歌DeepMind) Department of Sociology, Stanford University(斯坦福大学社会学系) Sciences Po(巴黎政治学院)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文研究了基于自我报告数据的LLM代理在模拟个体行为方面的有效性,通过不同数据源构建代理并验证其在多种任务中的准确性和跨群体公平性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30454 2026-06-30 physics.soc-ph cs.AI 92%

Collective cooperation without individual fidelity in LLM agents

LLM智能体中的集体合作无需个体忠诚

Henrique Ferraz de Arruda, Carlos Gracia Lázaro, Alberto Aleta, Yamir Moreno

机构 * ARAID Foundation(ARAID基金会) Institute for Biocomputation and Physics of Complex Systems(复杂系统生物计算研究所) University of Zaragoza(萨拉戈塔大学) Universidad San Jorge (USJ)(圣乔治大学) Department of Theoretical Physics(理论物理系)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究通过大规模网络囚徒困境实验,比较九种开源LLM与人类行为,发现LLM能复现宏观合作动态,但微观层面存在个体异质性和条件合作模式的差异,揭示了宏观-微观分离现象。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28450 2026-06-30 cs.CR cs.AI 92%

LLM agents security duality: a comprehensive survey of self-security and empowered cybersecurity

LLM代理安全二元性:自我安全与赋能网络安全的综合调查

Yiwei Xu, Yong Zhuang, Xuanming Liu, Tian Zhang, Bowen Xiao, Xiaoyang Xu, Delong Jiang, Juan Wang, Hongxin Hu

机构 * School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院) Department of Computer Science and Engineering, University at Buffalo(布法罗大学计算机科学与工程系)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 综述LLM代理的自我安全威胁与防御,以及其在网络攻防中的赋能作用,首次揭示两者间的正反馈协同效应。

Comments 73 pages,12 figures, 9 tables, Artificial Intelligence Review

Journal ref Artif Intell Rev (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30163 2026-06-30 eess.SY cs.SY 91%

End-to-End Abstraction-Based Control with LLM-Enhanced NL-to-LTL Translation

基于端到端抽象的控制与LLM增强的NL到LTL翻译

Amir Bayat, Necmiye Ozay, Alessandro Abate, Raphael M. Jungers

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 提出LLM增强的NL-to-LTL翻译管道,集成到ABCD框架中,并通过实验揭示翻译成功率与LTL公式复杂度的缩放规律。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29520 2026-06-30 cs.SE cs.AI cs.DB 91%

SAKE: Software Architectural Knowledge Evaluation Benchmark for Large Language Models

SAKE:大型语言模型的软件架构知识评估基准

Tiziano Santilli, Francesco Daghero, Mayhar Tourchi Moghaddam

机构 * University of Southern Denmark(丹麦南方大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(summary_cn,abstract_cn);分类 cs.AI

AI总结 提出SAKE基准,包含2154道专家策划的多选题,评估LLM在八个架构类别和四种上下文长度下的架构知识,揭示专业实践中的能力差距。

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28345 2026-06-30 cs.RO cs.AI cs.CL cs.CY 91%

Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients

审计具有文化特定道德梯度的LLM驱动社交机器人

Carmen Ng, Gjergji Kasneci

专题命中 评测与基准 :LLM(title,title_cn);prompting(abstract);分类 cs.CL、cs.AI

AI总结 针对LLM驱动社交机器人在跨文化场景中优先分配资源时的道德偏差,提出基于梯度的多语言审计框架,通过对称控制场景测试模型对文化偏好梯度的区分能力,发现提示工程无法可靠纠正的持续性不对称问题。

Comments Accepted for publication in Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19633 2026-06-30 cs.AI 90%

OptiMUS-0.3: Using Large Language Models to Model and Solve Optimization Problems at Scale

OptiMUS-0.3:利用大语言模型在大规模范围内建模和解决优化问题

Ali AhmadiTeshnizi, Wenzhi Gao, Herman Brunborg, Shayan Talaei, Connor Lawless, Madeleine Udell

机构 * School of Management Science and Engineering, Stanford University(斯坦福大学管理科学与工程学院) Institute for Computational and Mathematical Engineering, Stanford University(斯坦福大学计算与数学工程研究所)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

AI总结 本文提出OptiMUS-0.3系统,通过大语言模型自动建模和解决线性规划问题,提升优化工具的实用性与效率,实验证明其在易实例和实际案例中表现优异。

Comments This paper documents OptiMUS-0.3, improving on OptiMUS-0.1 (arXiv:2310.06116) and OptiMUS-0.2 (arXiv:2402.10172). arXiv admin note: text overlap with arXiv:2402.10172

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29273 2026-06-30 cs.CL cs.AI 90%

A Hybrid Framework for Song Lyric Annotation Based on Human-LLM Alignment

基于人类-大语言模型对齐的歌词注释混合框架

Rashini Liyanarachchi, Frank Tran, Md Mahmudul Hasan, Aditya Joshi, Erik Meijering

机构 * School of Computer Science and Engineering, University of New South Wales(新南威尔士大学计算机科学与工程学院)

专题命中 评测与基准 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出混合注释框架,通过预测人类与LLM在歌词情感标注中的潜在不一致,优化注释过程,并创建句子级歌词数据集。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28796 2026-06-30 cs.CL cs.LG 90%

Structure-Preserving Document Translation via Multi-Stage LLM Pipeline: A Case Study in Marathi

通过多阶段LLM流水线实现结构保持的文档翻译:以马拉地语为例

Manasi Waghe, Danish Chandargi, Mohammad Aamir Rayyan, Raviraj Joshi, A. R. Deshpande

机构 * Pune Institute of Computer Technology(普那计算机技术学院) L3Cube Labs(L3Cube实验室) Indian Institute of Technology Madras(印度理工学院马德拉斯分校)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 提出一种结构保持的马拉地语-英语政府文档翻译框架,集成布局感知OCR、坐标文本提取、LLM翻译和HTML重建,确保文档布局和结构一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09708 2026-06-30 cs.LG cs.AI cs.DC 90%

Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon

Metal-Sci: 一个用于苹果硅芯片上进化式LLM内核搜索的科学计算基准

Víctor Gallego

机构 * Komorebi AI Technologies(Komorebi人工智能技术)

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.AI、cs.LG

AI总结 Metal-Sci基准通过10个科学任务和六个优化领域,评估LLM在自动内核搜索中的性能,展示了不同模型在不同任务上的加速效果及结构化评估方法的可靠性。

Comments Published at the Fifth Workshop on Deep Learning for Code (DL4C) at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29602 2026-06-30 cs.CR 90%

An Empirical Evaluation of Prompt Injection Vulnerabilities in Large Language Models Across Multilingual and Obfuscated Attack Scenarios

多语言与混淆攻击场景下大语言模型提示注入漏洞的实证评估

Caglar Uysal, Baturay Birinci, Süha Orhun Mutluergil, Orçun Çetin

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn)

AI总结 本文实证评估六种大语言模型在多语言和混淆攻击下的提示注入漏洞,发现所有模型均易受攻击,非英语语言恶意合规率更高,需加强安全防御。

Comments Accepted to the AI-SS 2026 Workshop at the 21st European Dependable Computing Conference (EDCC 2026). To be published in the EDCC Companion Proceedings (EDCC-C)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28615 2026-06-30 cs.LG cs.AI cs.CL stat.ML 90%

What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs

LLM 解释的内容并非其信念:基于模型自身输入信念评估解释充分性

Nhi Nguyen, Shauli Ravfogel, Rajesh Ranganath

机构 * New York University(纽约大学)

专题命中 评测与基准 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出自洽充分性(SCSuff)指标,利用 LLM 自身生成替代输入来评估自由文本解释的充分性,发现 LLM 解释通常不充分且与模型大小、准确率等弱相关。

Comments 23 pages, 9 figures, 13 tables, Forty-Third International Conference on Machine Learning (ICML 2026)

Journal ref Forty-Third International Conference on Machine Learning (ICML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30383 2026-06-30 cs.AI 90%

Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents

你的智能体站在哪一边?LLM智能体中的多方委托人忠诚度

Bojie Li, Noah Shi

机构 * Pine AI University of Washington(华盛顿大学皮尼AI学院)

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.AI

AI总结 研究LLM智能体在多方对话中对委托人的忠诚度问题,提出基准测试PrincipalBench和两种机制(提示忠诚度框架和知识蒸馏),揭示泄漏与过度拒绝之间的权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29920 2026-06-30 cs.CL 90%

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

LLM-as-a-Judge 能否可靠地验证智能体场景中的评分标准?

Yangda Peng, Yunjia Qi, Hao Peng, Haotian Xia, Guanzhong He, Xintong Shi, Richeng Xuan, Songyuanyi Lu, Yixian Liu, Zhichao Hu, Yuhong Liu, Lei Hou, Bin Xu, Juanzi Li

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.CL

AI总结 针对 LLM 作为裁判在智能体场景中验证评分标准的可靠性问题,构建首个基准 RuVerBench,评估多种前沿模型,发现即使最强模型也存在显著噪声,并分析了提示设计、批处理和多数投票等策略的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29900 2026-06-30 cs.CV cs.AI 90%

LLM-based Multimodal Personality Recognition via Facial Action Unit-Text Semantic Fusion

基于面部动作单元-文本语义融合的LLM多模态人格识别

Tianyi Zhang, Wei Shan, Yuan Zong, Tianhua Qi, Wenming Zheng

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出一种基于大语言模型的面部动作单元与文本语义融合框架,通过将AU序列转化为可解释文本描述并与受访者文本回答融合,实现异步视频面试中的人格识别,在AVI-6基准上取得更优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14940 2026-06-30 cs.CL cs.AI 90%

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models

CASE-Bench:面向大语言模型的上下文感知安全基准

Guangzhi Sun, Xiao Zhan, Shutong Feng, Philip C. Woodland, Jose Such

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI

AI总结 本文提出CASE-Bench,通过整合上下文信息评估大语言模型的安全性,揭示上下文对人类判断的显著影响,并发现商业模型在安全情境下的响应与人类判断存在明显偏差。

Comments 24 pages. This paper has been accepted at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30109 2026-06-30 cs.RO 89%

TacEvo: Self-Evolving Architecture Discovery for Robotic Tactile Perception via LLM-Driven Quality-Diversity Search

TacEvo:通过LLM驱动的质量多样性搜索实现机器人触觉感知的自演化架构发现

Mohammed AbuSadeh, Lan Wei, Dandan Zhang

机构 * Department of Bioengineering, Imperial-X Initiative, Imperial College London(生物工程系、Imperial-X计划、伦敦帝国学院)

专题命中 评测与基准 :LLM(title,title_cn)

AI总结 提出TacEvo框架,利用LLM生成代码级变异和交叉,结合MAP-Elites质量多样性循环,自动发现高效触觉感知网络架构,在力回归和光栅分类任务上显著提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28340 2026-06-30 cs.IR 89%

Rethinking Fairness in LLM-Based Recommender Systems: A Survey

重新思考基于LLM的推荐系统中的公平性:一项综述

Song-Duo Ma, Chu-Yun Chen, Bang-An Li, Pin-Yu Chen, Shau-Yung Hsu, Yun-Nung Chen

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 本文系统综述了基于大语言模型的推荐系统中的公平性问题,从偏见机制和公平目标两个维度组织现有研究,并探讨了评估与缓解策略,旨在为未来研究提供结构化基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29366 2026-06-30 math.OC cs.AI 89%

Solver-Verified Formulation Generation and Selection for Multi-Warehouse Inventory Allocation Using Large Language Models

基于求解器验证的多仓库库存分配公式生成与选择——利用大语言模型

Jintao Xu, Yingzheng Ma, Jiong Dong, Yongzhi Qi, Jianshen Zhang, Dongyang Geng, Anni Zhang

机构 * Supply Chain Tech Team Y, JD.com(京东供应链技术团队Y)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 提出ORLA框架,利用大语言模型从自然语言需求自动生成运筹学公式,通过求解器反馈验证并选择最优公式,在京东29个生产批次上实现分配准确率提升4.5个百分点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28683 2026-06-30 cs.AI cs.CY stat.AP stat.ME 89%

Aristotelian Virtue Profiling of LLMs through Ethical Dilemmas

通过伦理困境对LLM进行亚里士多德式美德画像

Ioannis Tzachristas, John Pavlopoulos

机构 * Technical University of Munich(慕尼黑技术大学) Athens University of Economics and Business(雅典经济与商业大学) Archimedes, Athena Research Center(阿提卡研究中心)

专题命中 评测与基准 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出VirtueMap框架,通过七种非致命、非政治、非宗教的伦理困境,让人类或LLM对五种回应排序,基于95%共识的参考排序,使用归一化Borda对齐计算实践智慧、正义、诚实、勇气和节制的美德画像,在九个LLM家族中评估,平均排名一致性达90.3%。

Comments VirtueMap website: https://jtzach.github.io/Aristotle-Virtue-Map

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00835 2026-06-30 cs.CL 89%

Agentic Tool Use in Large Language Models

大语言模型中的代理工具使用

Jinchao Hu, Meizhi Zhong, Kehai Chen, Xuefeng Bai, Min Zhang

机构 * School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)计算机科学与技术学院)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL

AI总结 本文系统梳理了大语言模型中代理工具使用的三种范式,分析了其方法、优缺点及评估现状,旨在解决现有研究碎片化问题,提供更系统的进化视角。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21787 2026-06-30 cs.SE cs.AI 89%

Assessing the Business Process Modeling Competences of Large Language Models

评估大型语言模型的业务流程建模能力

Chantale Lauer, Peter Pfeiffer, Alexander Rombach, Nijat Mehdiyev

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出BEF4LLM框架,从语法、语用、语义和有效性四个维度评估LLM生成的BPMN模型,发现LLM在语法和语用质量上表现优异,但语义和有效性仍需提升,为未来模型优化提供指导。

Journal ref Information Systems, Vol. 142 (2026), Art. 102761

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10858 2026-06-30 cs.CR cs.AI 89%

Large Language Models for Security Operations Centers: A Comprehensive Survey

面向安全运营中心的大型语言模型:全面综述

Ali Habibzadeh, Farid Feyzi, Reza Ebrahimi Atani

机构 * University of Guilan(吉兰大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文综述了将生成式AI,特别是大型语言模型应用于安全运营中心的工作流程,探讨其能力、挑战及未来方向,为研究人员和运营管理者提供当前状态的全面视角。

Journal ref Journal of Electrical and Computer Engineering, 2026, 3383674 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30062 2026-06-30 cs.CL cs.AI 89%

Little Brains, Big Feats: Exploring Compact Language Models

小脑袋,大成就:探索紧凑型语言模型

Dari Baturova, Elena Bruches, Ivan Chernov, Roman Derunets, Arsenii Fomin, Andrey Kostin

机构 * Siberian Neuronets LLC(西伯利亚神经网络有限公司)

专题命中 评测与基准 :language model(title,abstract);SLM(abstract,abstract_cn);large language model(abstract);small language model(abstract)

AI总结 本研究探索小型语言模型在检索增强生成(RAG)系统中的生成性能,实验表明无需GPU即可在设备上运行,并提供了基准测试结果。

Comments Accepted to ECML PKDD 2026, Applied Data Science track. Author preprint; the definitive version will appear in the proceedings of ECML PKDD 2026, Springer LNCS

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28618 2026-06-30 cs.SE 89%

Evaluating LLMs on Java Code Snippet Adaptation Using a Mutation-Injection Framework

使用突变注入框架评估LLM在Java代码片段适配上的表现

Ali Aman, Muhammad Asaduzzaman, Shaowei Wang, Chanchal K. Roy

专题命中 评测与基准 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract)

AI总结 通过突变注入框架构建Java代码片段数据集,研究无指令条件下LLM对代码片段的适配能力,分析适配类型难度、复杂度影响及所需上下文范围。

Comments Accepted in the 42nd IEEE International Conference on Software Maintenance and Evolution (ICSME 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05366 2026-06-30 cs.CL cs.AI cs.LG 89%

Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models

在执行中迷失:大型语言模型在多语言环境下的工具调用鲁棒性研究

Zheng Luo, T Pranav Kutralingam, Ogochukwu N Okoani, Wanpeng Xu, Hua Wei, Xiyang Hu

机构 * University of Southern California(南加州大学) Arizona State University(亚利桑那州立大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究探讨了多语言环境下大型语言模型工具调用的鲁棒性,通过MLCL基准测试发现参数值语言不匹配是主要失败模式,尽管策略减少错误但无法恢复英语水平性能。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03546 2026-06-30 cs.CL cs.AI cs.HC cs.LG 89%

Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict

大语言模型在隐私与亲社会冲突下的价值-行动对齐

Guanyu Chen, Chenxiao Yu, Xiyang Hu

机构 * Arizona State University(亚利桑那州立大学) University of Southern California(南加州大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究探讨大语言模型在隐私与亲社会冲突中的价值-行动对齐,通过多组结构方程模型分析隐私关注与亲社会性对数据共享的影响,提出价值-行动对齐率(VAAR)作为评估指标。

Comments Findings of the Association for Computational Linguistics: ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30085 2026-06-30 cs.CL econ.GN q-fin.EC 89%

Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates

非完全人类品味:LLM调查替代者的程式化杂食性

Xiangyu Ma, Mengmi Zhang, Shannon Ang, Minne Chen

机构 * Nanyang Technological University Singapore(新加坡南洋理工大学)

专题命中 评测与基准 :LLM(title,title_cn);language model(abstract);分类 cs.CL

AI总结 本研究使用多个大型语言模型生成大量调查替代者,发现其品味存在系统性正向偏差、丧失真实品味结构的复杂关系,且扭曲了品味与社会空间的关联。

详情

展开后加载摘要…

URL PDF HTML 收藏