arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Annual Meeting of the Association for Computational Linguistics · 会议 · Natural Language Processing

共收录 10293
2608.12361 2026-08-14 cs.CL cs.CY 新提交

New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs

新术语,新毒性:基于共识的新词毒性检测——通过搜索增强型大语言模型(LLMs)

Shiyao Cui, QingLin Zhang, Di Wang, Yida Lu, Zhexin Zhang, Jinhua Gao, Jinglin Yang, Min He, Han Qiu, Minlie Huang

机构 * Tsinghua University(清华大学) ICT, CAS(中国科学院计算技术研究所) National Computer Network Emergency Response Technical Team Coordination Center of China(国家计算机网络应急技术处理协调中心) IIE, CAS(中国科学院工程热物理研究所) School of Cyber Security, UCAS(中国科学院大学网络空间安全学院) JCSS, Tsinghua University(清华大学软件学院) Science City (Guangzhou) Digital Technology Group Co., Ltd.(广州科学城数字科技集团有限公司)

AI总结 本文针对毒性新词的检测问题,提出捕捉毒性新词起源与共识标准的分类法,构建风险词表,引入搜索增强型框架SeTox,实验显示3B规模的SeTox性能优于近期大规模模型。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12332 2026-08-14 cs.CL cs.LG 新提交

Can Spectral-Clipping Enable Better Learning While Forgetting Less for Low-Rank Adaptation?

频谱裁剪能否在低秩适配中实现更好的学习同时减少遗忘?

Hyowon Wi, Noseong Park

AI总结 本研究提出SCLoRA方法,通过结合频谱裁剪与低秩适配(LoRA),在提升下游任务性能的同时缓解了灾难性遗忘,实现更好的学习且减少知识遗忘。

Comments ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05803 2026-08-14 cs.CL 版本更新

Human-like fleeting memory improves language learning but impairs reading time prediction in transformer language models

类人短暂记忆提升语言学习但损害变压器语言模型的阅读时间预测

Abishek Thamma, Micha Heilbron

机构 * University of Amsterdam, Amsterdam Brain and Cognition(阿姆斯特丹大学,阿姆斯特丹脑与认知中心) Vrije Universiteit Amsterdam, Department of Informatics(阿姆斯特丹自由大学,信息学院) Max Planck Institute for Psycholinguistics(马克斯·普朗克心理学语言学研究所)

AI总结 研究探讨了短暂记忆对语言学习和阅读时间预测的影响,发现短暂记忆提升语言学习但损害阅读时间预测,挑战了传统认知科学观点。

Comments v2: Revised after peer review. Accepted for publication in Transactions of the Association for Computational Linguistics v3: Added link to code repository. Code: https://github.com/drhanjones/fmt-llm

Journal ref Transactions of the Association for Computational Linguistics 14 (2026) 877-892

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13235 2026-08-14 cs.HC cs.AI cs.CL cs.CY cs.LG cs.SI

RubRIX: Rubric-Driven Risk Mitigation in Caregiver-AI Interactions

RubRIX:基于评分标准的风险缓解在护理人员-AI交互中

Drishti Goel, Jeongah Lee, Qiuyue Joy Zhong, Violeta J. Rodriguez, Daniel S. Brown, Ravi Karkar, Dong Whi Yoo, Koustuv Saha

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) OSF HealthCare(OSF医疗集团) Indiana University Indianapolis(印第安纳大学印第安纳波利斯分校)

AI总结 RubRIX提出了一种基于评分标准的框架,用于评估护理人员-AI交互中的风险,通过实证方法减少LLM响应的风险组件,为高负担情境下的AI支持系统提供评估方法。

Journal ref Findings of the Association for Computational Linguistics: ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11788 2026-08-13 cs.CL cs.AI 新提交

TELLME: Test-Enhanced Learning for Language Model Enrichment

TELLME:面向语言模型增强的测试增强学习

Minjun Kim, Inho Won, Hyeonseok Lim, MinKyu Kim, Junghun Yuk, Wooyoung Go, Jongyoul Park, Jungyeul Park, KyungTae Lim

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院) Seoul National University of Science and Technology(首尔科学技术大学) National Security Research Institute(国家安全研究院)

AI总结 本研究提出TELLME方法,将测试增强学习(TEL)原理与持续预训练(CPT)结合,缓解大语言模型领域自适应中数据与成本问题,在金融领域表现优于现有方法,长期记忆保留提升显著。

Comments Findings of the Association for Computational Linguistics: EACL 2026

Journal ref Findings of the Association for Computational Linguistics: EACL 2026, pages 1655-1677

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18452 2026-08-13 cs.CL

MedScore: Generalizable Factuality Evaluation of Free-Form Medical Answers by Domain-adapted Claim Decomposition and Verification

Heyuan Huang, Alexandra DeLucia, Vijay Murari Tiyyala, Mark Dredze

机构 * Center for Language and Speech Processing(语言与语音处理中心) Johns Hopkins University(约翰霍普金斯大学)

Comments Added generalizability experiment and examples on non-medical free-form answer. Added ablation study for MedCorp verification corpus and MedScore decomposition prompt

Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pages 14149-14180, San Diego, California, United States. Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11171 2026-08-12 cs.CL cs.AI cs.CY 新提交

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

从可解释性到控制:TrustNLP研讨会六年的启示

Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle, Anil Ramakrishna, Anubrata Das, Apurv Verma, Jwala Dhamala, Ninareh Mehrabi, Tharindu Kumarage, Yada Pruksachatkun, Yang Trista Cao, Kai-Wei Chang, Aram Galstyan

机构 * Meta Autodesk(欧特克公司) New Jersey Institute of Technology(新泽西理工学院) Salesforce(salesforce公司) University of California, Los Angeles(加州大学洛杉矶分校) Amazon AGI(亚马逊AGI)

AI总结 该研究基于TrustNLP研讨会六年论文,分析NLP可信领域从可解释性到生成式系统控制的转变,明确各信任维度的发展趋势及与领域整体的关联性,提出结构性见解与研究方向。

Comments 17 pages, 2 figures, 3 tables. Submitted to ACL ARR August 2026 cycle (EACL 2027)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09937 2026-08-12 cs.CL cs.CY 新提交

Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory

仔细考量文化:利用文化共识理论分析单文化与多文化场景下的大语言模型对齐

Krishna Pothugunta, John P. Lalor

AI总结 本研究利用文化共识理论,分析大语言模型在单/多文化场景下的对齐情况,发现模型存在文化结构误表征问题,该理论可用于区分模型反映人类多样性与算法同质化的情况。

Comments Accepted to ACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03645 2026-08-12 cs.CL cs.CY

LLM-MC-Affect: LLM-Based Monte Carlo Modeling of Affective Trajectories and Latent Ambiguity for Interpersonal Dynamic Insight

LLM-MC-Affect: 基于大语言模型的蒙特卡洛建模:情感轨迹与潜在模糊性的人际动态洞察

Yu-Zheng Lin, Bono Po-Jen Shih, John Paul Martin Encinas, Elizabeth Victoria Abraham Achom, Karan Himanshu Patel, Jesus Horacio Pacheco, Sicong Shao, Jyotikrishna Dass, Soheil Salehi, Pratik Satam

机构 * University of Arizona(亚利桑那大学) Pennsylvania State University(宾夕法尼亚州立大学) Universidad de Sonora(索尔纳大学) University of North Dakota(北达科他大学)

AI总结 本文提出LLM-MC-Affect框架,通过概率建模方法,将情感视为连续的潜在概率分布,从而捕捉人际互动中的情感轨迹和潜在模糊性,为动态分析提供新的视角和方法。

Comments Accepted to the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26148 2026-08-12 cs.HC cs.CL 版本更新

Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations

超越截图:评估VLMs对UI动画的理解

Chen Liang, Xirui Jiang, Naihao Deng, Eytan Adar, Anhong Guo

机构 * University of Michigan(密歇根大学)

AI总结 本文提出AniMINT数据集,评估VLMs对动态UI动画的理解能力,发现其在基础运动检测上可靠,但高阶解释仍不一致,揭示了模型性能的关键瓶颈。

Comments Published in ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08829 2026-08-11 cs.CL cs.AI 新提交

Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models

面向大型语言模型的可部署实例级多层激活调控

Muhammad Faishal Adly Nelwan, Alfan Farizki Wicaksono

AI总结 本研究提出可部署的实例级多层激活调控方案,解决现有全局固定层调控的缺陷,在两个8B模型和六个人格特质任务上,实现接近先验的性能,且避免流畅性崩溃。

Comments 43 pages, 24 figures, 30 tables. Under review at ACL Rolling Review (August 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17653 2026-08-11 cs.CL

Differences in Typological Alignment in Language Models' Treatment of Differential Argument Marking

语言模型处理差异论元标记中的类型学对齐差异

Iskar Deng, Nathalia Xu, Shane Steinert-Threlkeld

机构 * University of Washington(华盛顿大学)

AI总结 通过控制合成语料训练GPT-2模型,发现模型在标记方向(自然标记方向)上表现出类人偏好,但在论元角色偏好(宾语vs主语)上未复现人类语言的强烈宾语偏好。

Comments 16 pages, 8 figures, 7 tables. To appear at CoNLL 2026

Journal ref Proceedings of the 30th Conference on Computational Natural Language Learning (CoNLL 2026), pp. 268-283, San Diego, California, USA, Association for Computational Linguistics, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17239 2026-08-11 cs.CL cs.AI

SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging

SafeMERGE: 通过选择性层间模型融合在微调大语言模型中保持安全对齐

Aladin Djuhera, Swanand Ravindra Kadhe, Farhan Ahmed, Syed Zawad, Holger Boche

机构 * Technical University Munich(慕尼黑技术大学) IBM Research(IBM研究院)

AI总结 本文提出SafeMERGE,一种轻量级后微调框架,通过选择性融合层间模型来恢复安全对齐,同时保持下游性能,减少有害输出并提升实用性。

Journal ref Findings of the ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11741 2026-08-11 cs.AI 版本更新

Collaborative Multi-Agent Scripts Generation for Enhancing Imperfect-Information Reasoning in Murder Mystery Games

协作多智能体脚本生成以增强谋杀谜游戏中的不完全信息推理

Keyang Zhong, Junlin Xie, Hefeng Wu, Haofeng Li, Guanbin Li

机构 * Sun Yat-sen University(中山大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

AI总结 本文提出协作多智能体框架,通过生成丰富多模态上下文提升VLMs在叙事推理和欺骗鲁棒性中的表现,解决多玩家游戏中不完全信息下的复杂推理问题。

Comments 9 pages, 5 figures, Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03005 2026-08-11 cs.CL q-bio.PE 版本更新

Beyond cognacy

超越词源关系

Gerhard Jäger

机构 * University of Tübingen(蒂宾根大学)

AI总结 本文比较了传统词源集方法与两种自动方法在语言谱系分析中的效果,发现基于MSA的方法能产生更一致的树结构,预测 typological 变异更准确,具有更清晰的谱系信号。

Comments 10 pages, 1 figure. v3: substantially corrected version - retrained pHMM, corrected T-Coffee implementation, method-independent guide tree, full pipeline re-run on current Lexibank data (974 languages, 15 families); all reported numbers changed, qualitative conclusions unchanged (see title footnote). Data and code: doi:10.57754/FDAT.8n2vd-1bq37

Journal ref Proceedings of the 7th Workshop on Research in Computational Linguistic Typology and Multilingual NLP (SIGTYP 2025), Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03115 2026-08-11 cs.CL eess.AS 版本更新

Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models

在大型音频-语言模型中发现并因果验证情绪敏感神经元

Xiutian Zhao, Björn Schuller, Berrak Sisman

机构 * Center for Language and Speech Processing (CLSP)(语言与语音处理中心) Johns Hopkins University(约翰霍普金斯大学) Group on Language, Audio & Music (GLAM)(语言、音频与音乐小组) Imperial College London(伦敦帝国学院)

AI总结 本研究通过神经元层面的干预验证了大型音频-语言模型中情绪敏感神经元的存在,并揭示了情绪识别的因果机制。

Comments Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00549 2026-08-11 cs.CL cs.AI

Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages

Hyangsuk Min, Yuho Lee, Minjeong Ban, Jiaqi Deng, Nicole Hee-Yeon Kim, Taewon Yun, Hang Su, Jason Cai, Hwanjun Song

机构 * Korea Advanced Institute of Science and Technology(韩国先进科学研究院) AWS AI Labs(AWS人工智能实验室)

Comments 34 pages, 6 figures

Journal ref ACL 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06506 2026-08-10 cs.CL 新提交

Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand

测量跨语言理解差距:证据语言如何影响语言模型的理解能力

Rafael da Silva, Jeff Eicher

机构 * Eastern University(东方大学)

AI总结 该研究定义跨语言理解差距(CLCG),通过ParallelQA-18评估5种模型,发现英语能力无法同等迁移,低资源语言用户的模型质量可能被英语中心评估高估。

Comments 55 pages, 17 figures. Submitted to Computational Linguistics (MIT Press / ACL). Supplementary Material: 55 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06756 2026-08-10 cs.CL 版本更新

How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality

推理链长度如何影响LLM对答案事实性的判断

Minzhu Tu, Shiyu Ni, Keping Bi

机构 * State Key Laboratory of AI Safety(人工智能安全国家重点实验室) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学) Beijing University of Post and Telecommunications(北京邮电大学)

AI总结 研究探讨了推理链对LLM判断答案事实性的影响,发现弱判官易受推理存在影响,而强判官部分利用推理作为证据,但仍易被高质量推理链误导。

Comments ACL2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17181 2026-08-07 cs.SE cs.AI cs.CL 交叉投稿

A Study of LLMs' Preferences for Libraries and Programming Languages

对大型语言模型在库和编程语言偏好方面的研究

Lukas Twist, Mark Harman, Don Syme, Joost Noppen, Helen Yannakoudakis, Detlef Nauck, Jie M. Zhang

机构 * King’s College London(伦敦国王学院) University College London(伦敦大学学院) Digital AI Research, BT Group(BT集团数字人工智能研究)

AI总结 本研究探讨了大型语言模型在生成代码时对库和编程语言的选择偏好,通过实证研究分析了八种不同大型语言模型在库和语言选择上的倾向,发现模型倾向于使用广泛采用的库如NumPy,并且在某些情况下这种选择并非必要,同时也显示出对Python的偏好,尽管在某些高性能项目初始化任务中Python并非最优选择。

Comments 21 pages, 10 tables, 3 figures. Accepted to Findings of ACL 2026

Journal ref Findings of the Association for Computational Linguistics: ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11936 2026-08-07 cs.CL cs.AI cs.CV cs.LG

A Survey of Deep Learning for Geometry Problem Solving

深度学习在几何问题求解中的应用综述

Jianzhe Ma, Wenxuan Wang, Qin Jin

机构 * Renmin University of China(中国人民大学)

AI总结 本文综述了深度学习在几何问题求解中的应用,涵盖相关任务、方法、评估指标及未来方向,旨在提供实践参考以推动该领域发展。

Comments ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07832 2026-08-07 cs.CL cs.CV

LADDER: Language-Driven Slice Discovery and Error Rectification in Vision Classifiers

Shantanu Ghosh, Rayan Syed, Chenyu Wang, Vaibhav Choudhary, Binxu Li, Clare B. Poynton, Shyam Visweswaran, Kayhan Batmanghelich

机构 * Boston University(波士顿大学) Stanford University(斯坦福大学) Boston University Medical Campus(波士顿大学医学校区) University of Pittsburgh(匹兹堡大学)

Comments ACL 2025. Code: https://github.com/batmanlab/Ladder

Journal ref Findings of the Association for Computational Linguistics: ACL 2025, pages 22935-22970

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04899 2026-08-06 cs.CL 新提交

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification

基于大语言模型的分类置信度估计中的评估缺陷与稀疏性限制

Elena Merdjanovska, Omar Zaidan, Andreas Rücklé

机构 * Humboldt-Universität zu Berlin(柏林洪堡大学) Amazon(亚马逊公司)

AI总结 该研究指出 LLM 分类置信度估计中 verbalization 方法存在稀疏性缺陷,AUARC 评估的插值选择会影响排名,提出 verbalization logprobs 方法可解决稀疏性并提升 AUARC 且无额外推理成本。

Comments Published at Findings of ACL 2026

Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pages 33424-33435

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04576 2026-08-06 cs.CL 新提交

Causal Evidence Extraction and Triangulation in Crisis Reports using Large Language Models: A ReliefWeb-based Study

基于大型语言模型的危机报告中因果证据提取与三角验证:一项基于ReliefWeb的研究

Yuanjun Zhang, Mourad Oussalah

机构 * University of Oulu(奥卢大学) LUT University(拉普兰塔理工大学)

AI总结 该研究基于ReliefWeb数据,提出两阶段LLM流水线提取危机报告中结构化因果证据,结合上下文保留的三角验证方法,在100份专家标注报告中取得优异性能,为现金援助的食品相关结果提供了高收敛性证据。

Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pages 32478-32491, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04056 2026-08-06 cs.CL cs.CY cs.LG 新提交

Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization

基于多智能体视角偏好优化的性别歧视检测学习

Hadi Mohammadi, Tina Shahedi, Robert A. Bagheri, Mehdi Dastani, Masoume M. Raeissi

机构 * Utrecht University(乌得勒支大学) Wageningen University & Research(瓦赫宁根大学及研究中心)

AI总结 本研究针对性别歧视检测中标注分歧问题,提出MAP-PO框架,通过聚类标注者并训练对应智能体,结合个体与团队奖励协调智能体,实验验证了聚类训练及团队信号的必要性。

Comments 17 pages, 12 figures, 14 tables. Preprint; under review at EACL 2027 (ACL Rolling Review, August 2026 cycle). Code and data: https://github.com/mohammadi-hadi/MAP-PO

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02087 2026-08-06 cs.AI cs.CL cs.LG 版本更新

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy

基于非对称强化学习与自蒸馏的指令条件探索

Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein

机构 * University of Southampton(南安普顿大学)

AI总结 该研究针对LLM强化学习中的探索挑战,提出指令条件探索(ICE)方法,结合非对称RL/SD训练目标,使Qwen3-1.7B数学推理性能提升5.0%且长上下文下仍有效。

Comments Submitted to ACL Rolling Review (ARR) May 2026 cycle. OpenReview submission record at https://openreview.net/forum?id=PV945lekMa

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18199 2026-08-06 cs.CL

Linear-Time and Constant-Memory Text Embeddings Based on Recurrent Language Models

基于循环语言模型的线性时间与常数内存文本嵌入

Tobias Grantner, Emanuel Sallinger, Martin Flechl

机构 * Dynatrace Research(DynaTrace研究)

AI总结 本文提出基于循环架构的高效文本嵌入方法,通过垂直分块推理策略实现线性时间复杂度和常数内存使用,验证了Mamba2等模型在文本嵌入任务中的有效性。

Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2026, pages 41459-41481

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21692 2026-08-06 cs.CL cs.AI 版本更新

Revisiting Generalization Across Difficulty Levels: It's Not So Easy

重新审视不同难度层级间的泛化:这并不容易

Yeganeh Kordi, Nihal V. Nayak, Max Zuo, Ilana Nguyen, Stephen H. Bach

机构 * Brown University(布朗大学) Harvard University(哈佛大学)

AI总结 本文研究了LLMs在不同任务难度间泛化的能力,发现训练数据的难度对泛化效果影响有限,强调在训练和评估中需涵盖多种难度以避免风险。

Comments Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19358 2026-08-06 cs.CL cs.AI

Benchmarking and Improving LLM Robustness for Personalized Generation

Chimaobi Okite, Naihao Deng, Kiran Bodipati, Huaidian Hou, Joyce Chai, Rada Mihalcea

机构 * University of Michigan(密歇根大学)

Comments First draft. First camera-ready version

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03810 2026-08-05 cs.CL cs.AI 新提交

VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs

VIBE:面向大型语言模型输出的以实体为中心的情感分析基准,基于VAD(效价-唤醒度-支配度)

Andrei Chetvergov, Alexander Evseev, Timofei Sivoraksha, Stepan Ukolov, Mikhail Solovev, Danil Sazanakov, Sergey Bolovtsov

AI总结 本文提出基于VAD的VIBE基准,用于以实体为中心分析大型语言模型输出的情感,明确了测量契约,通过三个实证层面验证了相关假设,推动该领域成为规范化实践。

Comments 25 pages, 13 figures, 22 tables. Submitted to ACL Rolling Review, August 2026

详情

展开后加载摘要…

URL PDF HTML 收藏