arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Annual Meeting of the Association for Computational Linguistics · 会议 · Natural Language Processing

2026-04-20 至 2026-04-20 共收录 58
2604.14646 2026-04-20 cs.AI

Targeted Exploration via Unified Entropy Control for Reinforcement Learning

通过统一熵控制实现定向探索的强化学习

Chen Wang, Lai Wei, Yanzhi Zhang, Chenyang Shao, Zedong Dan, Weiran Huang, Ge Lan, Yue Wang

机构 * College of Software, Nankai University(南开大学软件学院) Zhongguancun Academy(中关村学院) Shanghai Jiao Tong University(上海交通大学) Chinese Academy of Sciences(中国科学院) Tsinghua University(清华大学) Sun Yat-sen Univeristy(中山大学)

AI总结 本文提出UEC-RL框架,通过定向探索机制和稳定器解决强化学习中的熵崩溃问题,提升大模型推理性能。

Comments Accepted for publication in Findings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13846 2026-04-20 cs.CL

Beyond Static Personas: Situational Personality Steering for Large Language Models

超越静态人设:面向大语言模型的情境性人格引导

Zesheng Wei, Mengxiang Li, Zilei Wang, Yang Deng

机构 * University of Science and Technology of China(中国科学技术大学) Singapore Management University(新加坡国立大学)

AI总结 本文提出IRIS框架,通过分析人设神经元揭示情境依赖性,实现无训练的神经元引导方法,在PersonalityBench和SPBench上验证了其在复杂情境下的泛化与鲁棒性。

Comments Accepted to Findings of ACL2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11182 2026-04-20 cs.CL

Evaluating Memory Capability in Continuous Lifelog Scenario

在连续生活日志场景中评估记忆能力

Jianjie Zheng, Zhichen Liu, Zhanyu Shen, Jingxiang Qu, Guanhua Chen, Yile Wang, Yang Xu, Yang Liu, Sijie Cheng

机构 * Southern University of Science and Technology(南方科技大学) Tsinghua University(清华大学) Shenzhen University(深圳大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 本文提出LifeDialBench基准,包含EgoMem和LifeMem两个子集,通过在线评估协议解决传统离线设置的时序泄漏问题,发现复杂记忆系统不如简单RAG基线表现,强调高保真上下文保存的重要性。

Comments 27 pages, 7 figures. ACL 2026 Findings camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05552 2026-04-20 cs.CL cs.AI

Context-Agent: Dynamic Discourse Trees for Non-Linear Dialogue

上下文代理:非线性对话的动态 discourse 树

Junan Hu, Shudan Guo, Wenqi Liu, Jianhua Yin, Yinwei Wei

机构 * Shandong University(山东大学)

AI总结 本文提出 Context-Agent 框架,通过动态树结构建模对话历史,解决非线性对话中的上下文管理问题,并引入 NTM 评估基准,提升多轮对话任务完成率和效率。

Comments 14 pages, 7 figures, ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11698 2026-04-20 cs.CV cs.AI cs.CL

OSCBench: Benchmarking Object State Change in Text-to-Video Generation

OSCBench:文本到视频生成中对象状态变化的基准测试

Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen

机构 * National University of Singapore(新加坡国立大学) Singapore Management University(新加坡管理大学) Carnegie Mellon University(卡内基梅隆大学) Fudan University(复旦大学)

AI总结 本文提出OSCBench基准,用于评估文本到视频模型在对象状态变化上的性能,揭示当前模型在处理新场景时的不足。

Comments ACL 2026 Main Conference, Project page: https://hanxjing.github.io/OSCBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01566 2026-04-20 cs.CL

FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based Agents

FS-Researcher:基于文件系统的长周期研究任务测试时扩展

Chiwei Zhu, Benfeng Xu, Mingxuan Du, Shaohan Wang, Xiaorui Wang, Zhendong Mao, Yongdong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Metastone Technology(MetaStone技术公司)

AI总结 本文提出FS-Researcher框架,通过持久化工作空间实现长周期研究任务的测试时扩展,采用双代理结构提升研究质量与扩展性。

Comments 22 pages, 6 figures; Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11374 2026-04-20 cs.CL

Reward Modeling for Scientific Writing Evaluation

科学写作评估的奖励建模

Furkan Şahinuç, Subhabrata Dutta, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab) Department of Computer Science and Hessian Center for AI (hessian.AI) Technical University of Darmstadt(普遍知识处理实验室(UKP实验室)计算机科学系海德堡人工智能中心(hessian.AI)达姆施塔特技术大学) Konrad Zuse School of Excellence in Learning and Intelligent Systems (ELIZA)(康拉德·祖斯卓越学习与智能系统学校(ELIZA))

AI总结 本文提出高效的开源奖励模型,通过两阶段训练框架提升科学写作评估的准确性和鲁棒性,实现跨任务泛化和动态评分标准适应。

Comments Accepted to ACL 2026 (Main). Project page: https://ukplab.github.io/acl2026-expert-rm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04377 2026-04-20 cs.CL cs.AI cs.LG

Disco-RAG: Discourse-Aware Retrieval-Augmented Generation

Disco-RAG: 语篇意识的检索增强生成

Dongqi Liu, Hang Ding, Qiming Feng, Xurong Xie, Zhucun Xue, Chengjie Wang, Jian Li, Jiangning Zhang, Yabiao Wang

机构 * Saarland University(萨尔兰大学) Shanghai Jiaotong University(上海交通大学) Fudan University(复旦大学) Zhejiang University(浙江大学) Tencent YouTu Lab(腾讯优图实验室)

AI总结 Discourse-aware RAG框架通过构建篇章内结构树和跨篇章修辞图,提升知识整合能力,在问答和长文档摘要任务中取得最佳效果。

Comments ACL 2026 Main & Long Conference Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02996 2026-04-20 cs.CL

Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners

大型推理模型是(尚未)多语言潜在推理器

Yihong Liu, Raoyuan Zhao, Hinrich Schütze, Michael A. Hedderich

机构 * cisnlp

AI总结 本文研究了多语言潜在推理在大型推理模型中的表现,发现其在资源丰富语言中表现较强,但在低资源语言中较弱,且在更难的基准测试中更不明显。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07515 2026-04-20 cs.CL cs.AI

TPA: Next Token Probability Attribution for Detecting Hallucinations in RAG

TPA:用于检测RAG中幻觉的下一个令牌概率归因

Pengqian Lu, Jie Lu, Anjin Liu, Guangquan Zhang

机构 * Australian Artificial Intelligence Institute (AAII)(澳大利亚人工智能研究所)

AI总结 TPA通过归因七个来源量化各组件对生成下一个令牌的影响,有效识别幻觉响应,实验显示其性能领先。

Comments Accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07173 2026-04-20 cs.LG

Improving the Throughput of Diffusion-based Large Language Models via a Training-Free Confidence-Aware Calibration

通过训练无关的置信度感知校准提高扩散式大语言模型的吞吐量

Jucheng Shen, Gaurav Sarkar, Yeonju Ro, Sharath Nittur Sridhar, Zhangyang Wang, Aditya Akella, Souvik Kundu

机构 * Rice University(里士大学) Intel(英特尔) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出CadLLM,一种无需训练的加速扩散式大语言模型(dLLMs)推理吞吐量的方法。通过动态调整生成块大小、步长和阈值,结合置信度控制,减少softmax开销,提升吞吐量。

Comments 12 pages, 3 figures. Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10262 2026-04-20 cs.CL cs.AI eess.AS

MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models

MTR-DuplexBench:面向全双工语音语言模型多轮对话全面评估的综合评测

He Zhang, Wenqian Cui, Haoning Xu, Xiaohui Li, Lei Zhu, Haoli Bai, Shaohua Ma, Irwin King

机构 * Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学) Huawei Technologies(华为技术)

AI总结 本文提出MTR-DuplexBench,用于评估全双工语音语言模型在多轮对话中的表现,涵盖对话质量、指令遵循和安全性等关键方面,揭示现有模型在多轮对话中的一致性问题。

Comments Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03056 2026-04-20 cs.CL cs.AI cs.LG

Reading Between the Lines: The One-Sided Conversation Problem

在话语之间阅读:单侧对话问题

Victoria Ebert, Rishabh Singh, Tuochao Chen, Noah A. Smith, Shyamnath Gollakota

机构 * Paul G. Allen School of Computer Science & Engineering, University of Washington(保罗·G·阿伦计算机科学与工程学院,华盛顿大学) Allen Institute for Artificial Intelligence(阿伦人工智能研究所) Hearvana AI

AI总结 本文提出单侧对话问题,研究如何从单侧对话中重建缺失发言和生成摘要,发现未来发言和发言长度信息有助于重建,而高质量摘要无需重建。

Comments 8 pages, 6 figures, 4 tables. Accepted to ACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13920 2026-04-20 cs.CL

FACTS: Table Summarization via Offline Template Generation with Agentic Workflows

FACTS: 通过离线模板生成实现基于表格的摘要的代理工作流

Ye Yuan, Mohammad Amin Shabani, Siqi Liu

机构 * McGill University(麦吉尔大学) Mila - Quebec AI Institute(魁北克人工智能研究所) RBC Borealis

AI总结 FACTS通过离线模板生成实现高效、准确且隐私合规的表格摘要,优于现有方法,适用于实际查询需求。

Comments Accepted by ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13829 2026-04-20 cs.CL cs.AI

A Linguistics-Aware LLM Watermarking via Syntactic Predictability

一种基于语言学的LLM水印方法:通过语法可预测性

Shinwoo Park, Hyejin Park, Hyeseon An, Yo-Sub Han

机构 * Yonsei University(延世大学) Rensselaer Polytechnic Institute(拉特格斯理工学院)

AI总结 本文提出STELA框架,通过语法可预测性动态调整水印信号,提升检测鲁棒性,适用于多种语言。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07774 2026-04-20 cs.CL

Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards

用评分奖励模型治愈大语言模型数学推理中的奇迹步骤

Youliang Yuan, Qiuyang Mang, Jingbang Chen, Hong Wan, Xiaoyuan Liu, Junjielong Xu, Jen-tse Huang, Wenxuan Wang, Wenxiang Jiao, Pinjia He

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen, China(数据科学学院,香港中文大学(深圳)) UC Berkeley(加州大学伯克利分校) Zhejiang University(浙江大学) Johns Hopkins University(约翰霍普金斯大学) Renmin University of China(中国人民大学) Xiaohongshu Inc.(小红书公司)

AI总结 本文通过评分奖励模型解决大语言模型数学推理中的奇迹步骤问题,通过系统分析和人类验证建立失败模式分类,提升推理准确性和可靠性。

Comments Accepted by ACL 2026 Main, 22 pages, 10 figures, 7 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06953 2026-04-20 cs.AI cs.CL

Revisiting the Uniform Information Density Hypothesis in LLM Reasoning

重新审视大语言模型推理中的均匀信息密度假说

Minju Gwak, Guijin Son, Jaehyung Kim

机构 * Yonsei University(延世大学) OneLine AI

AI总结 本文研究大语言模型推理中均匀信息密度假说的有效性,提出新框架量化局部和全局信息流均匀性,发现高质量推理在局部均匀但全局非均匀,证明均匀性优于其他内部信号预测推理质量。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25897 2026-04-20 cs.CL cs.AI cs.CY

RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity

RoleConflictBench: 一个用于评估LLM上下文敏感性的角色冲突场景基准

Jisu Shin, Hoyun Song, Juhyun Oh, Changgeon Ko, Eunsu Kim, Chani Jung, Alice Oh

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

AI总结 本文提出RoleConflictBench基准,通过生成13000个角色冲突场景,评估LLM在角色冲突中的上下文敏感性,发现模型更倾向于遵循角色偏好而非动态上下文线索。

Comments Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17292 2026-04-20 cs.CL cs.AI

Multi-View Attention Multiple-Instance Learning Enhanced by LLM Reasoning for Cognitive Distortion Detection

多视角注意力多实例学习增强的认知扭曲检测

Jun Seo Kim, Hyemi Kim, Woo Joo Oh, Hongjin Cho, Hochul Lee, Hye Hyeon Kim

机构 * Gachon University(加成大学) Korea Telecom Research(韩国电信研究所) Yonsei University(延世大学)

AI总结 本文提出结合大语言模型与多实例学习架构的方法,通过分解情绪、逻辑和行为成分提升认知扭曲检测的可解释性和推理能力。

Comments Accepted to the main conference of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21675 2026-04-20 cs.CL cs.CV cs.GR

Is this chart lying to me? Automating the detection of misleading visualizations

这张图表在欺骗我吗?自动化检测误导性可视化

Jonathan Tonglet, Jan Zimny, Tinne Tuytelaars, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, TU Darmstadt and National Research Center for Applied Cybersecurity ATHENE(普遍知识处理实验室(UKP实验室)、计算机科学系、图腾斯大学(TU Darmstadt)及应用网络安全国家研究中心ATHENE) Department of Electrical Engineering, KU Leuven(电子工程系、鲁文大学) Department of Computer Science, KU Leuven(计算机科学系、鲁文大学)

AI总结 研究提出Misviz基准测试集和合成数据集,用于自动化检测误导性可视化,发现该任务仍极具挑战性。

Comments Camera-ready version accepted at ACL 2026 Main conference. Code and data available at: https://github.com/UKPLab/acl2026-misviz

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16727 2026-04-20 cs.AI

Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with constraints

反思搜索器:通过带有约束的强化学习提高大语言模型的可靠性

Zhenyun Yin, Shujie Wang, Xuhong Wang, Xingjun Ma, Yinchun Wang

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

AI总结 本文提出反思搜索器框架,结合置信度校准与基于检索的搜索,通过强化学习优化准确性,提升模型输出的可靠性。

Comments Accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20020 2026-04-20 cs.AI cs.CL

Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning

具有人格分配的大型语言模型表现出类似人类的动机推理

Saloni Dash, Amélie Reymond, Emma S. Spiro, Aylin Caliskan

机构 * University of Washington(华盛顿大学)

AI总结 研究探讨了分配不同政治和社会人口属性的人格对大型语言模型动机推理的影响,发现人格分配的模型在评估信息真实性时表现下降,且政治人格在科学证据评估中更倾向于与自身身份一致的结论,传统去偏方法效果有限。

Comments ACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02261 2026-04-20 cs.IR cs.LG

What Makes LLMs Effective Sequential Recommenders? A Study on Preference Intensity and Temporal Context

是什么让大语言模型(LLMs)成为有效的序列推荐者?对偏好强度和时间上下文的研究

Zhongyu Ouyang, Qianlong Wen, Chunhui Zhang, Yanfang Ye, Soroush Vosoughi

机构 * Department of Computer Science, Dartmouth College(达特茅斯大学计算机科学系) ByteDance, US(字节跳动(美国)) Department of Computer Science and Engineering, University of Notre Dame(诺丁汉大学计算机科学与工程系)

AI总结 本文研究了LLMs在序列推荐中有效性的关键因素,提出RecPO框架通过整合显式和隐式反馈,提升推荐性能,强调偏好强度和时间上下文的重要性。

Comments Accepted The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21569 2026-04-20 cs.LG cs.AI cs.CL

ChemAmp: Amplified Chemistry Tools via Composable Agents

ChemAmp:通过可组合代理增强化学工具

Zhucong Li, Powei Chang, Jin Xiao, Zhijian Zhou, Qianyu He, Jiaqing Liang, Fenglei Cao, Xu Yinghui, Yuan Qi

机构 * Artificial Intelligence Innovation and Incubation Institute, Fudan University(复旦大学人工智能创新与孵化院) School of Data Science, Fudan University(复旦大学数据科学学院) College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) Department of Information and Intelligence Development, Zhongshan Hospital, Fudan University(复旦大学中山医院信息与智能发展部)

AI总结 ChemAmp通过优化动态协调提升化学工具能力,以有限数据构建任务专用超代理,在分子设计等任务中优于专用模型和通用LLM。

Comments Accepted to ACL 2026 Findings ; Code available at https://github.com/Chang-pw/ChemAmp

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16176 2026-04-20 cs.AI cs.CL

Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning

动态采样:自适应的自我意识迭代数据持久优化用于数学推理

Jun Rao, Xuebo Liu, Hexuan Deng, Zepeng Lin, Zixiong Yu, Jiansheng Wei, Xiaojun Meng, Min Zhang

机构 * Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学深圳研究院) Huawei Large Model Data Technology Lab(华为大模型数据技术实验室)

AI总结 本文提出SAI-DPO框架,通过动态调整数据分布以适应模型能力变化,提升数学推理效率,实验显示其在多个基准上优于静态方法。

Comments ACL2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13792 2026-04-20 cs.CL cs.AI

Interpretable Traces, Unexpected Outcomes: Investigating the Disconnect in Trace-Based Knowledge Distillation

可解释的轨迹,意外的结果:探讨基于轨迹的知识蒸馏中的断层

Siddhant Bhambri, Upasana Biswas, Subbarao Kambhampati

机构 * School of Computing & AI, Arizona State University(计算与人工智能学院,亚利桑那州立大学)

AI总结 本文通过规则问题分解设计实验,探讨基于轨迹的知识蒸馏中轨迹语义正确性与可解释性之间的断层,发现轨迹正确性并不能保证最终答案正确,且可解释的分解轨迹在准确性上不具优势。

Comments Accepted at The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01137 2026-04-20 cs.CL

Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models

跟随流动:文本到图像模型中文本令牌间信息流动的研究

Guy Kaplan, Michael Toker, Yuval Reif, Yonatan Belinkov, Roy Schwartz

机构 * Hebrew University of Jerusalem(海法大学) Technion – Israel Institute of Technology(技术学院 – 以色列理工学院) Kempner Institute, Harvard University(哈佛大学凯普纳研究所)

AI总结 本文研究文本到图像模型中令牌表示间的信息分布,发现信息常集中于少数几个令牌,且令牌间存在孤立与交互影响,改进编码可提升生成质量。

Comments Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04300 2026-04-20 cs.CV cs.AI

T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts

T2I-FactualBench:基于知识密集概念的文本到图像模型事实性评估基准

Ziwei Huang, Wanggui He, Quanyu Long, Yandi Wang, Haoyuan Li, Zhelun Yu, Fangxun Shu, Long Chan, Hao Jiang, Fei Wu, Leilei Gan

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团) Nanyang Technological University(南洋理工大学)

AI总结 本文提出T2I-FactualBench,通过知识密集概念评估文本到图像模型的事实性,揭示当前SOTA模型在事实性上的不足。

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 27501-27524, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏