arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2026-04-24 至 2026-04-24 共收录 79 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 20 篇

2604.21510 2026-04-24 cs.CL 57%

OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving

OptiVerse:一个面向优化问题求解的综合基准

Xinyu Zhang, Boxuan Zhang, Yuchen Wan, Lingling Zhang, YiXing Yao, Bifan Wei, Yaqiang Wu, Jun Liu

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) Ministry of Education Key Laboratory of Intelligent Networks and Network Security, China(教育部智能网络与网络安全重点实验室) Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, China(陕西省大数据知识工程重点实验室) Lenovo Research(联想研究院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 OptiVerse通过1000个跨领域问题评估LLM在复杂优化任务中的表现,揭示模型在难题上的性能下降,并提出Dual-View Auditor Agent提升准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21508 2026-04-24 cs.AI q-bio.BM 57%

BioMiner: A Multi-modal System for Automated Mining of Protein-Ligand Bioactivity Data from Literature

BioMiner:一种用于从文献中自动挖掘蛋白质-配体生物活性数据的多模态系统

Jiaxian Yan, Jintao Zhu, Yuhang Yang, Qi Liu, Kai Zhang, Zaixi Zhang, Xukai Liu, Boyan Zhang, Kaiyuan Gao, Jinchuan Xiao, Enhong Chen

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(中国科学技术大学认知智能国家重点实验室) Princeton University(普林斯顿大学) Huazhong University of Science and Technology(华中科技大学) Infinite Intelligence Pharma(无限智能制药)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 BioMiner通过多模态框架提取生物活性数据,解决手动整理与文献增长不匹配的问题,其核心方法结合直接推理和化学结构解析,实现生物活性三元组的高准确率提取。

Comments 20 pages, 5 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21420 2026-04-24 cs.AI 57%

FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation

FairQE:缓解翻译质量评估中性别偏见的多智能体框架

Jinhee Jang, Juhwan Choi, Dongjin Lee, Seunguk Yu, Youngbin Kim

机构 * Chung-Ang University(Chung-Ang 大学) AITRICS

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 FairQE通过多智能体机制检测性别线索并生成性别翻转翻译变体,结合传统QE评分与LLM偏见缓解推理,提升翻译质量评估的性别公平性。

Comments Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21308 2026-04-24 cs.CR cs.CL 57%

CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents

CI-Work: 企业LLM代理中情境完整性基准测试

Wenjie Fu, Xiaoting Qin, Jue Zhang, Qingwei Lin, Lukas Wutschitz, Robert Sim, Saravan Rajmohan, Dongmei Zhang

机构 * Huazhong University of Science and Technology(华中科技大学) Microsoft(微软)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 CI-Work基准测试评估企业LLM代理在密集检索中传达关键内容并隐藏敏感信息的能力,揭示隐私泄露普遍且高任务效用与隐私违规正相关,需转向以情境为中心的架构。

Journal ref The 64th Annual Meeting of the Association for Computational Linguistics (ACL'2026) -- Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21255 2026-04-24 cs.CL 57%

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors

当代理看起来相同:量化蒸馏诱导的工具使用行为相似性

Chenghao Yang, Yuning Zhang, Zhoufutu Wen, Tao Gong, Jiaheng Liu, Qi Chu, Nenghai Yu

机构 * School of Cyber Science and Technology, USTC(USTC计算机科学与技术学院) Anhui Province Key Laboratory of Digital Security(安徽省数字安全重点实验室) M-A-P

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 本文提出两种新指标RPS和AGS,用于区分任务成功必需行为与模型自主偏好,通过评估18个模型发现AGS能区分教师特定收敛与通用改进。

Comments Accepted by ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20848 2026-04-24 cs.IR cs.AI 57%

MATRAG: Multi-Agent Transparent Retrieval-Augmented Generation for Explainable Recommendations

MATRAG:多智能体透明检索增强生成用于可解释推荐

Sushant Mehta

机构 * ACM

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 MATRAG通过多智能体协作与知识图谱增强检索,提升推荐系统的透明性和可解释性,实验表明其在准确率和可信度上均优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21697 2026-04-24 cs.CR cs.AI cs.MM 57%

Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models

结构化视觉叙事削弱多模态大语言模型的安全对齐

Rui Yang Tan, Yujia Hu, Roy Ka-Wei Lee

机构 * Singapore University of Technology and Design(新加坡技术与设计大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 研究发现结构化视觉叙事在多模态大语言模型中导致安全对齐问题,通过漫画模板劫持攻击展示其危害,揭示现有防御方法在处理非有害敏感内容时的不可靠性。

Comments Code released at: https://github.com/Social-AI-Studio/ComicJailbreak

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11044 2026-04-24 cs.AI 57%

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

AgencyBench:在100万token现实场景中自主代理的前沿评估

Keyu Li, Junhao Shi, Yang Xiao, Mohan Jiang, Jie Sun, Yunze Wu, Dayuan Fu, Shijie Xia, Xiaojie Cai, Tianze Xu, Weiye Si, Wenjie Li, Dequan Wang, Pengfei Liu

机构 * SII Open Source(SII开源)

专题命中 推理评测 :self-correction(abstract);分类 cs.AI

AI总结 本文提出AgencyBench,通过138个现实场景评估6种核心代理能力,揭示闭源模型在资源效率和反馈修正上的优势,为下一代自主代理的发展提供测试平台。

Comments Accepted by ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10003 2026-04-24 cs.CL 57%

SocraticKG: Knowledge Graph Construction via QA-Driven Fact Extraction

SocraticKG:通过问答驱动的事实提取构建知识图谱

Sanghyeok Choi, Woosang Jeon, Kyuseok Yang, Taehyeong Kim

机构 * Department of Biosystems Engineering, Seoul National University(首尔国立大学生物系统工程系) Artificial Intelligence Institute, Seoul National University(首尔国立大学人工智能研究所) Interdisciplinary Program in Artificial Intelligence, Seoul National University(首尔国立大学人工智能跨学科项目) Interdisciplinary Program in Cognitive Science, Seoul National University(首尔国立大学认知科学跨学科项目)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 本文提出SocraticKG方法,通过引入问答对作为中间表示,系统展开文档语义,提升知识图谱的事实保留与结构连贯性,支持复杂多跳推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20697 2026-04-24 cs.SD cs.AI 57%

Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores

音乐乐谱理解基准:评估大型语言模型对完整乐谱的理解能力

Congren Dai, Yue Yang, Krinos Li, Huichi Zhou, Shijie Liang, Bo Zhang, Enyang Liu, Ge Jin, Hongran An, Haosen Zhang, Peiyuan Jing, Kinhei Lee, Z henxuan Zhang, Xiaobing Li, Maosong Sun

机构 * Central Conservatory of Music(中央音乐学院) Imperial College London(帝国理工学院) Tsinghua University(清华大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 本文提出MSU-Bench基准,评估大型语言模型和视觉-语言模型对完整乐谱的理解能力,发现模型在模态间存在显著差距,通过微调可提升多层级正确性,为多模态推理研究提供基础。

Comments Accepted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他推理 9 篇

2604.21357 2026-04-24 cs.AI cs.CL 84%

ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs

ReaGeo:基于大语言模型的推理增强端到端地名编码

Jian Cui, Zhiyuan Ren, Desheng Weng, Yongqi Zhao, Gong Wenbin, Yu Lei, Zhenning Dong

机构 * Amap, Alibaba Group(阿里巴巴集团阿地图) Tsinghua University(清华大学)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 ReaGeo利用大语言模型解决传统多阶段方法在地理数据库中依赖文本或向量相似性检索的局限性,通过将坐标转换为geohash序列,引入链式推理机制提升空间关系推理能力,并通过距离偏差奖励的强化学习优化生成精度。

Comments 12 pages, 8 figures, submitted to ACM SIGSPATIAL 2024 (under review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21264 2026-04-24 cs.AI 77%

Enhancing Online Recruitment with Category-Aware MoE and LLM-based Data Augmentation

通过类别感知的MoE和基于LLM的数据增强提升在线招聘

Minping Chen, Bing Xu, Zulong Chen, Chuanfei Xu, Ying Zhou, Zui Tao, Zeyi Wen

机构 * HKUST (GZ)(香港科技大学(广州)) HKUST(香港科技大学) Alibaba Group(阿里巴巴集团) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室(深圳)) Zhijiang Lab(浙江实验室) The Hong Kong Polytechnic University(香港理工大学)

专题命中 其他推理 :CoT(abstract,abstract_cn);chain-of-thought(abstract);分类 cs.AI

AI总结 本文提出基于LLM的数据增强和类别感知MoE方法,解决低质量职位描述和相似候选-职位对的问题,提升AUC、GAUC和点击转化率。

Comments Accepted to ACL Industry Track 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21043 2026-04-24 cs.CY cs.AI cs.LG 62%

Strategic Polysemy in AI Discourse: A Philosophical Analysis of Language, Hype, and Power

人工智能话语中的战略多义性:语言、炒作与权力的哲学分析

Travis LaCroix, Fintan Mallory, Sasha Luccioni

机构 * Durham University(杜伦大学)

专题命中 其他推理 :chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 本文探讨人工智能话语中语言的策略性使用,分析术语如'幻觉'、'思考链'等的多义性如何影响机构和话语实践,揭示语言作为社会技术机制塑造AI发展与治理的作用。

Comments Accepted in the Ninth Annual ACM Conference on Fairness, Accountability, and Transparency (ACM FAccT) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20487 2026-04-24 cs.CL cs.AI 62%

Knowledge Capsules: Structured Nonparametric Memory Units for LLMs

知识胶囊:用于大语言模型的结构非参数记忆单元

Bin Ju, Shenfeng Weng, Danying Zhou, Rongkai Xu, Kunkai Su

机构 * Zhejiang Angel Medical AI Technology Co., Ltd.(浙江天使医疗人工智能科技有限公司) Miti AI Technology Co., Ltd.(米蒂人工智能科技有限公司) China-Singapore Belt and Road Joint Laboratory on Translational Infection Biology for Diagnostics and Therapies(中英一带一路联合实验室(转化感染生物学诊断与治疗)) State Key Laboratory for Diagnosis and Treatment of Infectious Diseases(传染病诊断与治疗国家重点实验室) The First Affiliated Hospital(第一附属医院)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出知识胶囊,一种结构化的非参数记忆单元,通过冻结基础模型直接从文档语料构建,采用外部键值注入框架提升外部知识在注意力计算中的参与度,提升长上下文和多跳推理的稳定性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21536 2026-04-24 cs.IR cs.AI 57%

Pre-trained LLMs Meet Sequential Recommenders: Efficient User-Centric Knowledge Distillation

预训练大语言模型遇见序列推荐者:高效的用户导向知识蒸馏

Nikita Severin, Danil Kartushov, Vladislav Urzhumov, Vladislav Kulikov, Oksana Konovalova, Alexey Grishanov, Anton Klenitskiy, Artem Fatkulin, Alexey Vasilev, Andrey Savchenko, Ilya Makarov

机构 * Independent researcher(独立研究者) Sber AI Lab(Sber AI实验室) Innopolis University(Innopolis大学) HSE University(俄罗斯高等经济大学) ITMO University(ITMO大学) AIRI

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 本文提出一种高效的用户导向知识蒸馏方法,利用预训练大语言模型生成的文本用户资料提升序列推荐系统对用户语义的理解,无需实时调用大语言模型,保持传统序列模型的推理效率。

Comments Accepted to ECIR 2026. 7 pages. This version of the contribution has been accepted for publication, after peer review but is not the Version of Record and does not reflect post-acceptance improvements, or any corrections. The Version of Record is available online at: http://dx.doi.org/10.1007/978-3-032-21300-6_42

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21193 2026-04-24 cs.AI 57%

Trust but Verify: Introducing DAVinCI -- A Framework for Dual Attribution and Verification in Claim Inference for Language Models

信任但验证:介绍DAVinCI——一种用于语言模型声明推断中双属性和验证的框架

Vipula Rawte, Ryan Rossi, Franck Dernoncourt, Nedim Lipka

机构 * Adobe(Adobe公司) Adobe Research(Adobe研究院)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 本文提出DAVinCI框架,通过双属性和验证机制提升语言模型输出的事实可靠性与可解释性,实验表明其在分类准确率、属性精度、召回率和F1分数上提升5-20%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19811 2026-04-24 cs.CY cs.AI 57%

Model Capability Assessment and Safeguards for Biological Weaponization

生物武器化中的模型能力评估与安全措施

Michael Richter

机构 * Binghamton University(比灵顿大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

AI总结 本文通过测试多个AI模型在73个开放性STEM提示中的表现,评估其在生物武器化中的潜在风险,发现Gemini在某些情况下缺乏上下文意识,需加强安全措施以应对能力超越调节的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20064 2026-04-24 cs.LG 57%

Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs

非老虎机:在推测解码中具有证明无遗憾的草案模型选择

Hongyi Liu, Jiaji Huang, Zhen Jia, Youngsuk Park, Yu-Xiang Wang

机构 * Rice University(里士大学) Amazon Web Services(亚马逊网络服务) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 其他推理 :reasoning(abstract);分类 cs.LG

AI总结 本文提出一种在推测解码中能证明无遗憾的草案模型选择算法,通过评估所有草案模型而非仅选定模型,显著提升性能,适用于多种解码方法,并在多个领域优于现有方法。

Comments ICLR'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21455 2026-04-24 cs.HC 50%

The Privacy Guardian Agent: Towards Trustworthy AI Privacy Agents

隐私守护代理:迈向可信AI隐私代理

Vincent Freiberger

专题命中 其他推理 :reasoning(abstract)

AI总结 本文提出隐私守护代理,通过用户资料和上下文感知自动化隐私同意选择,同时在不确定或高风险情况时通知用户,确保透明和用户自主权。

Comments Position paper for the CHI26 Workshop "Moving Beyond Clicks: Rethinking Consent and User Control in the Age of AI"

详情

展开后加载摘要…

URL PDF HTML 收藏