arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-04-09 至 2026-04-09 共收录 277 信号源:cs.CL, cs.AI, cs.LG

1. 指令微调 14 篇

2512.14735 2026-04-09 q-fin.CP cs.AI cs.CV 57%

PyFi: Toward Pyramid-like Financial Image Understanding for VLMs via Adversarial Agents

PyFi: 通过对抗代理实现金字塔式金融图像理解的愿景语言模型

Yuqun Zhang, Yuxuan Zhao, Sijia Chen

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Yantai Research Institute, Harbin Engineering University(哈尔滨工程大学烟台研究院)

专题命中 指令微调 :language model(abstract);分类 cs.AI

AI总结 PyFi通过对抗代理生成的金字塔结构数据集,提升VLM在金融图像理解中的推理能力,实验显示模型在复杂金融问题上的准确率提升19.52%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09926 2026-04-09 cs.LG cs.CV 57%

LoFT: Parameter-Efficient Fine-Tuning for Long-tailed Semi-Supervised Learning in Open-World Scenarios

LoFT:长尾半监督学习中的参数高效微调:开放世界场景

Zhiyuan Huang, Jiahao Chen, Bing Su

机构 * Renmin University of China(中国人民大学)

专题命中 指令微调 :foundation model(abstract);分类 cs.LG

AI总结 LoFT通过参数高效微调提升长尾半监督学习在开放世界中的性能,理论证明基础模型降低假设复杂度并提升鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15527 2026-04-09 cond-mat.mtrl-sci 50%

NMR evidence for an antisite-induced magnetic moment on Bi in a topological insulator heterostructure MnBi$_2$Te$_4$/(Bi$_2$Te$_3$)$_n$

核磁共振证据表明Bi在拓扑绝缘体异质结MnBi$_2$Te$_4$/(Bi$_2$Te$_3$)$_n$中存在反位诱导的磁矩

R. Kalvig, E. Jedryka, A. Lynnyk, P. Skupinski, K. Grasza, M. Wojcik

专题命中 指令微调 :SFT(abstract)

AI总结 研究通过核磁共振和磁化测量揭示MnBi$_2$Te$_4$/(Bi$_2$Te$_3$)$_n$异质结中反位诱导的Bi磁矩,为理解拓扑绝缘体的磁相互作用提供新见解。

Comments 7 pages, 4 figures, Accepted for publication in Physical Review B

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 后训练与偏好优化 7 篇

2509.17183 2026-04-09 cs.CL cs.AI cs.LG 92%

LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization

LifeAlign: 为大型语言模型设计的终身对齐方法,结合记忆增强的聚焦偏好优化

Junsong Li, Jie Zhou, Bihao Zhan, Yutao Yang, Qianjun Pan, Shilian Chen, Tianyu Huai, Xin Li, Qin Chen, Liang He

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);preference optimization(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出LifeAlign,一种通过记忆增强的聚焦偏好优化实现终身对齐的方法,解决大型语言模型在连续学习任务中保持偏好对齐和知识保留的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06621 2026-04-09 cs.GL cs.LG stat.ML 81%

The Theorems of Dr. David Blackwell and Their Contributions to Artificial Intelligence

大卫·布莱克韦尔定理及其对人工智能的贡献

Napoleon Paxton

机构 * School of Information, University of California, Berkeley(加州大学伯克利分校信息学院)

专题命中 后训练与偏好优化 :LLM(abstract);large language model(abstract);language model(abstract);RLHF(abstract)

AI总结 本文综述了布莱克韦尔的三个重要定理及其对现代人工智能和机器学习的影响,包括马尔可夫链蒙特卡罗推断、自主机器人导航、生成模型训练等领域的应用。

Comments Survey article, 19 pages, 1 figure, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01840 2026-04-09 cs.AI 79%

Not All Tokens See Equally: Perception-Grounded Policy Optimization for Large Vision-Language Models

并非所有token都同等重要:基于感知的策略优化方法用于大型视觉-语言模型

Zekai Ye, Qiming Li, Xiaocheng Feng, Ruihan Chen, Ziming Li, Haoyu Ren, Kun Chen, Dandan Tu, Bing Qin

机构 * Harbin Institute of Technology(哈尔滨工业大学) Peng Cheng Laboratory(鹏城实验室) Huawei Technologies Co., Ltd(华为技术有限公司)

专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.AI

AI总结 本文提出基于感知的策略优化方法PGPO,通过动态调整token级别的优势,提升多模态推理的鲁棒性与稳定性,实验显示在多个基准上提升18.7%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06298 2026-04-09 cs.LG 77%

Limits of Difficulty Scaling: Hard Samples Yield Diminishing Returns in GRPO-Tuned SLMs

难度扩展的极限:在GRPO调优的SLMs中困难样本产生递减回报

Suraj Yadav, Siddharth Yadav, Parth Goyal

机构 * IIIT Delhi(德里印度理工学院)

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);preference optimization(abstract);分类 cs.LG

AI总结 本文研究了在资源受限条件下,GRPO与LoRA应用于SLMs进行数学推理时,难度增加对准确率的影响,发现高难度样本的提升效果递减。

Comments Accepted at ICLR Workshop 2026 ICBINB

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03647 2026-04-09 cs.CV cs.AI 77%

Stabilizing Unsupervised Self-Evolution of MLLMs via Continuous Softened Retracing reSampling

通过连续软化回溯重采样稳定多模态大语言模型的无监督自进化

Yunyao Yu, Zhengxian Wu, Zhuohong Chen, Hangrui Xu, Zirui Liao, Xiangwen Deng, Zhifang Liu, Senyuan Shi, Haoqian Wang

机构 * Tsinghua University(清华大学) Hefei University of Technology(合肥工业大学) University of Arizona(亚利桑那大学) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统实验室)

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);post-training(abstract);分类 cs.AI

AI总结 本文提出CSRS方法,通过回溯推理机制和软化频率奖励提升多模态大语言模型的无监督自进化稳定性,实验显示在MathVision等基准上表现优异。

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06771 2026-04-09 cs.CL cs.AI 62%

Multi-Faceted Self-Consistent Preference Alignment for Query Rewriting in Conversational Search

多维自一致偏好对齐的对话搜索查询重写

Zhiyu Cao, Peifeng Li, Qiaoming Zhu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院)

专题命中 后训练与偏好优化 :preference optimization(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MSPA-CQR方法,通过多维度自一致偏好对齐数据生成多样化查询,结合前缀引导的多维偏好优化,提升对话搜索中查询重写的效率与效果。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07210 2026-04-09 cs.CV 50%

VersaVogue: Visual Expert Orchestration and Preference Alignment for Unified Fashion Synthesis

VersaVogue: 视觉专家协作与偏好对齐的统一时尚合成

Jian Yu, Fei Shen, Cong Wang, Yi Xin, Si Shen, Xiaoyu Du, Jinhui Tang

机构 * Nanjing University of Science and Technology(南京理工大学) National University of Singapore(新加坡国立大学) Nanjing University(南京大学) Nanjing Forestry University(南京林业大学)

专题命中 后训练与偏好优化 :preference optimization(abstract)

AI总结 本文提出VersaVogue框架,通过 trait-routing attention 模块和自动化多视角偏好优化管道,实现多条件可控的时尚合成,提升视觉真实性和可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 长上下文与记忆 9 篇

2502.08875 2026-04-09 q-fin.GN 89%

Utilizing Pre-trained and Large Language Models for 10-K Items Segmentation

利用预训练和大语言模型进行10-K文件分段

Hsin-Min Lu, Yu-Tai Chien, Huan-Hsun Yen, Yen-Hsiu Chen

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract);prompting(abstract)

AI总结 本文提出两种先进分段方法,BERT4ItemSeg与GPT4ItemSeg,通过预训练模型与Bi-LSTM结合或大语言模型,提升10-K报告分段性能,核心指标达到宏F1值0.9825。

Comments Accepted for publication in the Journal of Information Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26083 2026-04-09 cs.LG cs.AI 79%

Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism

Nirvana:一种具有任务感知内存机制的专用通用模型

Yuhua Jiang, Shuang Cheng, Yihao Liu, Ermo Hua, Che Jiang, Weigao Sun, Yu Cheng, Feifei Gao, Biqing Qi, Bowen Zhou

机构 * Shanghai AI Laboratory(上海人工智能实验室) Tsinghua University(清华大学) Xiong’an Anying Technology Co., Ltd.(雄安安影科技有限公司) Zhejiang University(浙江大学)

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 Nirvana提出一种具备任务感知内存机制的专用通用模型,通过线性时间复杂度和测试时任务信息提取,在通用和专业领域均表现优异,尤其在医学、金融和法律等专业领域实现最低困惑度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07041 2026-04-09 cs.DB cs.AI cs.ET cs.HC cs.IR 77%

AV-SQL: Decomposing Complex Text-to-SQL Queries with Agentic Views

AV-SQL:通过代理视图分解复杂文本到SQL查询

Minh Tam Pham, Trinh Pham, Tong Chen, Hongzhi Yin, Quoc Viet Hung Nguyen, Thanh Tam Nguyen

机构 * Griffith University(格里菲斯大学) The University of Queensland(昆士兰大学)

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 AV-SQL通过代理视图分解复杂文本到SQL查询,提升复杂查询的执行准确率,其在Spider 2.0基准测试中达到70.38%的准确率,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06300 2026-04-09 cs.CY 75%

The End of Human Judgment in the Kill Chain? Relocating Initiative and Interpretation with Agentic AI

人类判断在杀链中的终结?通过代理AI重新定位主动性和解释力

Jovana Davidovic

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文探讨了基于大语言模型的代理AI在杀链中取代人类判断的问题,指出其主动性和解释力导致人类控制失效,并提出国际治理框架需应对这一挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06746 2026-04-09 cs.CL 70%

StructKV: Preserving the Structural Skeleton for Scalable Long-Context Inference

StructKV: 保持结构骨架以实现可扩展的长上下文推理

Zhirui Chen, Peiyang Liu, Ling Shao

机构 * UCAS-Terminus AI Lab, University of Chinese Academy of Sciences, China(中国科学院大学UCAS-Terminus AI实验室) National Engineering Research Center for Software Engineering, Peking University(北京大学国家软件工程研究中心)

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 StructKV通过全局入度中心性、动态枢轴检测和结构传播与解耦方法,有效保留长距离依赖性和检索鲁棒性,提升长上下文推理效率。

Comments Accepted to ACL 2026 Findings, 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06845 2026-04-09 cs.CL cs.AI 62%

HingeMem: Boundary Guided Long-Term Memory with Query Adaptive Retrieval for Scalable Dialogues

HingeMem: 基于边界的长期记忆与查询自适应检索用于可扩展对话

Yijie Zhong, Yunfan Gao, Haofen Wang

机构 * College of Design and Innovation, Tongji University(同济大学设计创意学院) Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University(同济大学上海自主智能无人系统科学中心)

专题命中 长上下文与记忆 :LLM(abstract);分类 cs.CL、cs.AI

AI总结 HingeMem通过边界引导的长期记忆和查询自适应检索,提升对话系统在多轮交互中的记忆效率与适应性,实验表明在不同规模的LLM上实现20%的性能提升并降低计算成本。

Comments Accepted by TheWebConf 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11471 2026-04-09 cs.LG 57%

Low-Rank Key Value Attention

低秩键值注意力

James O'Neill, Robert Clancy, Mariia Matskevichus, Fergal Reid

专题命中 长上下文与记忆 :pretraining(abstract);分类 cs.LG

AI总结 本文提出低秩键值注意力机制,通过利用注意力头间的冗余减少KV缓存内存消耗,同时保持计算效率,在多个基准测试中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06773 2026-04-09 cs.HC 50%

MemoryDiorama: Generating Dynamic 3D Diorama from Everyday Photos for Memory Recall

MemoryDiorama:从日常照片生成动态3D场景以增强记忆回忆

Keiichi Ihara, Tianle Li, Yasuhisa Shiino, Ryo Suzuki

专题命中 长上下文与记忆 :LLM(abstract)

AI总结 MemoryDiorama通过整合LLM场景分析与3D生成技术,将日常照片转化为动态3D场景,提升记忆回忆的丰富性与生动性。

Comments 11 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06506 2026-04-09 cs.CR cs.SE 50%

Guiding Symbolic Execution with Static Analysis and LLMs for Vulnerability Discovery

通过静态分析和LLMs引导符号执行以发现漏洞

Md Shafiuzzaman, Achintya Desai, Wenbo Guo, Tevfik Bultan

专题命中 长上下文与记忆 :LLM(abstract)

AI总结 SAILOR结合静态分析与LLM合成自动构建符号执行Harness,通过三个阶段发现内存安全漏洞,实验表明其在大规模代码库中显著提升漏洞检测效率。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 推理与问题求解 44 篇

2603.10512 2026-04-09 cs.AI cs.LG cs.NE 91%

Resource-constrained Amazons chess decision framework integrating large language models and graph attention

资源受限的亚马逊国际象棋决策框架整合大语言模型和图注意力

Tianhao Qian, Zhuoxuan Li, Jinde Cao, Xinli Shi, Leszek Rutkowski

机构 * School of Mathematics, Southeast University(东南大学数学学院) Systems Research Institute of the Polish Academy of Sciences(波兰科学院系统研究所) Institute of Computer Science, AGH University of Krakow(AGH科技大学计算机科学研究所) SAN University(SAN大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);foundation model(abstract)

AI总结 本文提出一种轻量级混合框架,结合图注意力学习和大语言模型生成能力,提升亚马逊国际象棋决策准确性,实验显示在资源受限条件下实现显著性能提升。

Comments 20 pages, 15 figures. Supported by the National Key Research and Development Project of China (No. 2020YFA0714300), NSFC (No. 61833005, 12061088), the Open Project of Key Laboratory of Transport Industry of Comprehensive Transportation Theory (Nanjing Modern Multimodal Transportation Laboratory) (MTF2023004), and the China Postdoctoral Science Foundation (2024T170129, GZC20240261)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01925 2026-04-09 cs.CL cs.AI 89%

Rectifying LLM Thought from Lens of Optimization

从优化角度纠正大语言模型的思维

Junnan Liu, Hongwei Liu, Songyang Zhang, Kai Chen

机构 * Monash University(莫纳什大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 本文通过优化视角分析LLM推理过程,提出RePro方法改进推理性能,减少过度思考等亚优行为。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06427 2026-04-09 cs.LG cs.AI cs.CL 88%

The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning

深度上限:大语言模型在发现潜在规划中的局限性

Yi Xu, Philipp Jettkant, Laura Ruis

机构 * University of Cambridge(剑桥大学) Imperial College London(帝国理工学院) MIT(麻省理工学院)

专题命中 推理与问题求解 :large language model(title);language model(title);prompting(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究探讨了大语言模型在无监督情况下发现多步规划策略的极限,发现模型在单次前向传递中能执行最多七步潜在规划,揭示了训练与测试时能力的差异。

Comments 10 pages, 3 figures, 1 table (30 pages, 9 figures, 10 tables including references and appendices)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06882 2026-04-09 cs.RO cs.SY eess.SP eess.SY 87%

Telecom World Models: Unifying Digital Twins, Foundation Models, and Predictive Planning for 6G

电信世界模型:统一数字孪生、基础模型和预测规划用于6G

Hang Zou, Yuzhi Yang, Lina Bariah, Yu Tian, Yuhuan Lu, Bohao Wang, Anis Bara, Brahim Mefgouda, Hao Liu, Yiwei Tao, Sergy Petrov, Salma Cheour, Nassim Sehad, Sumudu Samarakoon, Chongwen Huang, Samson Lasaulce, Mehdi Bennis, Mérouane Debbah

机构 * Research Institute for Digital Future, Khalifa University(哈利法大学数字未来研究所) College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息与电子工程学院) University of Oulu(奥卢大学) Aalto University(阿尔托大学) Université de Lorraine, CNRS, CRAN(洛林大学,法国国家科学研究中心,CRAN)

专题命中 推理与问题求解 :foundation model(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文提出电信世界模型(TWM),结合数字孪生与基础模型,解决6G网络中动态建模、不确定性处理和多层控制规划问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06834 2026-04-09 cs.CL cs.AI 86%

On the Step Length Confounding in LLM Reasoning Data Selection

在LLM推理数据选择中的步长混淆问题

Bing Wang, Rui Miao, Chen Shen, Shaotian Yan, Kaiyuan Liu, Ximing Li, Xiaosong Yuan, Sinan Fan, Jun Zhang, Jieping Ye

机构 * College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院) Key Laboratory of Symbolic Computation and Knowledge Engineering, MoE, Jilin University(吉林大学符号计算与知识工程教育部重点实验室) Alibaba Cloud Computing(阿里云) School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Department of Mathematics, University of Michigan(密歇根大学数学系) RIKEN Center for Advanced Intelligence Project(理化学研究所先进智能研究中心)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文分析了自然性基于的数据选择方法在LLM推理数据中系统性偏好长推理步骤的问题,并提出ASLEC-DROP和ASLEC-CASL方法以缓解此混淆。

Comments Accepted by Findings of ACL 2026. 15 pages, 9 figures. Code: https://github.com/wangbing1416/ASLEC

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06401 2026-04-09 cs.AI cs.CE cs.CV cs.LG 86%

ProofSketcher: Hybrid LLM + Lightweight Proof Checker for Reliable Math/Logic Reasoning

ProofSketcher: 混合LLM + 轻量级证明检查器用于可靠的数学/逻辑推理

Kranthi Kommuru, Kunal Khanvilkar, Gaurav Parekh

机构 * Amazon Web Services, Inc(亚马逊云科技)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出混合方法,结合LLM生成类型化证明草图与轻量级可信内核,以提升数学逻辑推理的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06204 2026-04-09 cs.CL cs.AI cs.HC 86%

SensorPersona: An LLM-Empowered System for Continual Persona Extraction from Longitudinal Mobile Sensor Streams

SensorPersona: 一种基于大语言模型的持续性人设提取系统,用于从长期移动传感器流中无感获取

Bufang Yang, Lilin Xu, Yixuan Li, Kaiwei Liu, Xiaofan Jiang, Zhenyu Yan

机构 * The Chinese University of Hong Kong(香港中文大学) Columbia University(哥伦比亚大学)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出SensorPersona系统,通过多模态长期传感器流持续提取稳定用户人设,提升LLM代理的个性化响应质量。系统通过上下文编码、分层人设推理和自适应更新,实现对物理行为、心理特质和生活经验的全面分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06421 2026-04-09 cs.CL 83%

State-of-the-Art Arabic Language Modeling with Sparse MoE Fine-Tuning and Chain-of-Thought Distillation

最先进的阿拉伯语言建模:基于稀疏MoE微调和链式思维蒸馏

Navan Preet Singh, Anurag Garikipati, Ahmed Abulkhair, Jyani Akshay Jagdishbhai, Atul Yaduvanshi, Amarendra Chaudhary, Madalina Ciobanu, Qingqing Mao, Ritankar Das

机构 * Incept Labs(因塞普特实验室)

专题命中 推理与问题求解 :language model(title);LLM(abstract);pretraining(abstract);分类 cs.CL

AI总结 本文提出阿拉伯-DeepSeek-R1,通过稀疏MoE架构和链式思维蒸馏,在开放阿拉伯LLM排行榜上实现SOTA,展示了稀疏MoE与文化导向的蒸馏方法在阿拉伯语言任务中的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08258 2026-04-09 cs.AI 83%

Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment

诊断和缓解大语言模型因果判断中的阿谀和怀疑

Edward Y. Chang

机构 * Stanford University(斯坦福大学)

专题命中 推理与问题求解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出CAUSALT3基准测试,通过三个轴评估分解性能,识别并缓解因果判断中的阿谀和怀疑陷阱,通过Regulated Causal Anchoring方法改进推理控制。

Comments 19 pages, 3 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06636 2026-04-09 cs.LG cs.AI cs.CL 82%

SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM Reasoning

SHAPE:基于潜在估计的阶段感知分层优势用于大语言模型推理

Zhengyang Ai, Zikang Shan, Xiaodong Ai, Jingxian Tang, Hangkai Hu, Pinyan Lu

机构 * Huawei Taylor Lab(华为泰勒实验室) Center for Data Science, Peking University(北京大学数据科学中心) Shanghai University of Finance and Economics(上海财经大学)

专题命中 推理与问题求解 :LLM(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 SHAPE通过潜在估计引入分层信用分配机制,提升大语言模型推理的效率与准确性,实验显示在数学推理任务中平均准确率提升3%,token消耗减少30%。

Comments ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06260 2026-04-09 cs.LG cs.AI 81%

$S^3$: Stratified Scaling Search for Test-Time in Diffusion Language Models

$S^3$:分层缩放搜索用于扩散语言模型的测试时间

Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Asad Aali, Muhammad Usman Khanzada, Muhammad Usman Rafique, Zihao He, Emily Fox, Dean F. Hougen

机构 * University of Oklahoma(俄克拉荷马大学) Stanford University(斯坦福大学) University of Göttingen(哥廷根大学) Zoox Meta

专题命中 推理与问题求解 :language model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出$S^3$方法,通过在去噪过程中重新分配计算资源提升生成质量,实验证明其在数学推理任务中表现最佳。

Comments Submitted to COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏