arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 4771 信号源:cs.CL, cs.AI, cs.LG

1. 长上下文与记忆 4771 篇

2512.04763 2025-12-05 cs.LG cs.CL cs.CV 82%

MemLoRA: Distilling Expert Adapters for On-Device Memory Systems

MemLoRA: 通过专家适配器实现设备端记忆系统

Massimo Bini, Ondrej Bohdal, Umberto Michieli, Zeynep Akata, Mete Ozay, Taha Ceritli

机构 * Samsung R&D Institute UK(三星英国研发中心) Technical University of Munich(慕尼黑技术大学) Helmholtz Munich(慕尼黑海德堡研究所)

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);small language model(abstract)

AI总结 MemLoRA通过为小型语言模型添加专用记忆适配器,实现设备端记忆系统的本地部署,并扩展至视觉理解任务,提升多模态场景下的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17800 2025-10-22 cs.CV cs.CL cs.LG 82%

Glyph: Scaling Context Windows via Visual-Text Compression

Jiale Cheng, Yusen Liu, Xinyu Zhang, Yulin Fei, Wenyi Hong, Ruiliang Lyu, Weihan Wang, Zhe Su, Xiaotao Gu, Xiao Liu, Yushi Bai, Jie Tang, Hongning Wang, Minlie Huang

机构 * The Conversational Artificial Intelligence (CoAI) Group, Tsinghua University(清华大学对话人工智能(CoAI)小组) Zhipu AI(智谱AI) The Knowledge Engineering Group (KEG), Tsinghua University(清华大学知识工程小组(KEG))

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);SFT(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15113 2025-08-05 cs.NE cs.AI cs.CL 82%

Associative memory inspires improvements for in-context learning using a novel attention residual stream architecture

Thomas F Burns, Tomoki Fukai, Christopher J Earls

机构 * SciAI Center, Cornell University, USA(SciAI中心,康奈尔大学,美国) Neural Coding and Brain Computing Unit, OIST, Japan(神经编码与脑计算单元,海洋研究所,日本) SciAI Center, Center for Applied Mathematics, School of Civil and Environmental Engineering, Cornell University, USA(SciAI中心,应用数学中心,土木与环境工程学院,康奈尔大学,美国)

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);small language model(abstract)

Comments 35 pages, 14 figures, 6 tables; accepted and published in TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21349 2025-06-23 cs.LG cs.AI cs.PF 82%

FALCON: Feedback-driven Adaptive Long/short-term memory reinforced Coding Optimization system

Zeyuan Li, Yangfan He, Lewei He, Jianhui Wang, Tianyu Shi, Bin Lei, Yuchen Li, Qiuwu Chen

机构 * School of Software, South China Normal University(南方科技大学软件学院) University of Minnesota - Twin Cities(明尼苏达大学双城分校) University of Electronic Science and Technology of China(电子科技大学) University of Toronto(多伦多大学) University of Connecticut(康涅狄格大学) AI center, AIGCode Inc.(AIGCode公司人工智能中心)

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);SFT(abstract);RLHF(abstract)

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12637 2025-04-18 cs.CL cs.AI 82%

Scaling Instruction-Tuned LLMs to Million-Token Contexts via Hierarchical Synthetic Data Generation

Linda He, Jue Wang, Maurice Weber, Shang Zhu, Ben Athiwaratkun, Ce Zhang

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);instruction tuning(abstract);post-training(abstract)

Comments 26 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22036 2025-04-04 cs.CL cs.AI 82%

Cognitive Prompts Using Guilford's Structure of Intellect Model

Oliver Kramer

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09597 2025-02-14 cs.LG cs.CL 82%

Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs

Siyan Zhao, Mingyi Hong, Yang Liu, Devamanyu Hazarika, Kaixiang Lin

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

Comments Accepted at ICLR 2025 as oral presentation. Code and data at: https://prefeval.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04139 2025-01-03 cs.CL cs.AI 82%

From Reading to Compressing: Exploring the Multi-document Reader for Prompt Compression

Eunseong Choi, Sunkyung Lee, Minjin Choi, June Park, Jongwuk Lee

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

Comments Findings of the Association for Computational Linguistics: EMNLP 2024; 21 pages; 10 figures and 7 tables. Code available at https://github.com/eunseongc/R2C

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15274 2024-12-23 cs.CL cs.AI 82%

Memory-Augmented Agent Training for Business Document Understanding

Jiale Liu, Yifan Zeng, Malte Højmark-Bertelsen, Marie Normann Gadeberg, Huazheng Wang, Qingyun Wu

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

Comments 11 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05317 2024-10-29 cs.LG cs.CL 82%

LoCoCo: Dropping In Convolutions for Long Context Compression

Ruisi Cai, Yuandong Tian, Zhangyang Wang, Beidi Chen

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);post-training(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07055 2024-08-14 cs.CL cs.LG 82%

LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs

Yushi Bai, Jiajie Zhang, Xin Lv, Linzhi Zheng, Siqi Zhu, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);SFT(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09727 2024-07-23 cs.CL cs.AI cs.IR 82%

A Human-Inspired Reading Agent with Gist Memory of Very Long Contexts

Kuang-Huei Lee, Xinyun Chen, Hiroki Furuta, John Canny, Ian Fischer

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

Comments Website: https://read-agent.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00029 2024-06-04 cs.CL cs.AI 82%

Clustered Retrieved Augmented Generation (CRAG)

Simon Akesson, Frances A. Santos

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13304 2023-05-23 cs.CL cs.LG 82%

RecurrentGPT: Interactive Generation of (Arbitrarily) Long Text

Wangchunshu Zhou, Yuchen Eleanor Jiang, Peng Cui, Tiannan Wang, Zhenxin Xiao, Yifan Hou, Ryan Cotterell, Mrinmaya Sachan

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29279 2026-08-03 cs.SD 新提交 82%

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition

ParaASR:面向快速长上下文基于大语言模型的语音识别的多令牌预测

Qingjian Lin, Yuxin Li, Haoyang Zhang, Jun Chen, Yechang Huang, Feng Tian, Xie Li, Xiangyu Tony Zhang, Daijiao Liu, Yuxin Zhang, Jinglan Gong, Bo Zhao, Fei Tian, Xuerui Yang, Gang Yu, Xiangyu Zhang, Daxin Jiang

机构 * StepFun(阶跃星辰) NTU(南洋理工大学) PKU(北京大学) UNSW(新南威尔士大学) SJTU(上海交通大学) USTC(中国科学技术大学)

专题命中 长上下文与记忆 :LLM(title,abstract);language model(abstract)

AI总结 ParaASR是一款基于大语言模型的ASR系统,通过多令牌预测技术,在保持低错误率、低延迟的同时,支持32K长上下文及30分钟音频的快速转录,兼顾识别质量、速度与上下文长度。

Comments 14 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25820 2026-07-29 cs.CV 新提交 82%

Food Image Segmentation with LLM-Derived Ingredient Labels and Multimodal Fusion

基于大语言模型衍生成分标签和多模态融合的食品图像分割

Jui-Feng Chi, Wei-Ta Chu, Sheng-Long Lin

机构 * National Cheng Kung University(国立成功大学)

专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract)

AI总结 针对现有食品图像分割模型在相似成分和罕见类别上表现不佳的问题,提出两个多模态模块LIM-F和LIM-Q,利用大语言模型衍生标签提升分割性能,在FoodSeg103基准测试中取得领先,且内存消耗增加适度,为细粒度食品理解提供实用方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21957 2026-07-27 cs.SE cs.CR 新提交 82%

KaPilot: LLM-Assisted Generation of Kani Specifications for Unsafe Rust Verification

KaPilot:用于不安全Rust验证的大语言模型辅助Kani规范生成

Minghua Wang, Yuxi Ling, Mingzhi Gao, Yuwei Liu, Lin Huang

专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract)

AI总结 研究针对不安全Rust内存安全验证规范编写难题,提出KaPilot多智能体框架,经轻量级分析、智能体协作及循环优化、策略筛选确定最佳规范,评估显示其在规范生成成功率和质量上优于AutoSpec。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01331 2026-07-14 cs.CL cs.AI cs.LG 版本更新 82%

MetaState: Persistent Working Memory Enhances Reasoning in Discrete Diffusion Language Models

MetaState: 持久工作记忆增强离散扩散语言模型的推理能力

Kejing Xia, Mingzhe Li, Lixuan Wei, Zhenbang Du, Xiangchi Yuan, Dachuan Shi, Qirui Jin, Wenke Lee

机构 * Georgia Institute of Technology(佐治亚理工学院) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Harvard University(哈佛大学)

专题命中 长上下文与记忆 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 MetaState通过引入轻量级循环增强模块,为冻结的离散扩散语言模型提供持久固定大小的工作记忆,提升推理性能,平均提升4.5%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27243 2026-07-07 cs.CV 版本更新 82%

Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models

检索头能看见图像吗?长上下文视觉语言模型中的多模态检索头

Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao, Yiming Du, Haobo Li, Xiyu Ren, Ginny Wong, Simon See, Lishu Luo, Haodong Duan, Pasquale Minervini, Yangqiu Song

机构 * HKUST(香港科技大学) University of Edinburgh(爱丁堡大学) CUHK(香港中文大学) NVAITC, NVIDIA, Santa Clara, USA(NVIDIA Santa Clara 分公司) Tsinghua University(清华大学)

专题命中 长上下文与记忆 :language model(title,abstract);large language model(abstract)

AI总结 本文提出一种多模态检索头检测方法,发现视觉语言模型中仅有4.4-10.2%的注意力头贡献了50%的正检索分数,这些头对长上下文推理至关重要,且可直接用于文档检索提升性能。

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00605 2026-07-02 cs.CL cs.AI cs.LG 新提交 82%

Auditing Forgetting in Limited Memory Language Models

审计有限记忆语言模型中的遗忘

Arya Raeesi, Hanna Roed

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 长上下文与记忆 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出因果审计框架,通过改变推理时数据库状态分解删除后行为,发现参数泄漏接近零,残留主要来自近邻检索,遗忘边界主要由数据库管理员决定。

Comments 17 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23670 2026-06-23 cs.LG cs.AI cs.CL 新提交 82%

Tapered Language Models

锥形语言模型

Reza Bayat, Ali Behrouz, Aaron Courville

机构 * Mila Cornell University(康奈尔大学) Université de Montréal(蒙特利尔大学) CIFAR AI Chair(CIFAR人工智能教席)

专题命中 长上下文与记忆 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 针对现有语言模型各层参数均匀分配的问题,提出锥形语言模型(TLM),在固定预算下单调递减层容量,实验表明MLP宽度按余弦调度锥形化可提升困惑度和下游性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11975 2026-06-23 cs.RO 版本更新 82%

M2HRI: An LLM-Driven Multimodal Multi-Agent Framework for Personalized Human-Robot Interaction

M2HRI:一种LLM驱动的多模态多智能体框架,用于个性化人机交互

Shaid Hasan, Breenice Lee, Sujan Sarker, Tariq Iqbal

机构 * University of Virginia(弗吉尼亚大学)

专题命中 长上下文与记忆 :LLM(title,title_cn)

AI总结 提出M2HRI框架,通过个性、长期记忆和情境化协调机制赋予每个机器人独特身份,实验表明该方法能提升交互自然性和协调性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16569 2026-06-16 cs.CV cs.RO 新提交 82%

PROSE: Training-Free Egocentric Scene Registration with Vision-Language Models

PROSE: 基于视觉语言模型的无训练自我中心场景配准

Zhiang Chen, Nahyuk Lee, Boyang Sun, Taein Kwon, Marc Pollefeys, Zuria Bauer, Sunghwan Hong

机构 * ETH Zurich(苏黎世联邦理工学院) VGG, University of Oxford(牛津大学VGG实验室) ETH AI Center(苏黎世联邦理工学院人工智能中心)

专题命中 长上下文与记忆 :language model(title,abstract);foundation model(abstract)

AI总结 提出PROSE方法,利用预训练视觉语言模型将RGB序列提升为对象级3D场景图,通过对象高度先验和相同/不同查询匹配实例,无需训练或深度传感器即可实现自我中心场景配准,在Aria基准上超越几何和场景图基线。

Comments Project page: https://rckola.github.io/prose/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12290 2026-06-11 cs.CR 新提交 82%

Selection Integrity for LLM Graph Memory: An Accumulability Criterion for Information-Flow-Blind Retrieval

LLM图记忆的选择完整性:面向信息流盲检索的可累积性准则

Zeming Fei, Hongming Fei, Xiaoyang Wang, Yang yang, Prosanta Gope, Biplab Sikdar, Ying Zhang

专题命中 长上下文与记忆 :LLM(title,title_cn)

AI总结 针对图记忆检索中信息流控制盲区,提出可累积性准则,证明无源结构写入可导致不可逆转账被误导,并通过重分配性而非依赖性预测漏洞,提出认证子图重计算防御。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01782 2026-06-02 cs.IR 82%

Whole-Pool Setwise Reranking with Long-Context Language Models

基于长上下文语言模型的整池集合重排序

Hang Li, Chuting Yu, Teerapong Leelanupab, Bevan Koopman, Guido Zuccon

专题命中 长上下文与记忆 :language model(title);LLM(abstract,abstract_cn)

AI总结 提出整池集合重排序方法,利用长上下文语言模型一次性处理所有候选段落,并通过DualEnd方法从两端同时排序,减少模型调用次数,提升效率。

Comments 4 pages main content, 10 page Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27810 2026-05-28 cs.IR 82%

LRanker: LLM Ranker for Massive Candidates

LRanker: 面向海量候选集的大语言模型排序器

Tao Feng, Zijie Lei, Zhigang Hua, Yan Xie, Shuang Yang, Ge Liu, Jiaxuan You

专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract)

AI总结 提出LRanker框架,通过候选聚合编码器和基于图的测试时缩放机制,解决大语言模型在海量候选集排序中受限于上下文长度和高计算成本的问题,在多个场景下实现显著性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24060 2026-05-26 cs.IR 82%

Same Ranking, Different Winner: How Scoring Targets Shape LLM Memory Benchmarks

相同排名,不同赢家:评分目标如何塑造LLM记忆基准

Sugam Panthi, Rabab Abdelfattah

专题命中 长上下文与记忆 :LLM(title,title_cn)

AI总结 本文揭示对话记忆评估中评分目标选择(原始、来源、规范)的隐含性,通过TIAP审计发现目标切换会显著改变基准结论,并建议明确报告评分目标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13466 2026-05-20 cs.CL cs.AI cs.LG 82%

Language Model Memory and Memory Models for Language

语言模型记忆与记忆模型用于语言

Benjamin L. Badger

机构 * IBM(IBM公司)

专题命中 长上下文与记忆 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究探讨了语言模型和记忆模型在信息存储中的能力差异,发现语言模型的嵌入向量信息较少,而自编码器在输入再生训练中能形成接近完美的记忆,提出了一种可并行的编码器-解码器记忆模型架构,并通过结合因果和信息保留目标函数来提升记忆形成和解码能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19684 2026-04-29 cs.CR 82%

LLM-Assisted Authentication and Fraud Detection

基于大语言模型的认证与欺诈检测

Emunah S-S. Chan, Aldar C-F. Chan

专题命中 长上下文与记忆 :LLM(title,abstract)

AI总结 本文提出基于大语言模型的认证机制和欺诈检测流水线,通过语义判断和文档分段提升认证准确率,减少欺诈误报,验证了大语言模型在安全流程中的应用价值。

Comments 20 pages, 7 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10161 2026-04-14 cs.SD 82%

From Speech to Profile: A Protocol-Driven LLM Agent for Psychological Profile Generation

从语音到档案:一种基于协议的LLM代理用于心理档案生成

Xingjian Yang, Yudong Yang, Zhixing Guo, Yongjie Zhou, Nan Yan, Lan Wang

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) Key Laboratory of Biomedical Imaging Science and System, Chinese Academy of Sciences(中国科学院生物医学成像科学与系统重点实验室)

专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract)

AI总结 本文提出StreamProfile框架,通过增量处理咨询语音,提取证据并生成可追溯的心理档案,有效避免幻觉和长上下文遗忘问题。

详情

展开后加载摘要…

URL PDF HTML 收藏