arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 4789 信号源:cs.CL, cs.AI, cs.LG

1. 长上下文与记忆 4789 篇

2402.19218 2024-03-01 cs.CL 70%

Memory-Augmented Generative Adversarial Transformers

Stephan Raaijmakers, Roos Bakker, Anita Cremers, Roy de Kleijn, Tom Kouwenhoven, Tessa Verhoef

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11681 2024-02-20 cs.CL cs.NA math.NA 70%

Opening the black box of language acquisition

Jérôme Michaud, Anna Jon-and

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14578 2024-01-23 cs.CL 70%

Connecting the Dots: What Graph-Based Text Representations Work Best for Text Classification Using Graph Neural Networks?

Margarita Bugueño, Gerard de Melo

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL

Comments Accepted to Findings of the Association for Computational Linguistics: EMNLP 2023 (Long Paper). 17 pages, 2 figures, 15 tables. The Appendix starts on page 12

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2023, pages 8943-8960

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08908 2024-01-18 cs.OS cs.LG 70%

Herding LLaMaS: Using LLMs as an OS Module

Aditya K Kamath, Sujay Yadalam

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.LG

Comments ASPLOS 2023, Wild and Crazy Ideas session

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17653 2024-01-01 cs.AI 70%

LARP: Language-Agent Role Play for Open-World Games

Ming Yan, Ruihao Li, Hao Zhang, Hao Wang, Zhilan Yang, Ji Yan

专题命中 长上下文与记忆 :language model(abstract);language agent(abstract);分类 cs.AI

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05417 2023-12-12 cs.IR cs.LG 70%

ESPN: Memory-Efficient Multi-Vector Information Retrieval

Susav Shrestha, Narasimha Reddy, Zongwang Li

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.LG

Comments 10 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16135 2023-10-26 cs.CL 70%

Can You Follow Me? Testing Situational Understanding in ChatGPT

Chenghao Yang, Allyson Ettinger

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL

Comments EMNLP 2023 Main Paper (Camera Ready)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12503 2023-08-29 cs.AI cs.HC cs.MA 70%

CGMI: Configurable General Multi-Agent Interaction Framework

Shi Jinxin, Zhao Jiabao, Wang Yilei, Wu Xingjiao, Li Jiawen, He Liang

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.AI

Comments 11 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02047 2023-08-08 cs.CL 70%

CAME: Confidence-guided Adaptive Memory Efficient Optimization

Yang Luo, Xiaozhe Ren, Zangwei Zheng, Zhuo Jiang, Xin Jiang, Yang You

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL

Comments Accepted by ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01169 2023-06-05 cs.CL 70%

Hybrid Long Document Summarization using C2F-FAR and ChatGPT: A Practical Study

Guang Lu, Sylvia B. Larcher, Tu Tran

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01070 2023-06-05 cs.LG 70%

Hierarchical Attention Encoder Decoder

Asier Mujika

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10436 2023-05-19 cs.CL 70%

SmartPhone: Exploring Keyword Mnemonic with Auto-generated Verbal and Visual Cues

Jaewook Lee, Andrew Lan

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL

Comments The 24th International Conference on Artificial Intelligence in Education (AIED 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03688 2023-05-18 cs.CL 70%

DAMO-NLP at SemEval-2023 Task 2: A Unified Retrieval-augmented System for Multilingual Named Entity Recognition

Zeqi Tan, Shen Huang, Zixia Jia, Jiong Cai, Yinghui Li, Weiming Lu, Yueting Zhuang, Kewei Tu, Pengjun Xie, Fei Huang, Yong Jiang

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL

Comments Accepted to SemEval 2023, winners for 9 out of 13 tracks, performance beyond ChatGPT

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03319 2023-05-16 cs.CL 70%

HiPool: Modeling Long Documents Using Graph Neural Networks

Irene Li, Aosong Feng, Dragomir Radev, Rex Ying

专题命中 长上下文与记忆 :language model(abstract);pretraining(abstract);分类 cs.CL

Journal ref ACL 2023 main proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.13246 2023-02-01 cs.SE cs.LG 70%

Conversational Automated Program Repair

Chunqiu Steven Xia, Lingming Zhang

专题命中 长上下文与记忆 :LLM(abstract);language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12687 2022-09-27 cs.CL cs.CY 70%

A Case Report On The "A.I. Locked-In Problem": social concerns with modern NLP

Yoshija Walter

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.00980 2022-07-20 cs.LG stat.ML 70%

Robust Training of Neural Networks Using Scale Invariant Architectures

Zhiyuan Li, Srinadh Bhojanapalli, Manzil Zaheer, Sashank J. Reddi, Sanjiv Kumar

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.LG

Comments 36 pages, 7 figures; ICML 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.03474 2021-07-09 cs.LG 70%

Differentiable Random Access Memory using Lattices

Adam P. Goucher, Rajan Troll

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.LG

Comments 11 pages, 3 figures, submitted to NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.15688 2021-05-25 cs.CL 70%

ERNIE-Doc: A Retrospective Long-Document Modeling Transformer

Siyu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang

专题命中 长上下文与记忆 :language model(abstract);pretraining(abstract);分类 cs.CL

Comments Accepted by ACL 2021 (main conference, long paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19181 2026-08-20 cs.LG cs.AI cs.CL 新提交 67%

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

超越教师似然:面向长上下文推理的组校准在线策略蒸馏

Zhu Zhang, Jixun Wang, Xiaoang Xu, Xiaorong Wang, Zihan Zhou, Zhiyuan Wang, Shuo Wang, Chaojun Xiao, Yuezhi Zhou

机构 * Tsinghua University(清华大学) Beijing University of Posts and Telecommunications(北京邮电大学) OpenBMB

专题命中 长上下文与记忆 :post-training(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 该研究针对长上下文推理中 OPD 的教师-验证器分歧问题,提出 GC-OPD 方法,结合验证器结果优化 OPD,在五个长上下文基准上显著提升 Qwen3 模型性能。

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19059 2026-08-20 cs.RO cs.CV 新提交 67%

LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene Understanding

LT-Mem:面向终身场景理解的感知时间动态的时空记忆

Yumin Lee, Hyoseok Ju, Giseop Kim

机构 * DGIST(大邱庆北科学技术院)

专题命中 长上下文与记忆 :LLM(abstract,abstract_cn)

AI总结 该研究针对机器人长期场景理解的时间遗忘问题,提出LT-Mem时空记忆框架,结合多会话SLAM与Tri-Memory结构,在LT-VQA数据集上性能优于基线且token消耗更少。

Comments 8 pages, 8 figures, 6 tables. Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04707 2026-08-20 cs.CR cs.AI cs.CL cs.LG 版本更新 67%

Jailbreaking in the Haystack

干草堆中的越狱攻击

Rishi Rajesh Shah, Chen Henry Wu, Shashwat Saxena, Ziqian Zhong, Alexander Robey, Aditi Raghunathan

专题命中 长上下文与记忆 :language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 该研究针对长上下文语言模型,提出 NINJA 越狱攻击方法,利用有害目标位置的关键作用,在 HarmBench 基准上提升了多款先进模型的攻击成功率,且该方法低资源、可迁移、难检测、计算最优,揭示了长上下文可引入模型安全漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16944 2026-08-19 cs.DS 新提交 67%

A Tight Linear Deterministic Competitive Ratio for Fully Online KV-Cache Scheduling

全在线KV缓存调度的紧线性确定性竞争比

Ian D'Ambrosio

专题命中 长上下文与记忆 :LLM(abstract,abstract_cn)

AI总结 该研究针对全在线KV缓存调度问题,通过构造困难实例得到确定性竞争比下界,利用均匀因果串行策略得到上界,在Lean 4中完成机器验证,证明其竞争比为紧线性的Θ(n)。

Comments 6 pages. The exact fully online model, fixed-memory-before-scheduler quantifier order, serial upper bound, and wide-short lower bound are checked in Lean 4. A separate reproducibility archive contains pinned-source bootstraps, exact finite controls, formal proofs, tests, and canonical SHA-256 manifests

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16844 2026-08-18 cs.LG cs.AI cs.CL 新提交 67%

Proteus: Incremental Memory Activation for Long-Context Sequence Modeling

Proteus:面向长上下文序列建模的增量式内存激活机制

Reza Bayat, Ali Behrouz, Vahab Mirrokni, Aaron Courville

机构 * Mila(米拉计算科学研究所) Google(谷歌公司)

专题命中 长上下文与记忆 :language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 Proteus 是一种增量式内存激活机制,可嵌入多种神经内存架构,应用于 SWLA 等模型后,能在长上下文相关任务上实现性能提升,证明调度有效容量对序列建模的适用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16551 2026-08-18 cs.CR 新提交 67%

What to Remember, What to Reveal: Privacy-Aware Memory for Conversational Agents

该记住什么,该透露什么:面向对话智能体的隐私感知记忆

Wenjie Wang, Wenhe Si, Xinyue Xu, Yue Xu

专题命中 长上下文与记忆 :LLM(abstract,abstract_cn)

AI总结 该研究针对对话智能体的隐私风险,提出SP-Mem隐私感知记忆架构,通过全生命周期隐私设计实现更强个性化并减少不必要隐私暴露,还构建了相关基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13578 2026-08-17 cs.CL cs.AI cs.LG 新提交 67%

BCMT: Blockwise Causal Memory Transformer

BCMT:分块因果记忆Transformer

Rachid Arezki

专题命中 长上下文与记忆 :language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 该研究提出BCMT架构,通过分块局部自注意力与指数因果记忆解耦长程依赖建模,在1024token上下文的语言建模中,性能与密集Transformer相当,且训练吞吐量更高、内存消耗更低。

Comments 19 pages. Official implementation: https://github.com/rachidlabs/BCMT

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13024 2026-08-17 cs.CE 版本更新 67%

TIEM: Temporal Integration of Hypergraph Evidence and Skill Memory for Event-Driven Financial Forecasting

TIEM:用于事件驱动金融预测的超图证据与技能记忆的时间整合框架

Wenjin Liu, Shen Pang, Chenxi Wang, Tiesunlong Shen, Zhe Cui, Xiaobao Wu, Anh Tuan Luu, Haoran Luo

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract)

AI总结 该研究针对事件驱动金融预测中大型语言模型智能体存在的证据鸿沟问题,提出带时间戳门控的TIEM框架,引入FinPURE基准,经多组金融基准验证其性能优于现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09181 2026-08-11 cs.SE cs.CR 新提交 67%

Memoir: Learning, Verifying, and Evolving False-Positive Memories for Static Application Security Testing Tools

Memoir:针对静态应用安全测试工具的误报记忆学习、验证与演化

Shenyuan Guan, Qiaodan Hou, Yanjun Chen, Xincheng Wen, Jia Feng, Keke Lian, Cuiyun Gao

专题命中 长上下文与记忆 :LLM(abstract,abstract_cn)

AI总结 本研究针对SAST工具的误报问题,提出Memoir框架,通过构建与演化语义记忆实现误报识别,在CWE-Bench-Java上表现优异,且记忆库可跨工具泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08995 2026-08-11 cs.MA 新提交 67%

Muscle Memory for Agents: Compile not Merely Retrieve

智能体的肌肉记忆:编译而非仅检索

Pouya Ghiasnezhad Omran, Soujanya Lanka, Qin Zhang, Tanya Dixit

专题命中 长上下文与记忆 :LLM(abstract,abstract_cn)

AI总结 本文提出将重复用户意图编译为专用专家智能体的肌肉记忆范式,通过四阶段流水线实现,在90个场景中专家触发时胜率达88.9%,优于传统检索模式,为智能体记忆设计提供新方向。

Comments 12 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08075 2026-08-11 cs.IR cs.CV cs.MM 新提交 67%

Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence

视觉世界的搜索:持久视觉记忆、分层索引与基于源的证据

Sankalp Nagaonkar, Rohit Garg, Ankit Raj, Ashish Choithani, Ashutosh Trivedi

专题命中 长上下文与记忆 :foundation model(abstract);pretraining(abstract)

AI总结 该研究针对智能体持续观测视觉数据的场景,提出基于持久视觉记忆等的视频检索基础设施模型,经9800+查询对比,通用组件流水线在Recall@1/@3/@10上优于商业视频原生引擎

Comments 33 pages, 5 figures, 17 tables. Technical report. Benchmark configurations and reproduction instructions: https://github.com/video-db/search-over-the-visual-world

详情

展开后加载摘要…

URL PDF HTML 收藏