arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-04-22 至 2026-04-22 共收录 327 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 74 篇

2512.09427 2026-04-22 cs.AR cs.AI 87%

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators

ODMA:面向LLM服务的LPDDR类加速器按需内存分配策略

Guoqiang Zou, Wanyu Wang, Hao Zheng, Longxiang Yin, Yinhe Han

机构 * University of Chinese Academy of Sciences(中国科学院大学) Beijing Information Science and Technology University(北京信息科技大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 ODMA针对LPDDR类加速器的随机访问带宽限制,提出按需内存分配策略,通过动态调整内存桶边界和安全池提升KV缓存利用率和吞吐量。

Comments 4 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06798 2026-04-22 cs.LG cs.AI 84%

MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization

MoBiE: 一种在后训练量化下高效混合二进制专家的推理方法

Zhixiong Zhao, Zukang Xu, Zhixuan Chen, Dawei Yang

专题命中 效率与部署 :post-training(title);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出MoBiE,一种专为基于混合专家(MoE)的大型语言模型(LLMs)设计的二进制化框架,通过减少交叉专家冗余、增强权重重要性估计和缓解路由扭曲,提升效率与性能。

Comments Although previously revised, per strict university regulations regarding incorrect affiliation, I am unauthorized to retain this manuscript. Furthermore, fundamental derivation errors in the NGES section compromise the mathematical framework, alongside misleading overlapping wording. The paper is therefore withdrawn

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19664 2026-04-22 cs.IR 83%

ECLASS-Augmented Semantic Product Search for Electronic Components

基于ECLASS的语义产品搜索增强

Nico Baumgart, Markus Lange-Hegermann, Jan Henze

专题命中 效率与部署 :LLM(summary_cn,abstract);foundation model(abstract)

AI总结 本文提出利用LLM密集检索和ECLASS层次语义提升电子元件语义搜索效果,实验显示其在准确率和效率上均优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16368 2026-04-22 cs.CL 83%

Cross-Family Speculative Decoding for Polish Language Models on Apple~Silicon: An Empirical Evaluation of Bielik~11B with UAG-Extended MLX-LM

跨家族推测解码用于波兰语言模型在苹果Silicon:对Bielik~11B与扩展MLX-LM的实证评估

Krzysztof Fonal

机构 * Wrocław University of Science and Technology(沃拉夫大学科学与技术学院)

专题命中 效率与部署 :language model(title);LLM(abstract,abstract_cn);分类 cs.CL

AI总结 本文评估了在苹果Silicon上使用扩展MLX-LM框架实现跨 tokenizer 推测解码的有效性,探讨了波兰语言模型的性能差异及统一内存架构下的实证结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04800 2026-04-22 cs.CL 83%

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights

语言模型的混合架构:系统分析与设计洞察

Sangmin Bae, Bilge Acun, Chien-Yu Lin, Haroun Habeeb, Seungyeon Kim, Liang Luo, Junjie Wang, Carole-Jean Wu

机构 * FAIR at Meta(Meta的FAIR) Meta

专题命中 效率与部署 :language model(title,abstract);large language model(abstract);分类 cs.CL

AI总结 本文系统分析了混合架构的设计,探讨了不同融合策略对语言模型性能、长上下文能力及效率的影响,提出优化设计方法。

Comments 41 pages, 8 figures, 22 tables;

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17789 2026-04-22 cs.CV cs.AI cs.CL 82%

DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization

DuQuant++: 细粒度旋转增强微缩放FP4量化

Haokun Lin, Xinle Jia, Haobo Xu, Bingchen Yao, Xianglong Guo, Yichen Wu, Zhichao Lu, Ying Wei, Qingfu Zhang, Zhenan Sun

机构 * CASIA(中国科学院自动化研究所) NJU(南京大学) THU(清华大学) ZJU(浙江大学) Harvard(哈佛大学) CityU(城市大学)

专题命中 效率与部署 :LLM(summary_cn,abstract);分类 cs.CL、cs.AI

AI总结 DuQuant++通过细粒度旋转优化MXFP4微缩放格式,解决激活异常带来的量化误差问题,提升LLM推理效率。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08899 2026-04-22 cs.CL cs.LG 82%

ConFu: Contemplate the Future for Better Speculative Sampling

ConFu:为更好的推测采样展望未来

Zongyue Qin, Raghavv Goel, Mukul Gagrani, Risheek Garrepalli, Mingu Lee, Yizhou Sun

机构 * University of California Los Angeles, United States(美国加州大学洛杉矶分校) Qualcomm AI Research, United States(高通人工智能研究)

专题命中 效率与部署 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 ConFu提出一种新的推测解码框架,通过引入展望令牌和软提示,使草案模型能利用目标模型的未来信号,提升生成速度和接受率。

Comments v3: Added stress test with long drafts (DL=12, top-k=1) and tail-acceptance (survival) analysis. Earlier versions added Qwen3-4B results and ablations

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18909 2026-04-22 cs.AR 82%

ChipLight: Cross-Layer Optimization of Chiplet Design with Optical Interconnects for LLM Training

ChipLight:基于光学互连的芯片片级设计跨层优化用于大语言模型训练

Kangbo Bai, Zhantong Zhu, Yifan Ding, Tianyu Jia

专题命中 效率与部署 :LLM(title,abstract)

AI总结 ChipLight通过跨层多目标设计优化方法,结合芯片片和光学互连技术,提升训练集群的通信效率与性能。

Comments Accepted by DATE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14170 2026-04-22 cs.CL cs.CY 81%

Comparing energy consumption and accuracy in text classification inference

在文本分类推理中比较能耗与准确性

Johannes Zschache, Tilman Hartwig

机构 * Application Lab for AI and Big Data(人工智能与大数据应用实验室)

专题命中 效率与部署 :LLM(summary_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究比较了文本分类推理中模型准确性和能耗的权衡,发现高准确度模型可能也具备能效,LLM在零样本分类中与传统模型有相似或更低的准确率,能耗受模型类型、规模和硬件影响显著。

Comments Key results in Figure 2, accepted in Nature Sci Rep, 32 pages

Journal ref Sci Rep 16, 12717 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18913 2026-04-22 cs.CL 81%

LogosKG: Hardware-Optimized Scalable and Interpretable Knowledge Graph Retrieval

LogosKG:硬件优化的可扩展且可解释的知识图谱检索

He Cheng, Yifu Wu, Saksham Khatwani, Maya Kruse, Dmitriy Dligach, Timothy A. Miller, Majid Afshar, Yanjun Gao

机构 * LARK Lab, University of Colorado Anschutz(洛克拉克实验室,科罗拉多大学安施图茨分校) University of Colorado Boulder(科罗拉多大学波德分校) Loyola University Chicago(芝加哥洛克拉克大学) Harvard Medical School(哈佛医学院) Boston Children’s Hospital(波士顿儿童医院) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 效率与部署 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 LogosKG通过符号知识图谱和硬件高效操作实现大规模知识图谱的多跳检索,提升效率与可解释性,展示出在生物医学知识与大语言模型推理对齐分析中的应用价值。

Comments Accepted to the ACL 2026 Main Conference. 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13327 2026-04-22 cs.DC cs.LG cs.PL 81%

Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel

事件张量:一种统一的编译动态巨核抽象

Hongyi Jin, Bohan Hou, Guanjie Wang, Ruihang Lai, Jinqi Chen, Zihao Ye, Yaxing Cai, Yixin Dong, Xinhao Cheng, Zhihao Zhang, Yilong Zhao, Yingyi Huang, Lijie Yang, Jinchen Jiang, Gabriele Oliaro, Jianan Ji, Xupeng Miao, Vinod Grover, Todd C. Mowry, Zhihao Jia, Tianqi Chen

机构 * Anonymous Authors(匿名作者)

专题命中 效率与部署 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出事件张量,一种统一的编译动态巨核抽象,通过编码 tiled 任务间的依赖关系,支持动态形状和数据依赖的动态性,从而提升大型语言模型服务的性能和减少系统预热开销。

Comments 16 pages. 18 figures. accepted in MLSys 2026. References corrected

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03261 2026-04-22 cs.CL cs.CY cs.HC 81%

VIGIL: An Extensible System for Real-Time Detection and Mitigation of Cognitive Bias Triggers

VIGIL:一个用于实时检测和缓解认知偏差触发的可扩展系统

Bo Kang, Sander Noels, Tijl De Bie

机构 * Ghent University(根特大学)

专题命中 效率与部署 :LLM(summary_cn,abstract);分类 cs.CL

AI总结 VIGIL通过浏览器扩展实时检测和缓解在线信息中的认知偏差触发,结合LLM重述和隐私分层推理,提供可扩展的插件系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23281 2026-04-22 cs.CL 81%

MCP vs RAG vs NLWeb vs HTML: A Comparison of the Effectiveness and Efficiency of Different Agent Interfaces to the Web (Technical Report)

MCP与RAG与NLWeb与HTML:不同网页代理接口的有效性和效率比较(技术报告)

Aaron Steiner, Ralph Peeters, Christian Bizer

机构 * Data and Web Science Group(数据与网络科学组) University of Mannheim(曼海姆大学)

专题命中 效率与部署 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文比较了四种网页代理接口的有效性和效率,发现RAG、MCP和NLWeb在效果和效率上均优于HTML。

Journal ref Proceedings of the ACM Web Conference 2026 (WWW '26), April 13-17, 2026, Dubai, United Arab Emirates. ACM, New York, NY, USA, pp. 8493-8496

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08240 2026-04-22 cs.CL 81%

The Alignment Waltz: Jointly Training Agents to Collaborate for Safety

对齐之舞:联合训练代理以安全协作

Jingyu Zhang, Haozhu Wang, Eric Michael Smith, Sid Wang, Amr Sharaf, Mahesh Pasupuleti, Benjamin Van Durme, Daniel Khashabi, Jason Weston, Hongyuan Zhan

机构 * Meta Superintelligence Labs(Meta超智能实验室) Johns Hopkins University(约翰·霍普金斯大学)

专题命中 效率与部署 :LLM(summary_cn,abstract);分类 cs.CL

AI总结 本文提出WaltzRL框架,通过联合训练对话代理和反馈代理,解决LLM在安全与帮助性间的平衡问题,实验表明其有效降低不安全响应和过度拒绝率。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03828 2026-04-22 cs.LG stat.ML 81%

IMPACT: Importance-Aware Activation Space Reconstruction

IMPACT: 基于重要性的激活空间重建

Md Mokarram Chowdhury, Daniel Agyei Asante, Ernie Chang, Yang Li

机构 * Department of Computer Science, Iowa State University, United States(爱荷华州立大学计算机科学系) Meta, United States(Meta公司)

专题命中 效率与部署 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出IMPACT框架,通过结合激活结构与梯度重要性,实现低秩压缩优化,实验显示在保持准确性的同时,模型大小可减少55.4%。

Comments To appear in the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18834 2026-04-22 cs.SE cs.SY eess.SY 80%

Structural Verification for Reliable EDA Code Generation without Tool-in-the-Loop Debugging

结构验证用于无需工具在回路调试的可靠EDA代码生成

Dinithi Jayasuriya, Aravind Saravanan, Nilesh Ahuja, Amanda Rios, Amit Trivedi

专题命中 效率与部署 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文提出通过强制执行结构正确性来提高EDA代码生成的可靠性与效率,采用结构依赖图作为显式执行合同,并通过验证引导合成框架进行图条件检索、约束生成和分阶段预执行验证,从而在单步和多步任务中显著提升通过率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18614 2026-04-22 cs.DC cs.CR cs.ET cs.MA 80%

HadAgent: Harness-Aware Decentralized Agentic AI Serving with Proof-of-Inference Blockchain Consensus

HadAgent:基于证明-推理区块链共识的去中心化代理AI服务

Landy Jimenez, Mariah Weatherspoon, Bingyu Shen, Yi Sheng, Jianming Liu, Boyang Li

专题命中 效率与部署 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 HadAgent通过引入证明-推理共识机制,替代传统工作量证明,实现去中心化代理AI服务,提升验证效率与安全性,实验显示高检测率与低误报率。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19642 2026-04-22 cs.CL 79%

Micro Language Models Enable Instant Responses

微语言模型实现即时响应

Wen Cheng, Tuochao Chen, Karim Helwani, Sriram Srinivasan, Luke Zettlemoyer, Shyamnath Gollakota

机构 * Paul G. Allen School of Computer Science & Engineering, University of Washington(保罗·G·艾伦计算机科学与工程学院,华盛顿大学)

专题命中 效率与部署 :language model(title,abstract);分类 cs.CL

AI总结 微语言模型通过在设备端生成前4-8个词,结合云端完成响应,克服了边缘设备计算和功耗限制,实现低延迟的助手体验。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19145 2026-04-22 cs.CV cs.AI 79%

ST-Prune: Training-Free Spatio-Temporal Token Pruning for Vision-Language Models in Autonomous Driving

ST-Prune:面向自动驾驶的视觉-语言模型中无训练的时空令牌修剪

Lin Sha, Haiyun Guo, Tao Wang, Cong Zhang, Min Huang, Jinqiao Wang, Qinghai Miao

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Carizon Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 效率与部署 :language model(title,abstract);分类 cs.AI

AI总结 针对自动驾驶中多视角相机和多帧视频输入的计算开销问题,ST-Prune提出无训练的时空令牌修剪框架,通过MTP和RSP模块有效压缩时空冗余,实现90%令牌减少下的近无损性能。

Comments 18 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18831 2026-04-22 cs.CV cs.RO 78%

Feasibility of Indoor Frame-Wise Lidar Semantic Segmentation via Distillation from Visual Foundation Model

基于视觉基础模型的知识蒸馏的室内点云语义分割可行性

Haiyang Wu, Juan J. Gonzales Torres, George Vosselman, Ville Lehtola

机构 * 1 Department of Earth Observation Science, Faculty of Geo-Information Science Earth Observation (ITC), University of Twente, 7522 NB Enschede, The Netherlands

专题命中 效率与部署 :foundation model(title,abstract)

AI总结 本文探讨了通过知识蒸馏从视觉基础模型中学习的方法,用于室内点云语义分割,展示了在无人工标注的情况下,该方法在室内场景中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19083 2026-04-22 cs.CR cs.AI 77%

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety

ProjLens: 揭示项目器在多模态模型安全中的作用

Kun Wang, Cheng Qian, Miao Yu, Lilan Peng, Liang Lin, Jiaming Zhang, Tianyu Zhang, Yu Cheng, Yang Wang

机构 * University of Science and Technology of China(中国科学技术大学) Beijing University of Aeronautics and Astronautics(北京航空航天大学) Nanyang Technological University(南洋理工大学) Southwest Jiaotong University(西南交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 效率与部署 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 ProjLens通过分析多模态大语言模型中的后门攻击机制,揭示了项目器在安全漏洞中的关键作用,发现后门注入参数编码于低秩子空间,并通过实验验证了激活机制的差异。

Comments 18 pages ,15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05608 2026-04-22 cs.CL 77%

A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Tasks

没有计划的愿望只是愿望:为长周期智能体任务高效且有效的全局规划器训练

Shuzheng Si, Haozhe Zhao, Kangyang Luo, Gang Chen, Fanchao Qi, Minjia Zhang, Baobao Chang, Maosong Sun

机构 * Tsinghua University(清华大学) Peking University(北京大学) DeepLang AI University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出EAGLET方法,通过两步流程训练高效规划器,提升智能体规划能力,实验显示其在长周期任务中优于现有方法,且训练成本降低8倍。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22811 2026-04-22 stat.ML cs.LG 77%

Highly Efficient and Effective LLMs with Multi-Boolean Architectures

高效且有效的多布尔架构大语言模型

Ba-Hien Tran, Van Minh Nguyen

机构 * Huawei Paris Research Center(华为巴黎研究中心)

专题命中 效率与部署 :large language model(abstract);language model(abstract);post-training(abstract);分类 cs.LG

AI总结 本文提出多布尔架构框架,通过布尔参数直接微调大语言模型,提升表示能力并降低复杂度,实验显示优于低比特量化和二值化方法。

Comments ICLR 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18005 2026-04-22 cs.MA cs.AI cs.CL 76%

Diversity Collapse in Multi-Agent LLM Systems: Structural Coupling and Collective Failure in Open-Ended Idea Generation

多智能体大语言模型系统中的多样性崩溃:结构耦合与开放性创意生成中的集体失败

Nuo Chen, Yicheng Tong, Yuzhe Yang, Yufei He, Xueyi Zhang, Qingyun Zou, Qian Wang, Bingsheng He

机构 * National University of Singapore(国立新加坡大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 效率与部署 :LLM(title);分类 cs.CL、cs.AI

AI总结 研究探讨了多智能体系统在开放性创意生成中多样性崩溃的原因,揭示了结构耦合导致的集体失败,强调了在设计创意任务时保持独立性和分歧的重要性。

Comments 56 pages, 15 figures; Accepted at ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19565 2026-04-22 cs.CL cs.AI cs.LG 75%

Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps

通过注意力图在推理时间检测语音大语言模型中的幻觉

Jonas Waldendorf, Bashar Awwad Shiekh Hasan, Evgenii Tsymbalov

机构 * University of Edinburgh(爱丁堡大学) Amazon AGI(亚马逊人工智慧实验室)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出四种基于注意力的指标,用于检测语音大语言模型中的幻觉,通过轻量级逻辑回归分类器实现高效推理检测,在多个任务中优于基线方法,展示了注意力模式在检测幻觉中的价值。

Comments Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23439 2026-04-22 cs.CL cs.AI cs.LG cs.SD eess.AS 75%

Speculative End-Turn Detector for Efficient Speech Chatbot Assistant

推测性终点检测器用于高效的语音聊天机器人助手

Hyunjong Ok, Suho Yoo, Jaeho Lee

机构 * POSTECH(POSTECH大学) HJ AILAB(HJ人工智能实验室) KAIST(韩国科学技术院)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出SpeculativeETD,通过轻量GRU模型和高性能Wav2vec模型结合,提升资源受限环境下的终点检测效率与准确性。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18756 2026-04-22 cs.LG cs.AI cs.CL cs.CR 75%

Towards Understanding the Robustness of Sparse Autoencoders

迈向稀疏自编码器鲁棒性的理解

Ahson Saiyed, Sabrina Sadiekh, Chirag Agarwal

机构 * University of Virginia(弗吉尼亚大学) Independent Researcher(独立研究者)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究通过在推理阶段集成预训练的稀疏自编码器,发现其能显著降低对抗攻击成功率,并揭示稀疏性与攻击成功率之间的剂量-反应关系,以及层间防御与性能的权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13654 2026-04-22 cs.LG cs.AI cs.CL 75%

Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency

输入时间扩展:添加噪声和无关信息显著提升推理性能和效率

Rapheal Huang, Weilong Guo

机构 * Independent Researcher(独立研究者)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 通过引入噪声和无关信息,输入时间扩展方法在推理性能和效率上取得突破,无需额外设计即可提升效果,适用于更广泛场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16535 2026-04-22 cs.LG cs.AI 73%

SCATR: Simple Calibrated Test-Time Ranking

SCATR: 简单校准的测试时间排名

Divya Shyamal, Marta Knežević, Lan Tran, Chanakya Ekbote, Vijay Lingam, Paul Pu Liang

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 SCATR通过利用基础模型的隐藏表示,从小型校准集学习轻量级评分器,提升了测试时间排名的效率和准确性,在多个基准测试中优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01455 2026-04-22 cs.CV cs.AI cs.CL cs.IR cs.MM 73%

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents

从逐字到概要:基于语义信息瓶颈的金字塔多模态记忆压缩用于长视界视频智能体

Niu Lian, Yuting Wang, Hanshu Yao, Jinpeng Wang, Bin Chen, Yaowei Wang, Min Zhang, Shu-Tao Xia

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳校区) Peng Cheng Laboratory(鹏城实验室)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MM-Mem架构,通过金字塔多模态记忆压缩,结合语义信息瓶颈目标,实现长视界视频理解中的高效记忆组织与任务相关信息保留。

Comments Accepted by ACL 2026 Main. 17 pages, 7 figures, 8 tables. TL;DR: We propose MM-Mem, a cognition-inspired, dual-trace hierarchical memory framework for long-horizon video understanding grounded in Fuzzy-Trace Theory. It features adaptive memory compression via the Information Bottleneck and employs an entropy-driven top-down retrieval to access fine-grained details only when necessary

详情

展开后加载摘要…

URL PDF HTML 收藏