arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-04-29 至 2026-04-29 共收录 226 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 65 篇

2603.29844 2026-04-29 cs.RO cs.AI cs.CV cs.LG 62%

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

DIAL: 通过潜在世界建模解耦意图与动作以实现端到端VLA

Yi Chen, Yuying Ge, Hui Zhou, Mingyu Ding, Yixiao Ge, Xihui Liu

机构 * The University of Hong Kong(香港大学) XPENG Robotics(小鹏机器人) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 评测与基准 :language model(abstract);分类 cs.AI、cs.LG

AI总结 DIAL通过潜在意图瓶颈解耦意图与动作,利用VLM进行潜在世界建模并结合轻量策略实现端到端VLA,实验表明其在RoboCasa GR1任务中优于现有方法,且在真实世界部署中表现稳健。

Comments Project page: https://xpeng-robotics.github.io/dial

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25584 2026-04-29 cs.AI 57%

DualFact+: A Multimodal Fact Verification Framework for Procedural Video Understanding

DualFact+: 一种用于过程视频理解的多模态事实验证框架

Cennet Oguz, Yasser Hamidullah, Josef van Genabith, Simon Ostermann

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) Saarland Informatics Campus(萨尔兰信息学校区)

专题命中 评测与基准 :language model(abstract);分类 cs.AI

AI总结 DualFact+通过双层多模事实验证框架,针对过程视频描述中的概念事实和上下文事实进行评估,揭示了多模态事实 grounding 的挑战。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24842 2026-04-29 cs.AI cs.MA cs.MM 57%

Co-Director: Agentic Generative Video Storytelling

Co-Director: 基于代理的生成视频叙事

Yale Song, Yiwen Song, Nick Losier, Nathan Hodson, Ye Jin, Rhyard Zhu, Yan Xu, Daniel Vlasic, Carina Claassen, Jasmine Leon, Khanh G. LeViet, Zack Chomyn, Joe Timmons, Brett Slatkin, Scott Penberthy, Tomas Pfister

机构 * Google(谷歌)

专题命中 评测与基准 :prompting(abstract);分类 cs.AI

AI总结 本文提出Co-Director框架,通过分层多代理方法解决视频生成的语义一致性问题,引入分层参数化和多模态自优化循环,实现叙事策略探索与有效配置的平衡,通过GenAD-Bench验证其在个性化广告中的优越性。

Comments Project Page: https://co-director-agent.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16877 2026-04-29 cs.CL 57%

Enhancing Financial Report Question-Answering: A Retrieval-Augmented Generation System with Reranking Analysis

增强财务报告问答:一种带有重排序分析的检索增强生成系统

Zhiyuan Cheng, Longying Lai, Yue Liu, Kai Cheng, Xiaoxi Qi

机构 * School of Engineering Stanford University Stanford, CA, USA Simon Business School University of Rochester Rochester, NY, USA Accounting \& Information Systems Rutgers University Newark, NJ, USA Institute for Social Economic Research Policy Columbia University New York, NY, USA Department of Economics Northeastern University Boston, MA, USA

专题命中 评测与基准 :language model(abstract);分类 cs.CL

AI总结 本文提出一种检索增强生成系统,通过重排序提升财务报告问答性能,实验表明重排序显著提高答案质量,正确率提升15.5个百分点。

Comments 7 pages, 2 figures. Accepted to ICECET 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25361 2026-04-29 cs.CV 50%

HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation

HuM-Eval: 一种以人为中心的视频评估框架

Bingzi Zhang, Kaisi Guan, Ruihua Song

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院)

专题命中 评测与基准 :language model(abstract)

AI总结 本文提出HuM-Eval框架,通过粗到细策略评估生成视频中的人体运动质量,实验显示其在人体相关性上达到58.2%的平均值,优于现有方法。

Comments Accepted to the 2026 IEEE International Conference on Multimedia and Expo (ICME 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19016 2026-04-29 cs.HC 50%

CHORUS: Effort-Aware Multi-Agent Human-AI Collaboration for Professional Translation

CHORUS:面向专业翻译的有意识多智能体人机协作

George X. Wang, Jiaqian Hu, Guande Wu Jing Qian

专题命中 评测与基准 :prompting(abstract)

AI总结 CHORUS通过多智能体协作系统支持翻译流程与个人风格,减少译者认知负担,提升翻译质量,降低完成时间33.8%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17492 2026-04-29 cs.CV 50%

MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding

MMLANDMARKS: 一种用于地理空间理解的跨视图实例级基准

Oskar Kristoffersen, Alba Reinders Sánchez, Morten Rieger Hannemose, Anders Bjorholm Dahl, Dim P. Papadopoulos

机构 * Technical University of Denmark(丹麦技术大学) Pioneer Center for AI(先锋人工智能中心)

专题命中 评测与基准 :foundation model(abstract)

AI总结 本文提出MMLandmarks基准,包含四类数据:高分辨率航拍图、地面视图、文本信息和地理坐标,用于跨视图地物检索、定位及文本到图像等任务,揭示多模态数据对地理空间理解的重要性。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 效率与部署 29 篇

2603.15954 2026-04-29 cs.LG cs.AI 90%

MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment

MobileLLM-Flash: 用于工业级部署的延迟引导设备端大语言模型设计

Hanxian Huang, Igor Fedorov, Andrey Gromov, Bernard Beckerman, Naveen Suda, David Eriksson, Maximilian Balandat, Rylan Conway, Patrick Huber, Chinnadhurai Sankar, Ayushi Dalmia, Zechun Liu, Lemeng Wu, Tarek Elgamal, Adithya Sagar, Vikas Chandra, Raghuraman Krishnamoorthi

机构 * Meta AI

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);foundation model(abstract)

AI总结 本文提出MobileLLM-Flash,通过硬件循环架构搜索在移动端延迟约束下设计高效的大语言模型,支持8k上下文长度,实现更快的预填和解码速度。

Comments Accepted to ACL Industry Track 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18030 2026-04-29 cs.CL cs.AI cs.LG 90%

From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models

从局部到全局:重新审视大型语言模型的结构剪枝范式

Ziyan Wang, Enmao Diao, Qi Le, Pu Wang, Minwoo Lee, Shu-ping Yeh, Evgeny Stupachenko, Hao Feng, Li Yang

机构 * University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校) DreamSoul University of Minnesota(明尼苏达大学) Intel Corporation(英特尔公司)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);post-training(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出GISP,一种基于全局重要性度量的迭代结构剪枝方法,通过去除注意力头和MLP通道提升模型效率和下游任务性能。

Comments 20 pages, 6 figures. Accepted by ACL2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25183 2026-04-29 cs.AR 89%

Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference

基于查找表加速器的硬件生成与探索:用于1.58位LLM推理

Robin Geens, Joran Heldens, Joren Dumoulin, Marian Verhelst

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 本文提出一种基于查找表的加速器设计框架,通过系统化分析揭示了三元权重计算的架构权衡,优化设计在面积上比乘法器基线减少了2.2倍,并通过基准测试证明了参数修正可进一步提升性能。

Comments Presented as ISPASS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11786 2026-04-29 cs.LG cs.AI 89%

Evaluating LLM Safety Under Repeated Inference via Accelerated Prompt Stress Testing

通过加速提示压力测试评估LLM在重复推理下的安全性

Keita Broadwater

机构 * Independent Researcher(独立研究者)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出APST框架,通过重复推理测试揭示模型在持续使用下的失败模式,展示单次评估无法反映实际可靠性差异。

Comments 23 pages, 9 figures; editorial and LaTeX revisions for clarity; improved presentation of methodology and results; updated figures, tables, and float placement; clarified temperature sensitivity and deployment-risk analysis; expanded reporting from the same experiments; results unchanged in substance

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17783 2026-04-29 cs.PF cs.AI cs.CL cs.LG 88%

Energy-Aware LLMs: A step towards sustainable AI for downstream applications

面向能源效率的大型语言模型:迈向可持续AI的一步

Nguyen Phuc Tran, Brigitte Jaumard, Oscar Delgado

机构 * Concordia University(康科迪亚大学)

专题命中 效率与部署 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出了一种端到端的管道,研究在通信网络故障工单分析中LLM的能效与性能的权衡,并通过两个真实数据集评估了根因分析和响应反馈任务的性能提升。

Comments This work has been submitted to V. International Conference on Electrical, Computer and Energy Technologies (ICECET 2025) for possible publication

Journal ref 2025 5th International Conference on Electrical, Computer and Energy Technologies (ICECET)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24808 2026-04-29 cs.MA cs.AI cs.CY cs.DC 87%

ITAS: A Multi-Agent Architecture for LLM-Based Intelligent Tutoring

ITAS:基于大语言模型的智能辅导系统多智能体架构

Iizalaarab Elhaimeur, Nikos Chrisochoides

机构 * Center for Real-Time Computing(实时计算中心) Computer Science Department(计算机科学系) Old Dominion University(旧 Dominion 大学) Physics Departments(物理系)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出ITAS多智能体架构,用于解决大语言模型在真实课程中运行的挑战,通过三层架构实现教学、操作和反馈功能,展示了系统在实际应用中的表现。

Comments Companion papers: arXiv:Q-ID (Quantum deployment), arXiv:L-ID (Latency analysis)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17697 2026-04-29 cs.LG cs.SE 87%

Pimp My LLM: Leveraging Variability Modeling to Tune Inference Hyperparameters

优化大语言模型:利用变异性建模来调整推理超参数

Nada Zine, Clément Quinton, Romain Rouvoy

机构 * Univ. Lille, CNRS, Inria, Centrale Lille, UMR 9189 CRIStAL, F-59000, Lille(里尔大学、国家科学研究中心、法国国家信息与自动化技术研究所、中央理工学院、UMR 9189 CRIStAL、法国里尔)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文通过变异性建模系统分析大语言模型推理配置,揭示超参数间的权衡关系,并支持基于有限测量预测推理行为,推动软件工程与机器学习的结合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25317 2026-04-29 cs.AR 87%

FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture

FusionCIM: 通过融合驱动的计算存内架构加速大语言模型推理

Zihao Xuan, Jia Chen, Yewen Li, Wei Xuan, Hegan Chen, Xiao Huo, Fengbin Tu

专题命中 效率与部署 :LLM(title,summary_cn)

AI总结 本文提出FusionCIM架构,通过融合驱动的计算存内架构提升LLM推理效率,核心方法包括混合CIM流水线、QO静态数据流和模式感知在线softmax机制,实现3.86倍能效提升和1.98倍加速。

Comments 7 Pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02452 2026-04-29 cs.CV 85%

Personalization Toolkit: Training Free Personalization of Large Vision Language Models

个性化工具包:无需训练的大型视觉语言模型个性化

Soroush Seifi, Vaggelis Dorovatas, Matteo Cassinelli, Fabien Despinoy, Daniel Olmeda Reino, Rahaf Aljundi

机构 * Toyota Motor Europe(丰田欧洲公司)

专题命中 效率与部署 :language model(title,abstract);foundation model(abstract);prompting(abstract)

AI总结 本文提出无需训练的LVLM个性化方法,通过预训练视觉基础模型提取特征,结合检索增强生成技术实现多概念个性化,适用于图像和视频。

Comments Accepted at Transactions on Machine Learning Research (TMLR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23108 2026-04-29 cs.CL cs.AI cs.LG 85%

Mixture of Heterogeneous Grouped Experts for Language Modeling

异质分组专家的混合模型用于语言建模

Zhicheng Ma, Xiang Liu, Zhaoxiang Liu, Ning Wang, Yi Shen, Kai Wang, Shuming Shi, Shiguo Lian

机构 * Data Science & Artificial Intelligence Research Institute, China Unicom(中国unicom数据科学与人工智能研究院) Unicom Data Intelligence, China Unicom(中国unicom数据智能)

专题命中 效率与部署 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出MoHGE,通过双层路由机制和分组辅助损失优化,实现资源感知的专家组合,减少参数量并平衡GPU负载,提升语言模型的效率与实用性。

Comments Accepted by ACL2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24971 2026-04-29 cs.LG cs.CL cs.DC 84%

PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference

PolyKV:一种用于多智能体LLM推理的共享非对称压缩KV缓存池

Ishan Patel, Ishan Joshi

机构 * Independent Researcher(独立研究者)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.CL、cs.LG

AI总结 PolyKV通过非对称压缩技术实现多个并发推理代理共享单一KV缓存池,提升内存效率并减少推理延迟。

Comments 10 pages, 6 tables. Code: https://github.com/ishan1410/PolyKV Keywords: KV cache compression, multi-agent LLM inference, asymmetric quantization, FWHT, TurboQuant, shared memory

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25903 2026-04-29 cs.SE cs.LG 83%

Carbon-Taxed Transformers: A Green Compression Pipeline for Overgrown Language Models

碳税变压器:用于过度语言模型的绿色压缩流程

Ajmain Inqiad Alam, Palash Roy, Chanchal K. Roy, Banani Roy, Kevin A. Schneider

机构 * University of Saskatchewan(萨斯喀彻温大学)

专题命中 效率与部署 :language model(title,abstract);large language model(abstract);分类 cs.LG

AI总结 本文提出碳税变压器,通过经济碳税原理设计压缩流程,提升软件工程中大语言模型的效率与环保性能,实现内存、时间及碳排放的显著降低,同时保持高精度。

Journal ref Proceedings of ACM Software Engineering 3, FSE, Article FSE047, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02657 2026-04-29 cs.IR cs.CL 83%

Less LLM, More Documents: Searching for Improved RAG

少用大模型,多用文档:寻找改进RAG的方法

Jingjie Ning, Yibo Kong, Yunfan Long, Jamie Callan

机构 * School of Computer Science, Carnegie Mellon University(计算机科学系,卡内基梅隆大学)

专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文研究了扩大检索器文档库对RAG性能的影响,发现增大文档库可提升性能,甚至在某些情况下优于使用更大模型,但存在边际效益递减的问题。

Comments Proceeding Version of ECIR 2026. In: Campos, R., et al. Advances in Information Retrieval. ECIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13766 2026-04-29 cs.SE cs.AI cs.CL 82%

A Blueprint for AI-Driven Software Quality: Integrating LLMs with Established Standards

AI驱动软件质量的蓝图:整合大语言模型与现有标准

Avinash Patil

机构 * Juniper Networks Inc.(Juniper网络公司)

专题命中 效率与部署 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文探讨了利用大语言模型提升软件质量保证的过程,结合现有标准,分析AI在需求验证、缺陷检测等任务中的应用,并提出未来发展方向。

Comments 16 pages, 2 Table, 7 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25080 2026-04-29 cs.DC 82%

CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration

CacheFlow: 高效的大语言模型服务中的3D并行KV缓存恢复

Sean Nian, Jiahao Fang, Qilong Feng, Zhiyu Wu, Fan Lai

专题命中 效率与部署 :LLM(title,abstract)

AI总结 本文提出CacheFlow框架,通过3D并行抽象优化KV缓存恢复,减少TTFT时间,提升大语言模型服务效率。

Comments 11 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20303 2026-04-29 cs.CL 81%

Citation Failure: Definition, Analysis and Efficient Mitigation

引用失败:定义、分析与高效缓解

Jan Buchmann, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab)(普遍知识处理实验室) Department of Computer Science(计算机科学系) Hessian Center for AI (hessian.AI)(黑森人工智能中心) Technical University of Darmstadt(达姆施塔特技术大学)

专题命中 效率与部署 :LLM(summary_cn,abstract);分类 cs.CL

AI总结 本文研究了LLM基于RAG系统的引用失败问题,提出CITECONTROL基准以分析失败模式,并通过CITENTION框架提升引用效率。

Comments Accepted to TACL in April 2024. Paper repository: https://github.com/UKPLab/tacl2026-citation-failure

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24933 2026-04-29 cs.AI cs.SD 79%

S-SONDO: Self-Supervised Knowledge Distillation for General Audio Foundation Models

S-SONDO:基于自监督的知识蒸馏用于通用音频基础模型

Mohammed Ali El Adlouni, Aurian Quelennec, Pierre Chouteau, Geoffroy Peeters, Slim Essid

专题命中 效率与部署 :foundation model(title,abstract);分类 cs.AI

AI总结 S-SONDO通过仅利用输出嵌入进行知识蒸馏,实现对通用音频模型的高效压缩,保留96%性能,适用于嵌入式教师模型。

Comments Accepted at IEEE ICASSP 2026. 5 pages, 2 figures, 3 tables. Equal contribution by first two authors. Code: https://github.com/MedAliAdlouni/ssondo | Models: https://huggingface.co/mohammedali2501/ssondo | Package: https://pypi.org/project/ssondo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21942 2026-04-29 physics.chem-ph cs.AI 79%

Suiren-1.0 Technical Report: A Family of Molecular Foundation Models

苏iren-1.0技术报告:分子基础模型家族

Junyi An, Xinyu Lu, Yun-Fei Shi, Li-Cheng Xu, Nannan Zhang, Chao Qu, Yuan Qi, Fenglei Cao

机构 * Shanghai Academy of AI for Science (SAIS)(上海人工智能科学研究院)

专题命中 效率与部署 :foundation model(title,abstract);分类 cs.AI

AI总结 苏iren-1.0通过结合3D构象几何与2D统计集合空间,提出分子基础模型家族,包含Suiren-Base、Suiren-Dimer和Suiren-ConfAvg三种变体,实现量子属性预测和分子表示学习的先进方法。

Comments 24 pages,5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25555 2026-04-29 cs.CR cs.AI 77%

From CRUD to Autonomous Agents: Formal Validation and Zero-Trust Security for Semantic Gateways in AI-Native Enterprise Systems

从CRUD到自主代理:面向AI原生企业系统的语义网关的正式验证与零信任安全

Ignacio Peyrano

机构 * Universidad Austral(乌尔苏拉大学)

专题命中 效率与部署 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出了一种基于模型上下文协议的语义网关,通过正式验证和实证评估,解决了AI原生系统中自主代理的安全验证问题,采用零信任安全模型和语义模糊化技术,实现了对动态状态转换系统的安全审计。

Comments 25 pages, 4 figures, 4 tables. Open-source proof-of-concept (47 automated tests, deterministic semantic fuzzer) available at https://github.com/PeyranoDev/semantic-gateway-poc

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25012 2026-04-29 cs.LG 77%

Why Search When You Can Transfer? Amortized Agentic Workflow Design from Structural Priors

为何需要搜索?从结构先验出发的 amortized 智能工作流设计

Shiyi Du, Jiayuan Liu, Weihua Du, Yue Huang, Jiayi Li, Yingtao Luo, Xiangliang Zhang, Vincent Conitzer, Carl Kingsford

机构 * Carnegie Mellon University(卡内基梅隆大学) Foundations of Cooperative AI Lab (FOCAL)(合作人工智能基础实验室) University of Notre Dame(圣约翰大学)

专题命中 效率与部署 :LLM(abstract,abstract_cn);foundation model(abstract);分类 cs.LG

AI总结 本文提出SWIFT框架,通过结构先验和跨任务工作流演示,实现工作流设计的 amortization,减少计算成本并提升泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25578 2026-04-29 cs.CL cs.AI 76%

Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling

Marco-MoE:开放多语言稀疏专家混合语言模型的高效再利用

Fan Jiang, Yu Zhao, Chenyang Lyu, Tianqi Shi, Yichao Du, Feihu Jiang, Longyue Wang, Weihua Luo

机构 * Alibaba International Digital Commerce(阿里巴巴国际数字商业)

专题命中 效率与部署 :language model(title);分类 cs.CL、cs.AI

AI总结 Marco-MoE通过极稀疏设计和密集模型再利用,在5T tokens上高效预训练,实现多语言和英语基准的领先性能,同时支持可扩展的语言扩展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24820 2026-04-29 cs.AR cs.AI 70%

Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding

Salca:一种面向高效长上下文注意力解码的稀疏性感知硬件加速器

Wang Fan, Wei Cao, Xi Zha, Kedi Ma, MingQian Sun, Jialin Chen, Fengzhe Zhang, Fan Zhang

机构 * Fudan University(复旦大学)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出Salca,通过软硬件协同设计,解决长上下文注意力解码中的计算与内存瓶颈,实现3.82倍速度提升和74.19倍能效提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09618 2026-04-29 cs.DC cs.AI cs.CR 70%

HearthNet: Edge Multi-Agent Orchestration for Smart Homes

HearthNet:边缘多智能体协调用于智能家居

Zhonghao Zhan, Krinos Li, Yefan Zhang, Hamed Haddadi

机构 * Imperial College London(伦敦帝国学院) Independent Researcher(独立研究员)

专题命中 效率与部署 :LLM(abstract,abstract_cn);分类 cs.AI

AI总结 HearthNet通过边缘多智能体系统解决智能家居中自然语言控制、设备故障和持续协调的问题,采用MQTT、Git共享状态和授权租赁实现设备管理。

Comments (CAIS 2026) Proceedings of the ACM Conference on AI and Agentic Systems, Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏