arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-04-08 至 2026-04-08 共收录 53 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 53 篇

2601.09726 2026-04-08 cs.CL 90%

Forgetting as a Feature: Cognitive Alignment of Large Language Models

作为特征的遗忘:大型语言模型的认知对齐

Alexandros Christoforos

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

AI总结 研究将遗忘视为大型语言模型的认知机制,通过建立基准测试评估时间推理、概念漂移适应和联想回忆,发现模型遗忘率与人类记忆效率的权衡相似,并提出概率记忆提示策略提升长期推理性能。

Comments arXiv admin note: This submission has been withdrawn by arXiv administrators due to incorrect authorship. Author list truncated

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06013 2026-04-08 cs.AI cs.CL 88%

Epistemic Blinding: An Inference-Time Protocol for Auditing Prior Contamination in LLM-Assisted Analysis

知识盲区:一种用于审计LLM辅助分析中先验污染的推理时协议

Michael Cuccarese

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出了一种在代理系统中用于审计LLM辅助分析中先验污染的推理时协议,通过替换实体标识符以匿名代码并比较输出与未盲对照组,以恢复审计性。

Comments code and LLM skill at: https://github.com/mcuccarese/epistemic-blinding 7 pages 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05887 2026-04-08 cs.AI 88%

HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model Inference

HybridKV: 为高效多模态大语言模型推理设计的混合KV缓存压缩

Bowen Zeng, Feiyang Ren, Jun Zhang, Xiaoling Gu, Ke Chen, Lidan Shou, Huan Li

机构 * The State Key Laboratory of Blockchain and Data Security, Zhejiang University(浙江大学区块链与数据安全国家重点实验室) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新区(滨江)区块链与数据安全研究院) Hangzhou Dianzi University(杭州电子科技大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出HybridKV框架,通过分层预算分配和不同压缩策略,有效减少KV缓存内存占用并提升解码速度,实验表明在11个多模态基准上内存减少达7.9倍,解码速度提升1.52倍,性能损失极小。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05502 2026-04-08 cs.CR cs.LG 88%

AttnDiff: Attention-based Differential Fingerprinting for Large Language Models

AttnDiff:基于注意力的差分指纹法用于大语言模型

Haobo Zhang, Zhenhua Xu, Junxian Li, Shangfeng Sheng, Dezhang Kong, Meng Han

机构 * Zhejiang University of Technology(浙江工业大学) Zhejiang University(浙江大学) Binjiang Institute of Zhejiang University(浙江大学滨江研究院) Shanghai Jiao Tong University(上海交通大学) University of Science and Technology of China(中国科学技术大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.LG

AI总结 本文提出AttnDiff,一种高效白盒框架,通过内在信息路由行为提取模型指纹,用于验证大语言模型的衍生关系,实现高相似度区分相关衍生模型与无关模型家族。

Comments Accepted at ACL2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09438 2026-04-08 cs.CL 88%

GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models

GrACE:一种生成方法,用于改进置信度 elicitation 和大语言模型在测试时的高效扩展

Zhaohan Zhang, Ziquan Liu, Ioannis Patras

机构 * Queen Mary University of London(伦敦玛丽女王大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 本文提出GrACE,一种生成方法,用于改进大语言模型的置信度 elicitation 和测试时的高效扩展,通过实时计算隐藏状态与特殊标记嵌入的相似性,实现可靠且可扩展的置信度评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13151 2026-04-08 cs.LG cs.CL 88%

Quantization-Robust LLM Unlearning via Low-Rank Adaptation

通过低秩适应实现抗量化的大语言模型反学习

João Vitor Boer Abitante, Joana Meneguzzo Pasquali, Luan Fonseca Garcia, Ewerton de Oliveira, Thomas da Silva Paula, Rodrigo C. Barros, Lucas S. Kupssinskü

机构 * Kunumi Institute, Brazil(Kunumi研究所,巴西)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 本文提出LoRA方法,通过冻结基础模型并集中反学习到可训练适配器中,提升4位量化下的模型性能,同时减少隐私泄露,适用于需要量化部署的场景。

Comments Accepted to IJCNN 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04816 2026-04-08 cs.OS cs.CL cs.DC 87%

Horizon-LM: A RAM-Centric Architecture for LLM Training

Horizon-LM:一种以RAM为中心的LLM训练架构

Zhengqing Yuan, Lichao Sun, Yanfang Ye

专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);instruction tuning(abstract)

AI总结 Horizon-LM通过重新定义CPU和GPU的角色,以主机内存为核心参数存储,实现大模型训练的内存优化,提升训练吞吐量并降低对多GPU集群的依赖。

Comments This paper contained an error in the throughput computation used in the experimental evaluation. Specifically, the TFLOPS calculation omitted the 12HL term in the training FLOPs formula, which led to systematic underestimation of the reported throughput numbers in the experimental results. We are withdrawing this version to correct the evaluation and avoid confusion for readers

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05137 2026-04-08 cs.PL cs.AI cs.CL cs.LG cs.SE 87%

EffiPair: Improving the Efficiency of LLM-generated Code with Relative Contrastive Feedback

EffiPair:通过相对对比反馈提升LLM生成代码的效率

Samira Hajizadeh, Suman Jana

机构 * Columbia University(哥伦比亚大学)

专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出EffiPair框架,通过相对对比反馈机制在测试时迭代优化代码效率,减少性能反馈开销,实验证明其在代码效率提升方面优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05629 2026-04-08 cs.CV 86%

A Unified Foundation Model for All-in-One Multi-Modal Remote Sensing Image Restoration and Fusion with Language Prompting

一种统一的多模态遥感图像修复与融合的全功能基础模型 with语言提示

Yongchuan Cui, Peng Liu

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院空天信息创新研究院) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院)

专题命中 效率与部署 :foundation model(title,abstract);prompting(title)

AI总结 本文提出LLaRS模型,通过最优传输和混合专家层实现多模态遥感图像修复与融合,结合大规模数据集提升模型性能与迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05460 2026-04-08 stat.ME cs.AI 85%

LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency

大语言模型评估作为张量补全:低秩结构与半参数效率

Jiachun Li, David Simchi-Levi, Will Wei Sun

机构 * Massachusetts Institute of Technology(麻省理工学院) Purdue University(普渡大学)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文将大语言模型评估视为低秩潜在评分张量的半参数推断问题,通过成对比较数据进行张量补全,提出了一种基于信息算子和影响函数的偏差修正估计器,解决了异质信息算子带来的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11652 2026-04-08 cs.DC cs.AI 85%

WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching

WISP: 通过动态草稿和SLO感知批处理实现边缘的废料和干扰抑制的分布式推测LLM服务

Xiangchen Li, Jiakun Fan, Qingyuan Wang, Dimitrios Spatharakis, Saeid Ghafouri, Hans Vandierendonck, Deepu John, Bo Ji, Ali R. Butt, Dimitrios S. Nikolopoulos

机构 * University College Dublin(都柏林大学学院) National Technical University of Athens(雅典国家技术大学) Queen's University Belfast(贝尔法斯特女王大学)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出WISP系统,通过动态草稿和SLO感知批处理优化边缘和云的负载平衡,提升系统容量和吞吐量。

Comments 31 Pages, 13 Figures, 13 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05375 2026-04-08 cs.MM 85%

DAT: Dual-Aware Adaptive Transmission for Efficient Multimodal LLM Inference in Edge-Cloud Systems

DAT:双重视觉与带宽感知的自适应传输,用于边缘-云计算系统中高效多模态大语言模型推理

Qi Guo, Zheming Yang, Yunqing Hu, Chang Zhao, Wen Ji

专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出DAT方法,通过轻量级边缘模型过滤非目标帧并触发MLLM推理,结合高效微调策略和多流自适应传输优化,实现高效多模态大语言模型推理,提升语义生成质量与低延迟警报性能。

Comments 10 pages, 6 figures. Submitted to ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05905 2026-04-08 cs.CL cs.AI cs.HC cs.LG cs.MA 85%

Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency

自信的幻觉?通过邻域一致性诊断LLM的真实性

Haoming Xu, Ningyuan Zhao, Yunzhi Yao, Weihong Xu, Hongru Wang, Xinle Deng, Shumin Deng, Jeff Z. Pan, Huajun Chen, Ningyu Zhang

机构 * Zhejiang University(浙江大学) University of Edinburgh(爱丁堡大学)

专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出邻域一致性信念(NCB)来评估LLM在上下文干扰下的信念鲁棒性,通过结构化测量提升模型在真实场景中的可靠性,并引入结构感知训练(SAT)减少知识脆性。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05865 2026-04-08 cs.AI cs.PL 84%

JTON: A Token-Efficient JSON Superset with Zen Grid Tabular Encoding for Large Language Models

JTON: 一种高效的JSON超集,采用Zen Grid表格编码以适应大语言模型

Gowthamkumar Nandakishore

专题命中 效率与部署 :large language model(title);language model(title);分类 cs.AI

AI总结 本文提出JTON,一种高效的JSON超集,通过Zen Grid表格编码减少冗余,提升大语言模型处理结构化数据的效率与准确性。

Comments 20 pages, 13 figures, 14 tables. Code and test suite available at https://github.com/gowthamkumar-nandakishore/JTON

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17766 2026-04-08 cs.CL cs.AI 84%

A State-Update Prompting Strategy for Efficient and Robust Multi-turn Dialogue

一种用于高效且鲁棒多轮对话的状态更新提示策略

Ziyi Liu

机构 * School of Artificial Intelligence Beijing University of Posts

专题命中 效率与部署 :prompting(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出一种无需训练的提示工程方法,通过状态重建和历史提醒机制优化多轮对话,提升信息过滤和问答性能,同时减少推理时间和token消耗。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05535 2026-04-08 cs.AI 83%

SignalClaw: LLM-Guided Evolutionary Synthesis of Interpretable Traffic Signal Control Skills

SignalClaw:基于大语言模型的进化合成可解释交通信号控制技能

Da Lei, Feng Xiao, Lu Li, Yuzhan Liu

机构 * Sichuan University(四川大学)

专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 SignalClaw通过大语言模型引导进化合成可解释的交通信号控制技能,结合模拟指标生成自然语言反馈,实现高效且可解释的交通信号控制策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05012 2026-04-08 cs.AR cs.AI 83%

Comparative Characterization of KV Cache Management Strategies for LLM Inference

KV缓存管理策略在大语言模型推理中的比较分析

Oteo Mamo, Olga Kogiou, Hyunjin Yi, Weikuan Yu

专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文通过实证研究比较了vLLM、InfiniGen和H2O三种KV缓存管理框架,分析了其在内存消耗和推理性能上的权衡,揭示了不同场景下最优的缓存策略配置。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20157 2026-04-08 cs.CV 82%

SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models

SigLino:高效多教师蒸馏用于聚类视觉基础模型

Sofian Chaybouti, Sanath Narayan, Yasser Dahou, Phúc H. Lê Khac, Ankit Singh, Ngoc Dung Huynh, Wamiq Reyaz Para, Hilde Kuehne, Hakim Hacid

机构 * Technology Innovation Institute(技术创新研究所) Tuebingen AI Center/University of Tuebingen(图宾根人工智能中心/图宾根大学) MIT-IBM Watson AI Lab(MIT-IBM Watson AI实验室)

专题命中 效率与部署 :foundation model(title,abstract);LLM(abstract)

AI总结 本文提出SigLino,一种高效的聚类视觉基础模型,通过同时蒸馏SigLIP2和DINOv3的知识到密集和专家混合学生中,提升了多教师蒸馏的效率与效果。

Comments 17 pages, 8 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05551 2026-04-08 cs.CL cs.AI cs.LG 82%

FastDiSS: Few-step Match Many-step Diffusion Language Model on Sequence-to-Sequence Generation--Full Version

FastDiSS: 少步匹配多步扩散语言模型在序列到序列生成中的应用——完整版本

Dat Nguyen-Cong, Tung Kieu, Hoang Thanh-Tung

机构 * FPT Software AI Center, FPT Corporation(FPT软件人工智能中心,FPT公司) Department of Computer Science, Aalborg University(奥尔堡大学计算机科学系) Quantum AI and Cyber Security Institute, FPT Corporation(FPT公司量子人工智能与网络安全研究所)

专题命中 效率与部署 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出FastDiSS模型,通过改进自条件机制和引入噪声感知机制,提升少步采样下的生成质量,并实现400倍更快的推理速度。

Comments camera-ready version, accepted by ACL Findings (ACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06169 2026-04-08 cs.LG cs.AI cs.CL stat.ML 80%

In-Place Test-Time Training

就地测试时间训练

Guhao Feng, Shengjie Luo, Kai Hua, Ge Zhang, Di He, Wenhao Huang, Tianle Cai

机构 * Peking University(北京大学)

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出了一种无需重新训练的就地测试时间训练框架,通过调整MLP块的最终投影矩阵实现模型在推理时的动态更新,提升大语言模型在长上下文任务中的性能。

Comments ICLR 2026 Oral Presentation; Code is released at https://github.com/ByteDance-Seed/In-Place-TTT

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04992 2026-04-08 cs.CR cs.AI 79%

FreakOut-LLM: The Effect of Emotional Stimuli on Safety Alignment

FreakOut-LLM:情绪刺激对安全对齐的影响

Daniel Kuznetsov, Ofir Cohen, Karin Shistik, Rami Puzis, Asaf Shabtai

机构 * Ben-Gurion University of the Negev(内盖夫本-古里安大学)

专题命中 效率与部署 :LLM(title,abstract);分类 cs.AI

AI总结 研究探讨情绪刺激是否影响大语言模型的安全对齐,通过实验发现压力情境显著增加逃逸成功率,而放松情境无显著影响,揭示情绪状态对AI攻击的可测量影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20309 2026-04-08 cs.LG 79%

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models

QuantVLA: 用于视觉-语言-动作模型的可扩展后训练量化

Jingxuan Zhang, Yunta Hsieh, Zhongwei Wan, Haokun Lin, Xin Wang, Ziqi Wang, Yingtie Lei, Mi Zhang

机构 * The Ohio State University(俄亥俄州立大学) University of Michigan(密歇根大学) City University of Hong Kong(香港城市大学)

专题命中 效率与部署 :post-training(title,abstract);分类 cs.LG

AI总结 QuantVLA是一种无需训练的后训练量化框架,首次应用于视觉-语言-动作系统,并成功量化了扩散变换器(DiT)动作头,通过三种可扩展组件实现了高效的低比特量化,显著提升任务成功率并降低内存消耗。

Comments CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15089 2026-04-08 cs.LG stat.ML 79%

Triplet Feature Fusion for Equipment Anomaly Prediction : An Open-Source Methodology Using Small Foundation Models

三元特征融合用于设备异常预测:一种使用小型基础模型的开源方法论

Takato Yasuno

专题命中 效率与部署 :foundation model(title,abstract);分类 cs.LG

AI总结 本文提出一种结合开源小型基础模型的三元特征融合方法,通过统计特征、时间序列嵌入和多语言文本嵌入的整合,提升设备异常预测的精度与鲁棒性。

Comments 15 pages, 8 figures, 7 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05961 2026-04-08 cs.CV cs.AI 79%

Aligned Vector Quantization for Edge-Cloud Collabrative Vision-Language Models

对齐的向量量化用于边缘-云协作的视觉-语言模型

Xiao Liu, Lijun Zhang, Deepak Ganesan, Hui Guan

专题命中 效率与部署 :language model(title,abstract);分类 cs.AI

AI总结 本文提出LLaVA-AlignedVQ系统,通过AlignedVQ算法实现中间特征高效压缩,减少数据传输开销,提升推理速度,保持高精度。

Comments I found a big mistake in the paper that causes significant bias on the results. The residual links are not taken into consideration when computing the transmission. All results about the compressed data size and transmission latency would be affected

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07663 2026-04-08 cs.DB cs.AI cs.LG 79%

Cortex AISQL: A Production SQL Engine for Unstructured Data

Cortex AISQL:面向非结构化数据的生产级SQL引擎

Paweł Liskowski, Benjamin Han, Paritosh Aggarwal, Bowei Chen, Boxin Jiang, Nitish Jindal, Zihan Li, Aaron Lin, Kyle Schmaus, Jay Tayade, Weicheng Zhao, Anupam Datta, Nathan Wiegand, Dimitris Tsirogiannis

机构 * Snowflake Inc.(雪花公司)

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 Cortex AISQL通过AI-aware优化、自适应模型级联和语义连接重写技术,解决大规模语义操作效率问题,提升非结构化数据查询性能。

Comments Published in SIGMOD Companion '26 (Industry Track), Bengaluru, India, May 31-June 5, 2026. ACM DOI: 10.1145/3788853.3803093. This version is the published ACM Version of Record under the Creative Commons Attribution 4.0 International (CC BY 4.0) license

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05688 2026-04-08 cs.CL cs.AI 79%

Attention Editing: A Versatile Framework for Cross-Architecture Attention Conversion

注意力编辑:一种跨架构注意力转换的通用框架

Zhen Cheng, Hao-Bo Yang, Wan-Yi Huang, Jin-Long Li

机构 * China Merchants Bank Artificial Intelligence Laboratory(招商银行人工智能实验室)

专题命中 效率与部署 :large language model(abstract);language model(abstract);pretraining(abstract);分类 cs.CL、cs.AI

AI总结 本文提出Attention Editing框架,通过渐进蒸馏实现无需重训练的注意力架构转换,提升大语言模型的效率与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02810 2026-04-08 cs.LG cs.AI cs.SE 79%

Dissecting Transformers: A CLEAR Perspective towards Green AI

解析Transformer:一种面向绿色AI的CLEAR视角

Hemang Jain, Shailender Goyal, Divyansh Pandey, Karthik Vaidhyanathan

机构 * IIIT Hyderabad(印度国际信息技术学院海得拉巴分校)

专题命中 效率与部署 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出CLEAR方法,通过组件级能耗评估,分析Transformer组件在不同参数下的能耗,揭示注意力机制的高能耗问题,为预测能效提供理论基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05934 2026-04-08 cs.CV eess.IV 78%

Leveraging Image Editing Foundation Models for Data-Efficient CT Metal Artifact Reduction

利用图像编辑基础模型实现数据高效的CT金属伪影减少

Ahmet Rasim Emirdagi, Süleyman Aslan, Mısra Yavuz, Görkay Aydemir, Yunus Bilge Kurt, Nasrin Rahimi, Burak Can Biner, M. Akın Yılmaz

机构 * Codeway AI Research(Codeway AI 研究院)

专题命中 效率与部署 :foundation model(title,abstract)

AI总结 本文提出利用视觉语言扩散基础模型进行金属伪影减少,通过参数高效低秩适应技术,仅需16-128对训练样本即可实现高效伪影抑制,并证明领域适应对减少幻觉至关重要。

Comments Accepted to CVPRW 2026 Med-Reasoner

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18117 2026-04-08 cs.CV 78%

Online In-Context Distillation for Low-Resource Vision Language Models

在线上下文蒸馏用于低资源视觉语言模型

Zhiqi Kang, Rahaf Aljundi, Vaggelis Dorovatas, Karteek Alahari

机构 * Inria(法国国家信息与自动化研究所) Toyota Motor Europe(丰田汽车欧洲公司)

专题命中 效率与部署 :language model(title,abstract)

AI总结 本文提出在线上下文蒸馏方法,通过小模型与强教师模型协作,利用稀疏演示高效缩小性能差距,提升小模型性能达33%,在低资源条件下实现与零样本性能相当的成果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05793 2026-04-08 cs.CR cs.CV 78%

BodhiPromptShield: Pre-Inference Prompt Mediation for Suppressing Privacy Propagation in LLM/VLM Agents

BodhiPromptShield: 预推理提示调解用于抑制LLM/VLM代理中的隐私传播

Bo Ma, Jinsong Wu, Weiqi Yan

机构 * Auckland University of Technology(奥克兰理工大学) University of Chile(智利大学)

专题命中 效率与部署 :LLM(title,abstract)

AI总结 本文提出BodhiPromptShield框架,通过检测敏感片段、路由和延迟恢复来抑制LLM/VLM代理中的跨阶段隐私传播,实验显示在控制提示隐私基准(CPPB)上效果显著。

详情

展开后加载摘要…

URL PDF HTML 收藏