arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7552 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7552 篇

2603.02865 2026-03-04 cs.CL cs.CV 79%

Nodes Are Early, Edges Are Late: Probing Diagram Representations in Large Vision-Language Models

节点早于边,边晚于节点:在大规模视觉-语言模型中探测图表示

Haruto Yoshida, Keito Kudo, Yoichi Aoki, Ryota Tanaka, Itsumi Saito, Keisuke Sakaguchi, Kentaro Inui

机构 * Tohoku University(东大大学) Human Informatics Labs.(人文学信息实验室) NTT, Inc.(NTT公司)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 研究发现大规模视觉-语言模型在处理图结构时,节点信息早于边信息被线性编码,导致模型在理解边关系时存在困难。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05535 2026-02-27 cs.LG 79%

Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty Quantification

通过证据不确定性量化检测大视觉-语言模型的误行

Tao Huang, Rui Wang, Xiaofei Liu, Yi Qin, Li Duan, Liping Jing

机构 * State Key Laboratory of Advanced Rail Autonomous Operation(先进轨道交通自主运行国家重点实验室) Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence(北京交通数据挖掘与具身智能重点实验室) School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院) School of Automation and Intelligence, Beijing Jiaotong University(北京交通大学自动化与智能学院) Beijing Key Laboratory of Security and Privacy in Intelligent Transportation(北京智能交通安全与隐私重点实验室)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.LG

AI总结 通过证据不确定性量化检测大视觉-语言模型的误行,识别内部冲突和无知以提高模型可靠性。

Comments Accepted to ICLR 2026. Code is available at https://github.com/HT86159/EUQ

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21704 2026-02-26 cs.CV cs.AI 79%

Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models

动态多模态激活引导用于大型视觉-语言模型中的幻觉缓解

Jianghao Yin, Qin Chen, Kedi Chen, Jie Zhou, Xingjiao Wu, Liang He

机构 * East China Normal University(华东师范大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 本文提出动态多模态激活引导方法,通过构建语义真实性引导向量数据库,实现幻觉缓解,提升模型性能。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21178 2026-02-25 cs.CV cs.AI 79%

XMorph: Explainable Brain Tumor Analysis Via LLM-Assisted Hybrid Deep Intelligence

XMorph: 通过LLM辅助混合深度智能进行可解释性脑肿瘤分析

Sepehr Salem Ghahfarokhi, M. Moein Esfahani, Raj Sunderraman, Vince Calhoun, Mohammed Alser

机构 * Department of Computer Science, Georgia State University(计算机科学系,佐治亚州立大学) TReNDS Center, Georgia State University(TReNDS中心,佐治亚州立大学)

专题命中 知识编辑与模型理解 :LLM(title,abstract);分类 cs.AI

AI总结 XMorph通过结合LLM和混合深度智能,实现了对脑肿瘤的高效可解释分类,准确率达96.0%。

Comments Accepted in ICCABS 2026: The 14th International Conference on Computational Advances in Bio and Medical Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19022 2026-02-24 cs.CV cs.AI 79%

An interpretable framework using foundation models for fish sex identification

基于基础模型的可解释框架用于鱼类性别识别

Zheng Miao, Tien-Chieh Hung

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.AI

AI总结 基于基础模型的可解释框架用于濒危鱼类三角洲虾虎鱼的性别识别,通过原型网络提升鲁棒性与可解释性,实现高准确率识别。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17532 2026-02-20 q-bio.GN cs.AI 79%

Systematic Evaluation of Single-Cell Foundation Model Interpretability Reveals Attention Captures Co-Expression Rather Than Unique Regulatory Signal

单细胞基础模型解释性系统评估揭示注意力捕捉共表达而非独特调控信号

Ihor Kendiukhov

机构 * Department of Computer Science University of Tübingen(计算机科学系图宾根大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.AI

AI总结 单细胞基础模型中注意力机制捕捉共表达而非独特调控信号,CSSI提升GRN恢复效率

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16980 2026-02-20 cs.LG cs.CR 79%

Discovering Universal Activation Directions for PII Leakage in Language Models

发现语言模型中PII泄露的通用激活方向

Leo Marchyok, Zachary Coalson, Sungho Keum, Sooel Son, Sanghyun Hong

机构 * Oregon State University, Corvallis OR, USA Korea Advanced Institute of Science \& Technology, Daejeon, South Korea

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.LG

AI总结 UniLeak通过识别语言模型中通用激活方向,提升PII泄露可能性,为隐私安全研究提供新视角。

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13494 2026-02-19 cs.CV cs.AI 79%

Language-Guided Invariance Probing of Vision-Language Models

语言引导的视觉-语言模型不变性探测

Jae Joong Lee

机构 * Department of Computer Science, Purdue University(普渡大学计算机科学系)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 本文提出LGIP基准,用于评估视觉-语言模型对语言扰动的鲁棒性,发现EVA02-CLIP和OpenCLIP表现优异,而SigLIP存在显著缺陷。

Comments Pattern Recognition Letters 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09378 2026-02-13 cs.SD cs.AI cs.MM eess.AS 79%

Prevailing Research Areas for Music AI in the Era of Foundation Models

基础模型时代音乐AI的主流研究领域

Megan Wei, Mateusz Modrzejewski, Aswin Sivaraman, Dorien Herremans

机构 * Brown University(布朗大学) Warsaw University of Technology(华沙技术大学) Indiana University(印第安纳大学) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.AI

AI总结 本文探讨了基础模型时代音乐AI的研究领域,涵盖基础模型、多模态系统、数据集、效率、生成模型应用及版权问题,旨在揭示未来研究方向。

Journal ref Proceedings of Machine Learning Research, PMLR 303:1-23, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08048 2026-02-10 cs.CL 79%

TDGNet: Hallucination Detection in Diffusion Language Models via Temporal Dynamic Graphs

TDGNet: 通过时间动态图检测扩散语言模型中的幻觉

Arshia Hemmat, Philip Torr, Yongqiang Chen, Junchi Yu

机构 * Department of Computer Science, University of Oxford(牛津大学计算机科学系) Department of Engineering Science, University of Oxford(牛津大学工程科学系) Carnegie Mellon University(卡内基梅隆大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 TDGNet通过时间动态图框架,利用演进的token级注意力图进行学习,实现对扩散语言模型中幻觉的高效检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07080 2026-02-10 cs.SE cs.AI 79%

CodeCircuit: Toward Inferring LLM-Generated Code Correctness via Attribution Graphs

CodeCircuit: 通过归因图推断LLM生成代码的正确性

Yicheng He, Zheng Zhao, Zhou Kaiyu, Bryan Dai, Jie Fu, Yonghui Yang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Edinburgh(爱丁堡大学) Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学) IQuest Research(IQuest研究)

专题命中 知识编辑与模型理解 :LLM(title,abstract);分类 cs.AI

AI总结 CodeCircuit通过归因图分析LLM内部结构,推断生成代码的正确性,验证内部动态信号的预测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10273 2026-02-09 cs.CV cs.AI 79%

Probing Perceptual Constancy in Large Vision-Language Models

探测大型视觉-语言模型中的知觉恒常性

Haoran Sun, Bingyang Wang, Suyang Yu, Yijiang Li, Qingying Gao, Haiyun Lyu, Lianyu Huang, Zelong Hong, Jiahui Ge, Qianli Ma, Hang He, Yifan Zhou, Lingzi Guo, Lantao Mei, Maijunxian Wang, Dezhi Luo, Hokin Deng

机构 * Johns Hopkins University(约翰霍普金斯大学) Emory University(埃默里大学) University of Washington(华盛顿大学) University of California San Diego(加州大学圣地亚哥分校) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Southern California(南加州大学) Washington University in St. Louis(圣路易斯华盛顿大学) Shanghai Jiao Tong University(上海交通大学) East China Normal University(华东师范大学) Stanford University(斯坦福大学) University of California, Berkeley(加州大学伯克利分校) University of Michigan(密歇根大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 本文研究了大型视觉-语言模型在颜色、大小和形状恒常性任务中的表现,发现模型在不同任务上的性能存在显著差异。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00377 2026-02-06 cs.CL 79%

DecompressionLM: Deterministic, Diagnostic, and Zero-Shot Concept Graph Extraction from Language Models

DecompressionLM:从语言模型中确定性、诊断性地提取零样本概念图

Zhaochen Hong, Jiaxuan You

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 DecompressionLM通过无状态框架实现零样本概念图提取,揭示语言模型编码内容,解决跨序列耦合、竞争解码效应和可扩展性限制问题,提升压缩模型的知识广度和事实基础评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04560 2026-02-06 cs.LG stat.ML 79%

GAMformer: Bridging Tabular Foundation Models and Interpretable Machine Learning

GAMformer: 联接表格基础模型与可解释机器学习

Andreas Mueller, Julien Siems, Harsha Nori, David Salinas, Arber Zela, Rich Caruana, Frank Hutter

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.LG

AI总结 GAMformer是首个针对GAMs的表格基础模型,通过上下文学习在一个前向传递中估计GAM形状函数,实现了基础模型的威力与关键应用的可解释性需求之间的平衡。

Comments 22 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02464 2026-02-03 cs.CL 79%

From Directions to Regions: Decomposing Activations in Language Models via Local Geometry

从方向到区域:通过局部几何分解语言模型中的激活

Or Shafran, Shaked Ronen, Omri Fahn, Shauli Ravfogel, Atticus Geiger, Mor Geva

机构 * Blavatnik School of Computer Science(Blavatnik计算机科学学院) AI, Tel Aviv University, Israel(人工智能,特拉维夫大学,以色列) New York University, New York, NY, USA(纽约大学,纽约,纽约州,美国)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 本研究提出MFA方法,通过局部几何分解语言模型中的激活,捕捉复杂非线性结构,优于无监督基线并优于稀疏自编码器的转向性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02029 2026-02-03 cs.AI cs.SE 79%

Canonical Intermediate Representation for LLM-based optimization problem formulation and code generation

基于LLM的优化问题建模与代码生成的规范中间表示

Zhongyuan Lyu, Shuoyu Hu, Lujie Liu, Hongxia Yang, Ming LI

机构 * The Hong Kong Polytechnic University, Hong Kong, China(香港理工大学)

专题命中 知识编辑与模型理解 :LLM(title,abstract);分类 cs.AI

AI总结 本研究提出规范中间表示(CIR)以提升LLM在复杂运营规则下的优化问题建模与代码生成能力,通过多智能体框架实现高精度建模。

Comments 41 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01148 2026-02-03 cs.CV cs.AI 79%

SocialFusion: Addressing Social Degradation in Pre-trained Vision-Language Models

SocialFusion: 解决预训练视觉-语言模型中的社会退化问题

Hamza Tahboub, Weiyan Shi, Gang Hua, Huaizu Jiang

机构 * Northeastern University(东北大学) Amazon.com, Inc.(亚马逊公司)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 SocialFusion通过统一框架解决预训练视觉-语言模型中的社会退化问题,实现跨社会任务的正迁移并提升整体性能。

Comments 22 pages, 10 figures. Published in Transactions on Machine Learning Research (TMLR)

Journal ref Transactions on Machine Learning Research, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00219 2026-02-03 cs.CR cs.AI 79%

Tri-LLM Cooperative Federated Zero-Shot Intrusion Detection with Semantic Disagreement and Trust-Aware Aggregation

三LLM协作联邦零日入侵检测:基于语义分歧与信任感知的聚合

Saeid Jamshidi, Omar Abdul Wahab, Foutse Khomh, Kawser Wazed Nafi

机构 * SWAT Laboratory, Polytechnique Montréal(Polytechnique Montréal SWAT 实验室) Department of Computer and Software Engineering, Polytechnique Montréal(Polytechnique Montréal 计算机与软件工程系)

专题命中 知识编辑与模型理解 :LLM(title,abstract);分类 cs.AI

AI总结 本文提出基于三LLM的联邦IDS框架,通过语义监督实现零日入侵检测,提升对未知攻击行为的识别能力,同时增强对异构和不可靠客户端的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22203 2026-02-02 q-bio.GN cs.AI 79%

Beyond Conditional Computation: Retrieval-Augmented Genomic Foundation Models with Gengram

超越条件计算:具有Gengram的检索增强基因组基础模型

Huinan Xu, Xuyang Feng, Junhong Chen, Junchen Liu, Kaiwen Deng, Kai Ding, Shengning Long, Jiaxue Shuai, Zhaorong Li, Shiping Liu, Guirong Xue, Zhan Xiao

机构 * Genos Team(基因组团队)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.AI

AI总结 Gengram通过基因组特定哈希方案引入高效查找原语,提升基因组基础模型在功能基因组任务中的性能和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22028 2026-01-30 cs.LG 79%

From Logits to Latents: Contrastive Representation Shaping for LLM Unlearning

从 logits 到 latents:用于大语言模型遗忘的对比表示塑造

Haoran Tang, Rajiv Khanna

机构 * Purdue University(普渡大学)

专题命中 知识编辑与模型理解 :LLM(title,abstract);分类 cs.LG

AI总结 CLReg 通过对比表示正则化减少大语言模型中遗忘与保留知识的纠缠,提升遗忘效果并降低隐私风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18901 2026-01-28 cs.CL 79%

Self-Aware Knowledge Probing: Evaluating Language Models' Relational Knowledge through Confidence Calibration

具有自意识的知识探测:通过置信度校准评估语言模型的关系知识

Christopher Kissling, Elena Merdjanovska, Alan Akbik

机构 * Humboldt-Universität zu Berlin(柏林洪堡大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 本文提出了一种新的校准探测框架,用于评估语言模型在关系知识上的置信度,揭示了大多数模型在置信度校准上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11739 2026-01-21 cs.CL 79%

Bridging Human Interpretation and Machine Representation: A Landscape of Qualitative Data Analysis in the LLM Era

弥合人类解释与机器表示:在大语言模型时代中定性数据分析的图景

Xinyu Pi, Qisen Yang, Chuong Nguyen, Hua Shen

专题命中 知识编辑与模型理解 :LLM(title,abstract);分类 cs.CL

AI总结 本文提出了一种4×4的图景,用于分析大语言模型在定性数据分析中的意义构建与建模层次,揭示现有系统在解释性和理论性推断方面的不足,并提出改进方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11666 2026-01-21 cs.CV cs.AI 79%

MATEX: Multi-scale Attention and Text-guided Explainability of Medical Vision-Language Models

MATEX: 医疗视觉-语言模型的多尺度注意力与文本引导可解释性

Muhammad Imran, Chi Lee, Yugyung Lee

机构 * Computer Science, School of Science and Engineering(科学与工程学院计算机科学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 MATEX通过多尺度注意力与文本引导方法提升医疗视觉-语言模型的可解释性,实现更精确和临床相关的梯度归因图。

Comments 12 pages, 3 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07863 2026-01-16 cs.CL 79%

QoSBERT: An Uncertainty-Aware Approach based on Pre-trained Language Models for Service Quality Prediction

QoSBERT:基于预训练语言模型的不确定性感知方法用于服务质量预测

Ziliang Wang, Xiaohong Zhang, Ze Shi Li, Meng Yan

机构 * Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education(高可信软件技术重点实验室(北京大学)) School of Computer Science, Peking University(北京大学计算机科学学院) Key Laboratory of Dependable Service Computing in Cyber Physical Society (Chongqing University), Ministry of Education, China(网络物理社会可信服务计算重点实验室(重庆大学)) School of Big Data and Software Engineering, Chongqing University(重庆大学大数据与软件工程学院) Department of Computer Science at the University of Victoria(维多利亚大学计算机科学系)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 QoSBERT基于预训练语言模型,通过不确定性估计提升服务质量预测的准确性和可靠性。

Journal ref IEEE Transactions on Services Computing ( Volume: 18, Issue: 6, Nov.-Dec. 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08860 2026-01-15 cs.CV cs.AI 79%

Bias Detection and Rotation-Robustness Mitigation in Vision-Language Models and Generative Image Models

视觉-语言模型和生成图像模型中的偏见检测与旋转鲁棒性缓解

Tarannum Mithila

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 本文提出旋转鲁棒缓解策略,通过数据增强、表征对齐和模型正则化,提升视觉-语言和生成图像模型在旋转和分布偏移下的鲁棒性和公平性。

Comments Preprint. This work is derived from the author's Master's research. Code and supplementary materials will be released separately

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08001 2026-01-14 cs.AI 79%

Interpreting Fedspeak with Confidence: A LLM-Based Uncertainty-Aware Framework Guided by Monetary Policy Transmission Paths

以信心解读Fedspeak:一种基于大语言模型的不确定性感知框架,由货币政策传导路径引导

Rui Yao, Qi Chai, Jinhai Yao, Siyuan Li, Junhao Chen, Qi Zhang, Hao Wang

专题命中 知识编辑与模型理解 :LLM(title,abstract);分类 cs.AI

AI总结 本文提出基于大语言模型的不确定性感知框架,用于解读Fedspeak并分类其货币政策立场,通过动态不确定性解码模块提升模型可靠性与准确性。

Comments Accepted by AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07006 2026-01-13 cs.AI 79%

LLM Performance Predictors: Learning When to Escalate in Hybrid Human-AI Moderation Systems

LLM性能预测器:在混合人机审核系统中学习何时升级

Or Bachar, Or Levi, Sardhendu Mishra, Adi Levi, Manpreet Singh Minhas, Justin Miller, Omer Ben-Porat, Eilon Sheetrit, Jonathan Morra

机构 * Reichman University(雷赫曼大学)

专题命中 知识编辑与模型理解 :LLM(title,abstract);分类 cs.AI

AI总结 本文提出了一种基于LLM性能预测器的框架,用于在混合人机审核系统中学习何时升级,通过提升准确性与成本权衡实现更高效的内容审核。

Comments Accepted as a full paper at the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06818 2026-01-13 cs.CL 79%

AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents

AgentHallu: 评估基于大语言模型的代理的自动幻觉归因

Xuannan Liu, Xiao Yang, Zekun Li, Peipei Li, Ran He

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Department of Computer Science & Technology, Tsinghua University(清华大学计算机科学与技术系) University of California, Santa Barbara(加州大学圣芭芭拉分校) Center for Research on Intelligent Perception and Computing, NLPR, CASIA(智能感知与计算研究中心,国家工程实验室)

专题命中 知识编辑与模型理解 :LLM(title,abstract);分类 cs.CL

AI总结 AgentHallu提出一个评估基于大语言模型的代理自动识别幻觉来源的任务,通过高质量轨迹和多级注释评估13个模型,发现顶级模型在定位幻觉步骤上表现有限,工具使用幻觉最难识别。

Comments Project page: https://liuxuannan.github.io/AgentHallu.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09503 2026-01-06 cs.LG 79%

Towards Fair In-Context Learning with Tabular Foundation Models

朝着基于表格基础模型的公平上下文学习

Patrik Kenfack, Samira Ebrahimi Kahou, Ulrich Aïvodji

机构 * ÉTS Montréal(蒙特利尔ÉTS) Mila - Quebec AI Institute(魁北克人工智能研究所) University of Calgary(卡尔加里大学) CIFAR(加拿大基础科学研究院)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.LG

AI总结 本文提出通过三种预处理方法提升基于表格基础模型的上下文学习公平性,实验表明基于不确定性的策略有效提高公平性指标且对预测准确性影响小。

Comments Published in Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23453 2025-12-30 cs.CV cs.AI 79%

CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models

CoFi-Dec:通过粗到细的生成反馈在大视觉-语言模型中实现抗幻觉解码

Zongsheng Cao, Yangfan He, Anran Liu, Jun Xie, Feng Chen, Zepeng Wang

机构 * Researcher(研究者)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 CoFi-Dec通过粗到细的生成反馈机制,在大视觉-语言模型中减少幻觉,提升解码的鲁棒性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏