arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7552 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7552 篇

2602.09825 2026-02-11 cs.CV 78%

SAKED: Mitigating Hallucination in Large Vision-Language Models via Stability-Aware Knowledge Enhanced Decoding

SAKED: 通过稳定性感知知识增强解码缓解大视觉-语言模型中的幻觉

Zhaoxu Li, Chenqi Kong, Peijun Bao, Song Xia, Yi Tu, Yi Yu, Xinghao Jiang, Xudong Jiang

机构 * ROSE Lab, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(ROSE实验室,电子工程学院,南洋理工大学,新加坡) ROSE Lab, Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore(ROSE实验室,跨学科研究生项目,南洋理工大学,新加坡) School of Physical and Mathematical Sciences, Nanyang Technological University, Singapore(物理与数学科学学院,南洋理工大学,新加坡) Shanghai Jiao Tong University, China(上海交通大学,中国)

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 SAKED通过引入稳定性感知知识增强解码方法,有效缓解大视觉-语言模型中的幻觉问题,无需训练即可集成至不同架构中,实现最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06964 2026-02-09 cs.LG cs.AI cs.CL 78%

Learning a Generative Meta-Model of LLM Activations

学习大语言模型激活的生成元模型

Grace Luo, Jiahai Feng, Trevor Darrell, Alec Radford, Jacob Steinhardt

专题命中 知识编辑与模型理解 :LLM(title);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究通过训练扩散模型学习大语言模型激活的分布,提出生成元模型以提升干预保真度和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06213 2026-02-09 eess.AS 78%

From Hallucination to Articulation: Language Model-Driven Losses for Ultra Low-Bitrate Neural Speech Coding

从幻觉到明确:基于语言模型的损失函数用于超低比特率神经语音编码

Jayeon Yi, Minje Kim

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 本文提出基于语言模型的损失函数,用于超低比特率神经语音编码,以缓解音素幻觉问题,提升语义一致性与输出质量。

Comments To appear in ICASSP 2026. Demo wavs, code, and checkpoints (currently) availble at https://github.com/stet-stet/lmloss-icassp2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09606 2026-02-03 cs.CV 78%

Feat2GS: Probing Visual Foundation Models with Gaussian Splatting

Feat2GS: 通过高斯点云探测视觉基础模型

Yue Chen, Xingyu Chen, Anpei Chen, Gerard Pons-Moll, Yuliang Xiu

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 Feat2GS通过高斯点云探测视觉基础模型的3D理解能力,无需3D数据即可分析几何和纹理意识。

Comments Project Page: https://fanegg.github.io/Feat2GS/

Journal ref Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21380 2026-01-30 cs.CR 78%

RerouteGuard: Understanding and Mitigating Adversarial Risks for LLM Routing

RerouteGuard: 理解并缓解LLM路由中的对抗风险

Wenhui Zhang, Huiyu Xu, Zhibo Wang, Zhichao Li, Zeqing He, Xuelin Wei, Kui Ren

专题命中 知识编辑与模型理解 :LLM(title,abstract)

AI总结 RerouteGuard通过动态嵌入检测和自适应阈值,有效检测并缓解LLM路由中的对抗风险,提升多模型AI系统的安全性。

Comments 15 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23077 2026-01-27 cs.RO 78%

Embodied Learning of Reward for Musculoskeletal Control with Vision Language Models

具身学习奖励以实现肌肉骨骼控制与视觉语言模型

Saraswati Soedarmadji, Yunyue Wei, Chen Zhang, Yisong Yue, Yanan Sui

机构 * Tsinghua University(清华大学) California Institute of Technology(加州理工学院)

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 本文提出MoVLR框架,利用视觉语言模型实现高维肌肉骨骼系统的具身学习,通过迭代交互发现和改进奖励函数。

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08876 2026-01-15 cs.CV 78%

The Semantic Lifecycle in Embodied AI: Acquisition, Representation and Storage via Foundation Models

具身AI中的语义生命周期:通过基础模型实现获取、表示与存储

Shuai Chen, Hao Chen, Yuanchen Bei, Tianyang Zhao, Zhibo Zhou, Feiran Huang

机构 * College of Information Science and Technology, Jinan University(信息科学与技术学院,暨南大学) Faculty of Data Science, City University of Macau(数据科学学院,澳门城市大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Zhongguancun Laboratory(中关村实验室) College of Cyber Security, Jinan University(网络安全学院,暨南大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 本文提出语义生命周期框架,通过基础模型在具身AI中实现语义信息的获取、表示与存储,探讨了语义处理的连续流动与维护,并总结了当前挑战与未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07585 2026-01-13 cs.CV 78%

Robust Multicentre Detection and Classification of Colorectal Liver Metastases on CT: Application of Foundation Models

鲁棒多中心结直肠肝转移的CT检测与分类:基础模型的应用

Shruti Atul Mali, Zohaib Salahuddin, Yumeng Zhang, Andre Aichert, Xian Zhong, Henry C. Woodruff, Maciej Bobowicz, Katrine Riklund, Juozas Kupčinskas, Lorenzo Faggioni, Roberto Francischello, Razvan L Miclea, Philippe Lambin

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 本研究提出基于基础模型的AI流程,用于多中心CT中结直肠肝转移的稳健检测与分类,通过整合不确定性量化和可解释性,提升了分类准确率和病变检测效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03422 2026-01-13 cs.RO cs.CV 78%

What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models

机器人中最佳的3D场景表示是什么?从几何到基础模型

Tianchen Deng, Yue Pan, Shenghai Yuan, Dong Li, Chen Wang, Mingrui Li, Long Chen, Lihua Xie, Danwei Wang, Jingchuan Wang, Javier Civera, Hesheng Wang, Weidong Chen

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 本文探讨了机器人中最佳的3D场景表示方法,对比了传统和神经表示的优劣,并展望了基础模型在机器人应用中的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01984 2026-01-06 cs.CV 78%

Thinking with Blueprints: Assisting Vision-Language Models in Spatial Reasoning via Structured Object Representation

基于蓝图的思考:通过结构化物体表示协助视觉-语言模型进行空间推理

Weijian Ma, Shizhao Sun, Tianyu Yu, Ruiyu Wang, Tat-Seng Chua, Jiang Bian

机构 * National University of Singapore(新加坡国立大学) Microsoft Research, Asia(微软亚洲研究院) Tsinghua University(清华大学) University of Toronto(多伦多大学)

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 通过结构化物体表示提升视觉-语言模型的空间推理能力,引入蓝图嵌入推理轨迹、蓝图意识奖励和反捷径数据增强技术。

Comments Preprint. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00659 2026-01-05 cs.CV 78%

CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models

CRoPS:一种用于视觉-语言模型的无训练幻觉缓解框架

Neeraj Anand, Samyak Jha, Udbhav Bamba, Rahul Rahaman

机构 * Indian Institute of Technology (ISM)(印度理工学院(ISM)) Transmute AI National University of Singapore(新加坡国立大学)

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 CRoPS通过选择性移除关键文本标记和广义对比解码,有效缓解视觉-语言模型的幻觉问题,提升CHAIR分数20%并优于现有无训练方法。

Comments Accepted at TMLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17967 2025-12-23 cs.DB 78%

Memelang: An Axial Grammar for LLM-Generated Vector-Relational Queries

Memelang:一种用于LLM生成向量关系查询的轴性语法

Bri Holt

专题命中 知识编辑与模型理解 :LLM(title,abstract)

AI总结 Memelang是一种紧凑的查询语言,通过轴性语法实现LLM生成的向量关系查询的结构化生成,支持坐标稳定引用、变量绑定和隐式上下文传递,以提高查询效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15708 2025-12-18 cs.CV 78%

Multi-View Foundation Models

多视图基础模型

Leo Segre, Or Hirschorn, Shai Avidan

机构 * Tel Aviv University(特拉维夫大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 本文提出一种将基础模型转换为多视图基础模型的方法,通过引入3D感知注意力层提升多视角特征一致性,应用于表面法线估计和多视角分割任务,实验表明其在特征匹配上有显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09321 2025-12-16 cs.CR 78%

ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data

ObliInjection: 面向多源数据LLM代理的顺序无关提示注入攻击

Reachal Wang, Yuqi Jia, Neil Zhenqiang Gong

专题命中 知识编辑与模型理解 :LLM(title,abstract)

AI总结 ObliInjection是一种针对多源数据LLM代理的新型提示注入攻击,通过顺序无关损失和顺序GCG算法有效污染输入数据以误导模型执行攻击者指定任务。

Comments To appear in NDSS 2026. For slides, see https://people.duke.edu/~zg70/code/PromptInjection.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11104 2025-12-15 cs.CV 78%

Information-driven Fusion of Pathology Foundation Models for Enhanced Disease Characterization

基于信息驱动的病理基础模型融合以增强疾病表征

Brennan Flannery, Thomas DeSilvio, Jane Nguyen, Satish E. Viswanath

机构 * Case Western Reserve University(凯斯西储大学) Department of Biomedical Engineering(生物医学工程系) Cleveland Clinic(克利夫兰诊所) Department of Pathology(病理学系) Emory University(埃默里大学) Department of Pediatrics(儿科学系) Louis Stokes VA Cleveland Medical Center(路易斯·斯托克斯退伍军人医疗中心)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 本研究提出基于信息驱动的病理基础模型融合方法,通过智能融合提升癌症分级和分期的预测性能与可解释性。

Comments 29 Pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08439 2025-12-10 cs.CV 78%

LapFM: A Laparoscopic Segmentation Foundation Model via Hierarchical Concept Evolving Pre-training

LapFM:通过分层概念演化的预训练构建腹腔镜分割基础模型

Qing Xu, Kun Yuan, Yuxiang Luo, Yuhao Zhai, Wenting Duan, Nassir Navab, Zhen Chen

机构 * School of Computer Science, University of Lincoln, UK(英国林肯大学计算机科学学院) University of Nottingham, UK(英国诺丁汉大学) University of Nottingham Ningbo China, China(中国宁波诺丁汉大学) University of Strasbourg, France(法国斯特拉斯堡大学) Technical University of Munich, Germany(德国慕尼黑技术大学) Graduate School of Information, Production and Systems, Waseda University, Japan(日本早稻田大学信息、生产与系统研究生院) Department of Gastrointestinal Surgery, The Second Qilu Hospital, Shandong University, China(中国山东大学第二齐鲁医院胃肠外科) School of Engineering and Physical Science, University of Lincoln, Lincoln LN6 7TS, UK(英国林肯大学工程与物理科学学院) Yale University, New Haven, CT 06510, USA(美国耶鲁大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 LapFM通过分层概念演化预训练方法,构建了基于腹腔镜手术图像的大型基准,实现了对复杂手术场景的高效分割和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22019 2025-12-09 cs.CV 78%

Intra-Class Probabilistic Embeddings for Uncertainty Estimation in Vision-Language Models

类内概率嵌入用于视觉-语言模型中的不确定性估计

Zhenxiang Lin, Maryam Haghighat, Will Browne, Dimity Miller

机构 * Queensland University of Technology(昆士兰理工大学)

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 本研究提出一种无需训练的后处理方法,通过类内概率嵌入提升视觉-语言模型的不确定性估计,有效检测错误预测。

Comments Accepted at the IEEE/CVF Winter Conference on Applications of Computer Vision 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19532 2025-12-08 q-bio.BM 78%

Toward the Explainability of Protein Language Models

迈向蛋白质语言模型的可解释性

Andrea Hunklinger, Noelia Ferruz

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 本文探讨了XAI在蛋白质语言模型中的应用,提出了XAI在蛋白质研究中的五个潜在角色,并呼吁推动可解释性的发展。

Comments 15 pages, 6 figures; version 4: Additional revision of the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17162 2025-12-08 cs.CR 78%

Analyzing PDFs like Binaries: Adversarially Robust PDF Malware Analysis via Intermediate Representation and Language Model

像二进制一样分析PDF:通过中间表示和语言模型实现对抗鲁棒的PDF恶意软件分析

Side Liu, Jiang Ming, Guodong Zhou, Xinyi Liu, Jianming Fu, Guojun Peng

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 通过中间表示和语言模型实现对抗鲁棒的PDF恶意软件分析,利用语义和结构特征提取提升检测性能。

Comments Accepted by ACM CCS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00493 2025-12-02 cs.CV 78%

CC-FMO: Camera-Conditioned Zero-Shot Single Image to 3D Scene Generation with Foundation Model Orchestration

CC-FMO:基于相机的零样本单图像到3D场景生成与基础模型协调

Boshi Tang, Henry Zheng, Rui Huang, Gao Huang

机构 * Tsinghua University(清华大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 CC-FMO通过结合语义感知和结构化潜在表示,实现基于相机的零样本单图像到3D场景生成,提升场景连贯性和实例保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22171 2025-12-02 cs.CV 78%

HARMONY: Hidden Activation Representations and Model Output-Aware Uncertainty Estimation for Vision-Language Models

HARMONY:隐藏的激活表示和模型输出感知的不确定性估计用于视觉-语言模型

Erum Mushtaq, Zalan Fabian, Yavuz Faruk Bakman, Anil Ramakrishna, Mahdi Soltanolkotabi, Salman Avestimehr

机构 * University of Southern California(南加州大学) Amazon AGI(亚马逊人工智能实验室)

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 HARMONY通过整合生成token、模型输出不确定性分数和隐藏表示,提升视觉-语言模型的不确定性估计性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22961 2025-12-01 cs.CV 78%

HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model

HMR3D:用于大视觉-语言模型的层次多模态表示以实现3D场景理解

Chen Li, Eric Peh, Basura Fernando

机构 * Institute of High-Performance Computing, Agency for Science, Technology and Research, Singapore(高性能计算研究所,科技研究局,新加坡) Centre for Frontier AI Research, Agency for Science, Technology and Research, Singapore(前沿人工智能研究中心,科技研究局,新加坡) College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院,南洋理工大学,新加坡)

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 HMR3D通过层次化多模态表示,结合多视图图像和文本描述,提升3D场景理解的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22664 2025-12-01 cs.CV 78%

VaMP: Variational Multi-Modal Prompt Learning for Vision-Language Models

VaMP:用于视觉-语言模型的变分多模态提示学习

Silin Cheng, Kai Han

机构 * Visual AI Lab, The University of Hong Kong(香港大学视觉人工智能实验室)

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 VaMP提出了一种变分多模态提示学习框架,通过实例条件提示和不确定性建模提升视觉-语言模型在少样本和领域泛化任务中的性能。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21614 2025-11-27 q-bio.QM 78%

Automated Protein Motif Localization using Concept Activation Vectors in Protein Language Model Embedding Space

利用蛋白质语言模型嵌入空间中的概念激活向量实现蛋白质motif自动定位

Ahmad Shamail, Claire D. McWhite

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 本文提出利用蛋白质语言模型嵌入空间中的概念激活向量实现蛋白质motif的自动化定位,通过训练线性分类器和计算内积实现高效准确的motif识别。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18416 2025-11-25 cs.CV 78%

4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation

4D-VGGT:一种具有时空意识的通用基础模型,用于动态场景几何估计

Haonan Wang, Hanyu Zhou, Haoyue Liu, Luxin Yan

机构 * National Key Lab of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(多谱信息智能处理国家实验室,人工智能与自动化学院,华中科技大学) School of Computing, National University of Singapore(计算学院,新加坡国立大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 4D-VGGT通过分而治之的时空表示方法,提升动态场景几何估计的准确性和通用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05923 2025-11-20 cs.CV 78%

Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation

Qiming Li, Zekai Ye, Xiaocheng Feng, Weihong Zhong, Weitao Ma, Xiachong Feng

专题命中 知识编辑与模型理解 :language model(title,abstract)

Comments AAAI2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09592 2025-11-14 eess.IV q-bio.QM 78%

Segment Any Tumour: An Uncertainty-Aware Vision Foundation Model for Whole-Body Analysis

Himashi Peiris, Sizhe Wang, Gary Egan, Mehrtash Harandi, Meng Law, Zhaolin Chen

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08978 2025-11-13 cs.MM cs.CV 78%

Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding

Jingtian Ma, Jingyuan Wang, Wayne Xin Zhao, Guoping Liu, Xiang Wen

机构 * School of Computer Science and Engineering, and the MOE Engineering Research Center of Advanced Computer Application Technology, Beihang University(计算机科学与工程学院,以及教育部先进计算机应用技术工程研究中心,北京航空航天大学) School of Computer Science and Engineering, the School of Economics and Management, and the MIIT Key Laboratory of Data Intelligence and Management, Beihang University(计算机科学与工程学院,经济管理学院,以及工信部数据智能与管理重点实验室,北京航空航天大学) Gaoling School of Artificial Intelligence, Renmin University of China(中关村人工智能学院,中国人民大学) DiDi Global Inc.(滴滴出行公司)

专题命中 知识编辑与模型理解 :language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25175 2025-10-30 cs.CV 78%

Test-Time Adaptive Object Detection with Foundation Model

Yingjie Gao, Yanan Zhang, Zhi Cai, Di Huang

机构 * State Key Laboratory of Complex and Critical Software Environment, Beihang University(复杂与关键软件环境国家重点实验室,北京航空航天大学) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19911 2025-10-28 cs.CV 78%

Attention! Your Vision Language Model Could Be Maliciously Manipulated

Xiaosen Wang, Shaokang Wang, Zhijin Ge, Yuyang Luo, Shudong Zhang

专题命中 知识编辑与模型理解 :language model(title,abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏