arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-06-30 至 2026-06-30 共收录 47 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 47 篇

2606.30571 2026-06-30 cs.LG cs.CL 92%

Attractor States Emerge in Multi-Turn LLM Conversations

多轮LLM对话中吸引子状态的出现

Ting-Wen Ko, Jonas Geiping

机构 * Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) ELLIS Institute Tübingen(图宾根ELLIS研究所) Tübingen AI Center(图宾根人工智能中心)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 研究多轮LLM对话中是否出现吸引子行为,通过自对弈和混合对弈辩论发现模型特定的吸引子不对称地影响对话伙伴的风格和立场。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29545 2026-06-30 cs.CL 91%

AURORA: Asymmetry and Update-Induced Rotation for Robust Hallucination Detection in Large Language Models

AURORA:用于大型语言模型中鲁棒幻觉检测的不对称性与更新诱导旋转

Zishuai Zhang, Hainan Zhang, Zhiming Zheng

机构 * School of Artificial Intelligence, Beihang University, China(北京航空航天大学人工智能学院) Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University, China(北京航空航天大学未来区块链与隐私计算先进创新中心)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(summary_cn,abstract_cn);分类 cs.CL

AI总结 提出AURORA框架,利用权重梯度动态(不对称性和旋转比)检测LLM幻觉,跨模型和数据集表现鲁棒。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28737 2026-06-30 cs.CL cs.AI 90%

5ting at SemEval-2026 Task 8: Strong End-to-End Multi-Turn RAG via LLM-Based Reranking and Faithfulness Control

5ting在SemEval-2026任务8中:基于LLM重排序和忠实性控制的强端到端多轮RAG

Thien-Qua-T-Nguyen, Chi Hoang, Nguyen Tran, Tri Le, Khanh Truong, Chinh Trong Nguyen

机构 * University of Information Technology, Ho Chi Minh City, Vietnam(信息技术大学,胡志明市,越南) Vietnam National University Ho Chi Minh City, Ho Chi Minh City, Vietnam(越南胡志明市国家大学,胡志明市,越南)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);分类 cs.CL、cs.AI

AI总结 提出5ting系统,结合BGE-M3稠密检索、FAISS索引、双查询合并检索和LLM重排序,通过角色分离生成约束于检索证据,解决多轮RAG中的上下文漂移和幻觉问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28798 2026-06-30 cs.AI stat.AP 90%

Primary ICD Category Prediction using LLM-based Probing

基于LLM探针的主要ICD类别预测

Chengyuan Liu, Xinyue Zhang, Yao Li, Guanting Chen

机构 * Department of Statistics, Pennsylvania State University(宾夕法尼亚州立大学统计学系) Department of Biostatistics, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校生物统计学系) Department of Statistics and Operations Research, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校统计学与运筹学系)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究利用冻结的医学大语言模型表示作为共享嵌入空间,通过线性探针融合结构化变量和临床叙述,实现多模态主要诊断类别预测,在MIMIC-IV上达到87.69%的严格准确率。

Comments 9 pages, 2 figures. Supplementary materials provided as an ancillary file

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03335 2026-06-30 cs.CL 89%

Compressed Sensing for Capability Localization in Large Language Models

压缩感知在大语言模型能力定位中的应用

Anna Bair, Yixuan Even Xu, Mingjie Sun, J. Zico Kolter

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL

AI总结 研究通过压缩感知方法识别大语言模型中特定能力依赖的稀疏注意力头,发现关闭少量头可显著降低特定能力表现,揭示了模型模块化组织原则。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28358 2026-06-30 cs.IR cs.AI cs.CL 88%

How Do LLMs Cite? A Mechanistic Interpretation of Attribution in Retrieval-Augmented Generation

LLM如何引用?检索增强生成中归因的机制解释

Ian van Dort, Maria Heuss

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 知识编辑与模型理解 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 通过激活修补方法,发现LLM的引用机制并非单一组件,而是由注意力头和MLP层组成的分布式“归因集成”,调控这些组件可修复大部分错误引用。

Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution is published in Advances in Information Retrieval, ECIR 2026, Lecture Notes in Computer Science, vol. 16485, pp. 458-473, and is available online at https://doi.org/10.1007/978-3-032-21324-2_35

Journal ref Advances in Information Retrieval, ECIR 2026. Lecture Notes in Computer Science, vol. 16485, pp. 458-473. Springer, Cham (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08831 2026-06-30 cs.AI 新提交 88%

Inference-Time Conformal Reasoning with Valid Factuality Control for Large Language Models

面向大语言模型的推理时保形推理与有效事实性控制

Ting Wang, Yuanjie Shi, Yan Yan, Huan Zhang

机构 * Machine Learning, ICML(机器学习,国际机器学习大会)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 提出推理时保形推理框架,将保形预测集成到推理图生成中,通过图级不确定性校准生成停止阈值,实现有效事实性控制。

Comments Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29490 2026-06-30 cs.LG cs.AI 87%

Reported Confidence in LLMs Tracks Commitment More Than Correctness

LLM中的报告置信度追踪承诺而非正确性

Dharshan Kumaran

机构 * Google DeepMind(谷歌DeepMind)

专题命中 知识编辑与模型理解 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本研究通过两阶段弃权范式,发现LLM的言语置信度预测弃权决策远优于预测答案正确性,而令牌对数概率则相反,表明言语置信度是内部承诺准备状态的行为输出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12418 2026-06-30 cs.CR cs.CL cs.LG 84%

Sparse Autoencoders are Capable LLM Jailbreak Mitigators

稀疏自编码器是LLM劫持攻击缓解器

Yannick Assogba, Jacopo Cortellazzi, Javier Abad, Pau Rodriguez, Xavier Suau, Arno Blaas

专题命中 知识编辑与模型理解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文提出基于SAE的CC-Delta方法,通过比较带有和无劫持上下文的有害请求的token级表示,识别劫持相关的稀疏特征,从而在稀疏SAE特征空间中实现更优的安全-效用权衡。

Comments Accepted at the Mechanistic Interpretability Workshop, ICML 2026. 31 pages, 20 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29600 2026-06-30 cs.CV cs.AI 83%

One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models

一场景,两深度:探究单目基础模型中的几何歧义性

Xiaohao Xu, Feng Xue, Xiang Li, Haowei Li, Shusheng Yang, Tianyi Zhang, Matthew Johnson-Roberson, Xiaonan Huang

机构 * University of Michigan(密歇根大学) Carnegie Mellon University(卡内基梅隆大学) New York University(纽约大学) Vanderbilt University(范德比大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);prompting(abstract);分类 cs.AI

AI总结 本文提出稀疏双层序数基准MD-3k,用于测量单目深度基础模型的深度层偏好和多层空间关系准确性,发现不同模型对同一分层几何结构有不同解析,且拉普拉斯视觉提示可改变冻结模型的输出层。

Comments 49 pages, 25 figures; Accepted by European Conference on Computer Vision (ECCV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30578 2026-06-30 cs.CL cs.LG 82%

Uncertainty-Aware Generation and Decision-Making Under Ambiguity

模糊性下的不确定性感知生成与决策

Nico Daheim, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, Technical University of Darmstadt(普遍知识处理实验室(UKP实验室),计算机科学系,达姆施塔特技术大学) National Research Center for Applied Cybersecurity ATHENE, Germany(应用网络安全国家研究中心ATHENE,德国)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 基于贝叶斯决策理论和风险规避决策,提出不确定性感知算法用于辅导和同行评审任务,通过共形预测提供策略和分数的保证,实验表明贝叶斯方法优于风险规避规则。

Comments Code available under https://github.com/UKPLab/arXiv2026-uncertainty-aware

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04385 2026-06-30 cs.CL cs.AI cs.LG 82%

How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models

对齐路由:在语言模型中本地化、扩展和控制策略电路

Gregory N. Frank

机构 * Independent Researcher(独立研究者)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究通过本地化策略路由机制,探讨在语言模型中扩展和控制策略电路的方法,发现路由机制在安全性和性能上的关键作用。

Comments Code and data: https://github.com/gregfrank/how-alignment-routes. Accepted at the Mechanistic Interpretability Workshop at the 43rd International Conference on Machine Learning (ICML), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29049 2026-06-30 cs.LG 81%

MOSAIC: Orchestrating Collaborative Knowledge Tracing with Hierarchical Semantic Alignment

MOSAIC: 通过层次语义对齐编排协作知识追踪

Xinjin Li, Mengyue Wang, Yuzhen Lin, Pengbin Feng, Ziqi Sha, Yeyang Zhou, Yu Ma

机构 * Columbia University(哥伦比亚大学) University of California, Berkeley(加州大学伯克利分校) School of Information Systems and Management, Carnegie Mellon University(信息系统与管理学院,卡内基梅隆大学) Department of Mathematics, University of Southern California(数学系,南加州大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Computer Science Department, UC San Diego(计算机科学系,UCSD)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);分类 cs.LG

AI总结 提出MOSAIC框架,利用冻结LLM生成动态嵌入和层次预测提示,结合跨粒度一致性目标,在协作知识追踪中实现多粒度掌握估计,在多个数据集上取得SOTA。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28770 2026-06-30 cs.AI 81%

Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions

LLMs人格的机械论分析:通过潜在特征干预引导人格

David Courtis, Ting Hu

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出一种机械可解释性方法,通过稀疏自编码器和对比激活分析识别残差流中的潜在方向,并施加加法干预向量来增强目标OCEAN人格特质,同时保持语言建模性能。

Comments Written in 2024; submitted to arXiv 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.24050 2026-06-30 cs.LG stat.ML 81%

A Mechanistic Study of Transformers Training Dynamics

Transformer训练动态的机制研究

Ambroise Odonnat, Wassim Bouaziz, Vivien Cabannes

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);foundation model(abstract);pretraining(abstract)

AI总结 本文通过可控实验研究Transformer训练动态,发现梯度下降可实现聚类头解决稀疏模块加法任务,并揭示训练过程中两阶段学习及归一化层高曲率导致的损失尖峰现象。

Comments Accepted at ICML 2026 Mechanistic Interpretability workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23353 2026-06-30 cs.LG cs.AI 81%

SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport

SOTAlign:通过最优传输实现半监督的单模态视觉与语言模型对齐

Simon Roschmann, Paul Krzakala, Sonia Mazelet, Quentin Bouniot, Zeynep Akata

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出SOTAlign框架,通过少量配对数据和大量未配对数据实现视觉与语言模型的半监督对齐,利用最优传输理论提升对齐效果,优于传统监督和半监督方法。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29865 2026-06-30 math.RT 80%

On the structure of the singular triplet monoid and its virtual extension

关于奇异三元组幺半群及其虚拟扩展的结构

Carmen Caprau, Mohamad N. Nasser

专题命中 知识编辑与模型理解 :SLM(summary_cn,abstract)

AI总结 本文引入与n股三元组群L_n相关的奇异三元组幺半群SLM_n及其虚拟扩展VSLM_n,通过生成元和关系定义,并开发了k-局部型和Φ-型两种表示扩展方法,证明所有2-局部表示均可扩展,并应用于具体表示μ。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29406 2026-06-30 q-fin.RM math.OC 80%

Adaptive AI Delegation under Uncertainty: A Bayesian Governance Policy for Sequential Decision Authority

不确定性下的自适应AI授权:序贯决策权限的贝叶斯治理策略

Matthew Francis Dixon

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 针对组织在不确定性下动态分配AI决策权限的问题,提出基于贝叶斯推断的治理感知POMDP框架,通过序贯优化实现自适应授权,实验表明该方法在异构AI质量场景下优于五种基准策略。

Comments 48 manuscript pages, 17 figures, and 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06748 2026-06-30 cs.CL cs.AI cs.LG 新提交 80%

Evidence Graph Consistency in Retrieval-Augmented Generation: A Model-Dependent Analysis of Hallucination Detection

检索增强生成中的证据图一致性:基于模型的幻觉检测分析

Jianru Shen

机构 * University of Montana(蒙大拿大学)

专题命中 知识编辑与模型理解 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出证据图一致性(EGC)框架,通过构建局部证据图并计算五种结构一致性指标检测幻觉,发现不同模型族间一致性特征方向相反,表明嵌入图一致性不能作为模型无关的检测信号。

Comments Accepted at the International Conference on Advanced Machine Learning and Data Science; to appear in the IEEE Xplore proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21815 2026-06-30 cs.CV cs.LG 79%

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models

高熵标记作为视觉-语言模型中的多模态失败点

Mengqi He, Xinyu Tian, Xin Shen, Jinhong Ni, Shu Zou, Zhaoyuan Yang, Jing Zhang

机构 * The Australia National University(澳大利亚国立大学) The University of Queensland(昆士兰大学) GE research(GE研究)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.LG

AI总结 本研究揭示视觉-语言模型中约20%的高熵标记集中了不成比例的对抗性影响,并提出基于熵引导的稀疏攻击方法(EGA),实现高攻击成功率与有害率。

Comments 19 Pages,11 figures,8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29069 2026-06-30 cs.AI cs.CL cs.CV 79%

Low-cost concept-based localized explanations: How far can we get with training-free approaches?

低成本基于概念的可解释性:无训练方法能走多远?

Darian Fernández-Gutiérrez, Rafael Bello, Marilyn Bello, Natalia Díaz-Rodríguez

机构 * Dept. of Computer Science and Artificial Intelligence, University of Granada (UGR)(计算机科学与人工智能系,格拉纳达大学(UGR))

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.CL、cs.AI

AI总结 本文提出零样本概念命名协议,利用中等规模多模态大模型对局部区域进行概念标注,无需训练即可实现62%-88%的物体级精确匹配,为低成本可解释AI提供新思路。

Comments 6 pages, 2 figures, 4 tables. Accepted at the 2026 IEEE International Conference on Artificial Intelligence (CAI), 8-10 May 2026, Granada, Spain. Code: https://github.com/darianfgUgr/CoNa

Journal ref 2026 IEEE International Conference on Artificial Intelligence (CAI), Granada, Spain, 2026, pp. 1405-1410

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30020 2026-06-30 cs.CV 78%

Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning

病理基础模型中的不确定性估计:基于深度互学习

Gbègninougbo Aurel Davy Tchokponhoue, Sevda Öğüt, Ali Idri, Dorina Thanou, Pascal Frossard

机构 * UM6P(摩洛哥穆萨-阿卜杜勒-阿齐兹大学) EPFL(瑞士联邦理工学院) UM5(摩洛哥穆萨-阿卜杜勒-阿齐兹大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract)

AI总结 提出DICE框架,通过集成多个冻结的病理基础模型并利用深度互学习对齐,以模型分歧作为不确定性代理,实现可靠的不确定性估计、异常定位,并在分类、校准和定位上匹配或超越现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29441 2026-06-30 cs.CR cs.AI cs.CL cs.ET cs.LG 75%

Closing the Activation-Cone Blind Spot: Response-Time Probing and Unified Defense

关闭激活锥盲点:响应时间探测与统一防御

Subhadip Mitra

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 针对大语言模型推理时安全方法,发现提示时激活防御对预填充攻击存在结构性盲点,提出响应时间探测(线性探针)结合停止机制,将预填充攻击成功率降至0,并与AlphaSteer组合实现正交防御。

Comments 27 pages, 12 figures, 18 tables. Code and data: https://github.com/bassrehab/response-time-probing

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25013 2026-06-30 cs.CL cs.AI cs.LG 75%

Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers

注意力-only变换器中间接对象识别最小电路的涌现

Rabin Adhikari

机构 * Saarland University(萨尔大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究通过从头训练小型注意力-only变换器,在符号化的间接对象识别任务中发现,单层模型仅需两个注意力头即可实现完美准确率,揭示了任务特定训练如何诱导可解释的最小电路。

Comments Published at ACL (Volume 4: Student Research Workshop) ISBN: 979-8-89176-393-7 URL: https://aclanthology.org/2026.acl-srw.4

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28667 2026-06-30 cs.CL 74%

Phonological Perception of Sign Language Models

手语模型的音系感知

Kayo Yin, Jessica Carter, Alex Xijie Lu, Annemarie Kocab

机构 * University of California, Berkeley(加州大学伯克利分校) Johns Hopkins University(约翰霍普金斯大学) Microsoft Research(微软研究院)

专题命中 知识编辑与模型理解 :language model(title);分类 cs.CL

AI总结 本研究通过最小对测试和表征对齐评估手语识别模型的音系感知能力,发现模型具有涌现音系敏感性但存在架构权衡:姿态模型对手形敏感,像素模型对位置敏感。

Comments Accepted to CogSci 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30609 2026-06-30 cs.LG cs.AI 73%

C$^{2}$R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders

C$^{2}$R: 跨样本一致性正则化缓解稀疏自编码器中的特征分裂与吸收

Haoran Jin, Xiting Wang, Shijie Ren, Hong Xie, Defu Lian

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 提出跨样本一致性正则化(C$^2$R),通过惩罚方向相似潜变量的共激活,缓解稀疏自编码器中的特征分裂与吸收问题,提升潜变量可解释性而不损失重构保真度。

Comments 24 pages, 6 figures. Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29357 2026-06-30 cs.CV cs.AI cs.LG 73%

Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking

利用视觉语言模型动态解析和更新自然语言规范以实现鲁棒的视觉语言跟踪

Xiao Wang, Liye Jin, Dan Xu, Yuehang Li, Lan Chen, Yaowei Wang, Yonghong Tian, Jin Tang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳分校) Peng Cheng Laboratory, Shenzhen(鹏城实验室深圳分部) Shenzhen Graduate School, Peking University(北京大学深圳研究生院)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 提出语言依赖解析机制提取跟踪核心成分,并利用Qwen-VL进行成分感知的文本更新,在多个基准上取得一致优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29463 2026-06-30 cs.CV 67%

CellDETR: A Detection-Guided Framework for Scalable Cell Representation Learning from Histopathology Images

CellDETR: 一种面向组织病理学图像的可扩展细胞表示学习的检测引导框架

Shikang Zhang, Guojun Li, Yicong Mao, Chulin Sha

专题命中 知识编辑与模型理解 :foundation model(abstract);pretraining(abstract)

AI总结 提出基于Deformable DETR的检测引导框架CellDETR,通过位置特征解耦和框约束注意力机制实现可扩展的细胞表示学习,在细胞分类任务上超越现有方法,并支持无标签数据的预训练和跨数据集迁移。

Comments 12pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20606 2026-06-30 cs.CV 67%

Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking

探查并利用视频扩散变换器特征以实现稳健的点跟踪

Soowon Son, Honggyu An, Jisu Nam, Hyunah Ko, Chaehyun Kim, Dahyun Chung, Siyoon Jin, Jung Yi, Junhwa Hur, Seungryong Kim

机构 * KAIST AI(韩国科学技术院人工智能) Google DeepMind(谷歌DeepMind)

专题命中 知识编辑与模型理解 :foundation model(abstract);pretraining(abstract)

AI总结 本文探讨了视频扩散变换器在点跟踪中的优势,提出DiTracker框架,通过整合视频DiT特征提升跟踪鲁棒性,实验证明其在挑战性场景下表现优异。

Comments Project Page: https://cvlab-kaist.github.io/DiTracker/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23220 2026-06-30 cs.CL cs.LG 62%

Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders

模型方向,而非词语:使用稀疏自编码器的机制性主题模型

Carolina Zheng, Nicolas Beltran-Velez, Sweta Karlekar, Claudia Shi, Achille Nazaret, Asif Mallik, Amir Feder, David M. Blei

机构 * Columbia University(哥伦比亚大学) Google Research(谷歌研究院) Independent(独立研究者)

专题命中 知识编辑与模型理解 :LLM(abstract);分类 cs.CL、cs.LG

AI总结 本文提出机制性主题模型(MTMs),利用稀疏自编码器学习可解释特征,以揭示深层概念主题,并通过topic judge评估框架验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏