Phase structure of the Random Language Model
随机语言模型的相结构
专题命中 预训练与数据 :language model(title,abstract);large language model(abstract)
AI总结 通过双标度极限分析,揭示随机语言模型存在一系列相变,包括符号关联涌现、单符号边缘分布非均匀化以及规则使用的玻璃态冻结,并导出与大型语言模型一致的标度律。
AI 大模型
大语言模型、预训练、指令微调、后训练和语言模型应用。
随机语言模型的相结构
专题命中 预训练与数据 :language model(title,abstract);large language model(abstract)
AI总结 通过双标度极限分析,揭示随机语言模型存在一系列相变,包括符号关联涌现、单符号边缘分布非均匀化以及规则使用的玻璃态冻结,并导出与大型语言模型一致的标度律。
用于金融收益预测的预训练时间序列基础模型
专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract)
AI总结 本文在保守的金融设置下,基准测试预训练时间序列基础模型(TSFMs)与从头训练的神经基线,发现TSFMs在排名上占优,但相对于随机游走基准的预测提升微弱且稀疏,表明TSFMs作为实用先验可降低模型开发成本,但并非统计可靠的alpha生成引擎。
提示扩散模型用于零样本实例分割
机构 * Technical University of Munich (TUM)(慕尼黑工业大学) ; Munich Center of Machine Learning (MCML)(慕尼黑机器学习中心) ; Mercedes-Benz AG(梅赛德斯-奔驰集团) ; Visualais
专题命中 预训练与数据 :prompting(title);foundation model(abstract);pretraining(abstract)
AI总结 提出Prompt2Seg框架,通过空间条件化增强扩散模型,利用2D高斯或置信度图作为提示,实现零样本实例分割,在多个数据集上优于基线。
Comments Under review
基因组投毒:针对DNA基础模型的目标后门攻击
专题命中 预训练与数据 :foundation model(title,abstract);language model(abstract)
AI总结 本研究首次系统研究基因组语言模型的训练数据投毒,通过在预训练和微调阶段注入少于1%的对抗序列,可选择性破坏目标基因组上下文的生成性能,并实现条件后门攻击和下游任务分类破坏。
Comments 23 pages, double column format
没有人知道地理空间基础模型的现状
机构 * Taylor Geospatial(泰勒地理空间公司) ; Technical University of Munich(慕尼黑技术大学) ; Microsoft AI for Good Research Lab(微软AI for Good研究实验室) ; Allen Institute for AI(艾伦人工智能研究所) ; Vector Institute(向量研究所) ; Carleton University(卡尔顿大学) ; Clark University(克拉克大学) ; University of British Columbia(不列颠哥伦比亚大学) ; Arizona State University(亚利桑那州立大学)
专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract)
AI总结 本文指出地理空间基础模型缺乏标准化评估和训练协议,提出六项具体期望以促进社区共识。
CLAP: 从人类视频中学习视觉-语言-动作模型的对比潜在动作预训练
机构 * Tsinghua University(清华大学) ; Astribot ; University of Hong Kong(香港大学) ; Massachusetts Institute of Technology(麻省理工学院)
专题命中 预训练与数据 :pretraining(title,abstract);post-training(abstract)
AI总结 提出CLAP框架,通过对比学习将人类视频与机器人动作词汇对齐,利用伪标签训练VLA模型,实现从人类视频到机器人执行的有效技能迁移。
Comments The code is available at: https://github.com/LinShan-Bin/OpenCLAP
APT: 动作专家预训练提升视觉-语言-动作策略的指令泛化能力
机构 * Zhejiang University(浙江大学) ; Zhejiang Humanoid Robot Innovation Center(浙江人形机器人创新中心)
专题命中 预训练与数据 :pretraining(title,abstract);language model(abstract)
AI总结 针对连续动作专家模型对分布外语言指令泛化差的问题,提出APT两阶段训练方法,先预训练动作专家作为视觉-动作先验,再通过门控融合注入语言,显著提升泛化性能。
AnyPPG:基于心电引导的PPG基础模型,在超过10万小时记录上训练,用于全面健康分析
专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract)
AI总结 提出AnyPPG,一种基于心电引导预训练的光电容积描记(PPG)基础模型,在超10万小时数据上训练,首次开展覆盖1468种疾病表型的全表型关联研究,证明PPG可超越传统心血管应用,对307种表型(含230种非循环系统疾病)实现有效判别。
保障代码理解安全:检测代码语言模型中的自然后门漏洞
专题命中 预训练与数据 :language model(title,abstract);post-training(abstract)
AI总结 本研究系统探究代码语言模型中的自然后门漏洞,通过44种场景证明其普遍性,分析其与注入后门的差异及迁移性,并提出检测方法ScanNBT。
Comments Accepted to IEEE Transactions on Software Engineering (TSE)
统一像素标记与词标记的生成语言模型
机构 * Buaa.edu.cn(北京航空航天大学)
专题命中 预训练与数据 :language model(title,abstract);pretraining(abstract)
AI总结 本文提出一种统一像素标记和词标记的生成语言模型,通过引入图像无监督预训练、颜色折叠、全局条件注意力近似等方法,提升模型在图像细节识别上的能力,实验表明该模型在小模型和有限数据下仍表现优异。
Comments 13 pages, 6 figures
我们需要多少模型?遥感基础模型中的冗余与可瘦身性
专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract)
AI总结 通过后验瘦身(均匀减少编码器Transformer块宽度)评估8个遥感基础模型的表示冗余,发现遥感模型在激进宽度缩减下仍保持69%-109%相对精度,而自然图像预训练模型性能急剧下降,表明遥感模型存在冗余编码且可有效瘦身。
将视觉语言模型的预训练扩展到一千亿数据
机构 * Google DeepMind(谷歌DeepMind)
专题命中 预训练与数据 :language model(title,abstract);pretraining(abstract)
AI总结 本文通过实验探究将视觉语言模型预训练数据扩展到一千亿规模的效果,发现传统基准性能饱和,但文化多样性任务和低资源语言受益显著,并指出质量过滤可能减少文化多样性。
Comments v2: CVPR Findings'26
机器遗忘在视觉-语言模型中的鲁棒性研究
机构 * Xiamen University(厦门大学)
专题命中 预训练与数据 :language model(title,abstract);prompting(abstract)
AI总结 本文首次系统调查了视觉-语言模型机器遗忘的鲁棒性,通过提出三种攻击范式揭示现有方法往往隐藏而非彻底移除目标知识。
InHabit: 利用图像基础模型实现可扩展的3D人体放置
机构 * Bosch Center of Artificial Intelligence(博世人工智能中心) ; Max Planck Institute for Informatics(马克斯·普朗克信息学院)
专题命中 预训练与数据 :foundation model(title,abstract);language model(abstract)
AI总结 提出InHabit方法,通过渲染-生成-提升流程利用2D基础模型知识自动生成3D场景中与几何一致的人体交互数据,并构建大规模数据集InHabitants,显著提升3D人体-场景重建和接触估计性能。
利用语言模型在多语言场景下生成日志语句:我们走了多远?
专题命中 预训练与数据 :language model(title,abstract);large language model(abstract)
AI总结 本文通过构建包含五种编程语言的多语言基准,比较评估了三种最先进的日志语句生成方法和五个大型语言模型,发现UniLog方法在多语言环境中表现最佳,并揭示了不同语言间日志生成难度的差异源于日志插入分布和语言特定日志习惯,指出单纯扩大模型规模或训练数据量不足以解决多语言日志生成问题,需针对目标语言特性设计方法。
C3P: 对比启动子-蛋白质预训练产生捕捉细菌基因调控的表征
专题命中 预训练与数据 :pretraining(title,abstract);language model(abstract)
AI总结 提出对比启动子-蛋白质预训练(C3P),通过对齐启动子与对应蛋白质来学习细菌调控序列表征,在调控注释推断和零样本共调控基因检索中显著优于基因组语言模型。
从解剖到疾病表型的通用CT表示:通过聚合预训练
机构 * Wallace H. Coulter Department of Biomedical Engineering, Georgia Institute of Technology and Emory University(沃森·H·库勒生物医学工程系,佐治亚理工学院和埃默里大学) ; Department of Radiation Oncology and Winship Cancer Institute, Emory University(放射肿瘤学系和Winship癌症研究所,埃默里大学) ; Department of Electrical and Computer Engineering, Duke University(电气与计算机工程系,杜克大学) ; Department of Computer Science and Informatics, Emory University(计算机科学与信息学系,埃默里大学) ; Department of Materials Science & Engineering, Nuclear Engineering Program, University of Florida(材料科学与工程系、核工程项目,佛罗里达大学)
专题命中 预训练与数据 :pretraining(title,abstract);foundation model(abstract)
AI总结 提出FlexiCT系列CT基础模型,通过三阶段聚合连续预训练(二维轴向、三维解剖、报告引导语义对齐)统一CT分析,在分割、分类、配准、视觉语言理解和临床检索等任务上达到或超越专用模型,并捕获与肿瘤分期相关的影像特征。
面向临床的大脑MRI基础模型:来自FOMO25挑战赛的发现
机构 * organization= Department of Computer Science, University of Copenhagen , city= Copenhagen , country= Denmark ; organization= Pioneer Centre for AI , city= Copenhagen , country= Denmark ; organization= Copenhagen Research Centre for Biological ; Precision Psychiatry, Mental Health Centre Copenhagen, Copenhagen University Hospital , region= Capital Region of Denmark , city= Copenhagen , country= Denmark ; organization= Athinoula A. Martinos Center for Biomedical Imaging, Massachusetts General Hospital ; Harvard Medical School , city= Boston , state= Massachusetts , country= USA ; Artificial Intelligence Laboratory, Massachusetts Institute of Technology , city= Boston , state= Massachusetts , country= USA ; organization= Johns Hopkins University , city= Baltimore , state= Maryland , country= USA ; organization= Radiological AI Testcenter (RAIT) , region= Capital Region of Denmark , city= Copenhagen , country= Denmark ; organization= Copenhagen University Hospital, Rigshospitalet , region= Capital Region of Denmark , city= Copenhagen , country= Denmark ; organization= Copenhagen University Hospital, Bispebjerg \& Frederiksberg Hospital , region= Capital Region of Denmark , city= Copenhagen , country= Denmark ; organization= Department of Clinical Medicine, Faculty of Health ; Medical Sciences, University of Copenhagen , city= Copenhagen , country= Denmark ; organization= Division of Medical Image Computing, German Cancer Research Center (DKFZ) , city= Heidelberg , country= Germany ; organization= University of British Columbia , city= Vancouver , state= British Columbia , country= Canada ; organization= Hawkes Institute, Department of Computer Science, University College London , city= London , country= United Kingdom ; Lung Institute, Faculty of Medicine, Imperial College London , city= London , country= United Kingdom ; organization= Department of Applied Mathematics, Technical Medical Centre, University of Twente , city= Enschede , country= Netherlands ; organization= IISLAB, Technical University of Košice , city= Košice , country= Slovakia ; organization= 2nd Department of Internal Medicine, Pavol Jozef Safarik University ; L Pasteur University Hospital , city= Košice , country= Slovakia ; organization= Fudan University , city= Shanghai , country= China ; organization= Shenzhen Technology University , city= Shenzhen , country= China ; organization= Department of Radiology, Lausanne University Hospital ; University of Lausanne , city= Lausanne , country= Switzerland ; organization= Louvain Neuroinflammation Imaging Lab (NIL), Université Catholique de Louvain , city= Brussels , country= Belgium ; organization= University of Applied Sciences ; organization= CIBM Center for Biomedical Imaging , city= Lausanne , country= Switzerland ; organization= Department of Radiation Oncology (Maastro), GROW Research Institute for Oncology ; Reproduction, Maastricht University Medical Centre+ , city= Maastricht , country= The Netherlands ; organization= Department of Biomedical Engineering, Medical Image Analysis, Eindhoven University of Technology , city= Eindhoven , country= The Netherlands ; organization= Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences , city= Shenzhen , country= China ; organization= McGill University ; Mila - Quebec AI Institute , city= Montreal , country= Canada ; organization= Hotchkiss Brain Institute ; Department of Radiology, University of Calgary , city= Calgary , state= Alberta , country= Canada ; organization= Department of Radiology, University of Calgary , city= Calgary , state= Alberta , country= Canada ; organization= Alberta Children's Hospital Research Institute, Department of Clinical Neuroscience, University of Calgary , city= Calgary , state= Alberta , country= Canada ; organization= The Wallace H. Coulter Department of Biomedical Engineering, Georgia Tech ; Emory University , city= Atlanta , state= Georgia , country= USA ; organization= SGGS College of Engineering ; organization= Seoul National University , city= Seoul , country= South Korea ; organization= The D-Lab, Department of Precision Medicine, GROW Research Institute for Oncology ; Reproduction, Maastricht University , city= Maastricht , country= The Netherlands ; organization= Artificial Intelligence in Medicine (AIM) Program, Mass General Brigham, Harvard Medical School , city= Boston , state= Massachusetts , country= USA ; Nuclear Medicine, CARIM \& GROW, Maastricht University , city= Maastricht , country= The Netherlands ; organization= Department of Radiation Oncology, Dana-Farber Cancer Institute, Brigham ; Women’s Hospital, Harvard Medical School , city= Boston , state= Massachusetts , country= USA ; Learning Group, Heidelberg University Hospital , city= Heidelberg , country= Germany
专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract)
AI总结 针对临床脑MRI数据异质且标注成本高的问题,FOMO25挑战赛通过自监督预训练(FOMO60K数据集)评估了16个团队的基础模型,发现自监督预训练能提升域迁移泛化性,但不同任务需不同预训练目标,且模型规模扩展收益有限。
弥合冷启动缺口:利用大语言模型生成合成数据用于Airbnb的自然语言搜索
专题命中 预训练与数据 :LLM(title);large language model(abstract);language model(abstract)
AI总结 本文提出利用大语言模型生成合成查询和标签的方法,解决自然语言搜索系统中冷启动问题,通过生成真实用户数据的过渡来提升模型训练和评估效果。
SPA-MAE:一种基于物理的CSI基础模型用于无线物理层
专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract)
AI总结 本文提出了一种基于物理的CSI基础模型SPA-MAE,通过结合适应性的MAE骨干网络和通道知识,提升了无线物理层任务的泛化能力,实验表明其在参数较少的情况下,在低SNR和有限数据条件下表现更优。
预训练写入,对齐读取:Transformer权重空间的不对称性
机构 * Intuition Machines
专题命中 预训练与数据 :pretraining(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 研究揭示了预训练和对齐在Transformer权重空间中的不对称性,通过分析权重变化在残差流激活子空间和预测子空间中的对齐情况,发现读路径权重集中于注意力输入激活的主方向,而写路径权重在预测子空间中保持各向同性。
通过合成神经符号监督学习结构化机器人策略
机构 * University of Padova, Dept. of Information Engineering(帕多瓦大学信息工程系) ; Fraunhofer Italia Research(弗劳恩霍夫意大利研究所) ; Polytechnic of Bari Dept. of Electrical and Information Engineering(巴里理工学院电气与信息工程系)
专题命中 预训练与数据 :language model(title,abstract);foundation model(abstract)
AI总结 本文提出通过合成神经符号监督方法,利用视觉语言模型生成结构化机器人策略,结合多模态感知与符号控制,实现高维学习与符号控制的结合。
手写解码作为EEG基础模型中的一个具有挑战性的运动任务
专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract)
AI总结 本文提出将手写解码作为EEG基础模型的挑战性运动任务,发现现有数据集可能存在问题,并引入更严谨的评估数据集,显示基础模型在手写解码任务中不如专门模型表现优异。
BrainAnytime: 基于解剖结构的跨模态预训练用于脑图像分析,支持任意模态可用性
机构 * Department of Biomedical Engineering, The Hong Kong Polytechnic University, Hong Kong SAR, China(生物医学工程系,香港理工大学,香港特别行政区,中国) ; Department of Technology Management for Innovation, The University of Tokyo, Japan(创新技术管理系,东京大学,日本) ; Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University, Hong Kong SAR, China(数据科学与人工智能系,香港理工大学,香港特别行政区,中国)
专题命中 预训练与数据 :pretraining(title,abstract);foundation model(abstract)
AI总结 BrainAnytime通过跨模态蒸馏和解剖引导课程掩码,在共享的3D掩码自动编码器中学习MRI与PET的结构-分子对应关系,实现对任意模态可用性的统一预训练,提升多任务性能。
Comments Early accepted by MICCAI 2026
MaskTab: 适用于工业分类的可扩展遮蔽表格预训练与缩放定律及蒸馏
机构 * Zhejiang University(浙江大学) ; MyBank, Ant Group(蚂蚁集团MyBank)
专题命中 预训练与数据 :pretraining(title);foundation model(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 MaskTab通过结合监督预训练和MoE增强损失,提升工业表格数据的分类性能,实现更高的AUC和KS值,并在轻量化模型中保持鲁棒性。
通过FCP释放可扩展的上下文并行性以实现基础模型预训练
专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract)
AI总结 本文提出FCP,一种灵活的上下文并行范式,通过块级粒度划分和调度序列,实现高效计算和负载平衡,提升基础模型预训练的可扩展性。
Comments Accepted by MLSys 2026
超越ViT标记:面向细胞级密集预测的掩码扩散预训练卷积病理基础模型
机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) ; Research Institute of Tsinghua, Pearl River Delta(清华大学 Pearl River Delta 研究院) ; The Chinese University of Hong Kong, ShenZhen(香港中文大学(深圳))
专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract)
AI总结 本文提出ConvNeXt Masked-Diffusion模型,通过卷积生成预训练框架提升病理图像细胞级密集预测性能,实验表明其在有限标注条件下表现更优,优于现有ViT模型和端到端分割方法。
通过补丁标记利用视觉基础模型特征进行AI生成图像检测
机构 * Mannheim University of Applied Sciences(曼海姆应用科学大学)
专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract)
AI总结 本文通过评估多种视觉基础模型家族,提出利用可调注意池化(TAP)改进分类器头,提升AI生成图像检测性能,发现最佳模型在准确率上超越CLIP超过12%。
Comments This paper has been accepted at IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2026
随机底限:在语言模型令牌分布中测量固有非随机性
机构 * Institute of Computer Science(计算机科学研究所) ; Faculty of Mathematics and Computer Science(数学与计算机科学学院) ; Jagiellonian University(雅盖隆大学)
专题命中 预训练与数据 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文通过系统分析31200次生成数据,引入熵偏差(ED)衡量语言模型令牌分布与均匀分布的KL散度,发现即使在中性提示下,Transformer仍表现出约0.30的ED,表明非随机性主要源于模型权重而非上下文。
Comments 13 pages, 4 figures, 5 tables
JoyAI-RA 0.1:一种用于机器人自主性的基础模型
机构 * Joy Future Academy(京东探索研究院)
专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract)
AI总结 本文提出JoyAI-RA,一种面向通用机器人操作的视觉-语言-动作基础模型,通过多源多层级预训练框架提升跨身体行为学习能力,在仿真和现实任务中表现优异。