arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 138791 信号源:cs.CL, cs.AI, cs.LG

1. 预训练与数据 12393 篇

2604.17739 2026-05-08 cs.LG cs.CL 81%

Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model

用免费的8B语言模型完全模拟环境来民主化工具学习

Chenming Tang, Hsiu-Yuan Huang, Weijie Liu, Junqiang Zheng, Saiyong Yang, Yunfang Wu

机构 * National Key Laboratory for Multimedia Information Processing, Peking University(北京大学多媒体信息处理国家重点实验室) School of Computer Science, Peking University(北京大学计算机科学学院) LLM Department, Tencent(腾讯LLM部门)

专题命中 预训练与数据 :language model(title,abstract);分类 cs.CL、cs.LG

AI总结 本文提出TRUSTEE方法,利用免费开源的8B语言模型模拟动态环境,结合自适应课程学习机制,有效训练工具调用代理,实验证明其在多数情况下优于需额外资源的基线方法。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21106 2026-05-08 cs.LG cs.CL 81%

How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models

循环一次值有多大?循环语言模型的等深度缩放定律

Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis

机构 * Chair for AI in Healthcare and Medicine, Technical University of Munich(慕尼黑技术大学人工智能在医疗和医学中的主任) Department of Computing, Imperial College London(伦敦帝国理工学院计算机系) Munich Center for Machine Learning (MCML), Germany(慕尼黑机器学习中心(MCML)) Hasso Plattner Institute for Digital Engineering, University of Potsdam, Germany(波茨坦大学数字工程霍普夫研究所)

专题命中 预训练与数据 :language model(title);pretraining(abstract);分类 cs.CL、cs.LG

AI总结 研究通过等深度预训练测量循环变换器中一次循环的价值,推导出缩放定律并发现循环等价指数φ=0.46,表明共享循环比独特块更差,展示了φ作为诊断工具的实用性。

Comments v3: substantially refined framing + minor corrections v2: added case studies on truncated-BPTT and hyperconnections

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02534 2026-05-08 cs.CL cs.AI 81%

ANGOFA: Leveraging OFA Embedding Initialization and Synthetic Data for Angolan Language Model

ANGOFA: 利用OFA嵌入初始化和合成数据进行安哥拉语言模型

Osvaldo Luamba Quinjica, David Ifeoluwa Adelani

机构 * Masakhane NLP Department of Computer Science University College London(计算机科学系伦敦大学学院) University College London(伦敦大学学院)

专题命中 预训练与数据 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出四个针对安哥拉语言定制的预训练语言模型,采用多语言自适应微调方法,通过有意识的嵌入初始化和合成数据提升模型性能,优于SOTA AfroXLMR-base和OFA模型。

Comments Accepted at AfricaNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00933 2026-05-05 cs.LG cs.AI 81%

CGM-JEPA: Learning Consistent Continuous Glucose Monitor Representations via Predictive Self-Supervised Pretraining

CGM-JEPA:通过预测自监督预训练学习一致的连续葡萄糖监测表示

Hada Melino Muhammad, Zechen Li, Flora Salim, Ahmed A. Metwally

机构 * University of New South Wales(新南威尔士大学) Google Research(谷歌研究)

专题命中 预训练与数据 :pretraining(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出CGM-JEPA框架,通过预测掩码的潜在表示来提升多模态数据的迁移能力,实验显示其在不同场景下均优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00937 2026-05-01 cs.RO cs.AI cs.CV cs.LG 81%

CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining

CLAMP: 基于对比学习的3D多视角动作条件机器人操控预训练

I-Chun Arthur Liu, Krzysztof Choromanski, Sandy Huang, Connor Schenck

机构 * Google DeepMind(谷歌DeepMind) University of Southern California(南加州大学)

专题命中 预训练与数据 :pretraining(title,abstract);分类 cs.AI、cs.LG

AI总结 CLAMP通过3D点云和机器人动作进行预训练,利用对比学习提升机器人操控精度与效率,优于现有基线方法。

Comments Accepted to the Robotics: Science and Systems (RSS) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20012 2026-04-23 cs.CV cs.AI cs.CL 81%

EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training

EmbodiedMidtrain: 通过中训练弥合视觉语言模型与视觉语言动作模型之间的差距

Yiyang Du, Zhanqiu Guo, Xin Ye, Liu Ren, Chenyan Xiong

机构 * Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所) Bosch Research North America & Bosch Center for Artificial Intelligence (BCAI)(博世北美研究部及博世人工智能中心(BCAI))

专题命中 预训练与数据 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出EmbodiedMidtrain方法,通过中训练弥合视觉语言模型与视觉语言动作模型之间的差距,实验表明中训练能提升不同VLM骨干网络的性能,且在初始化阶段对VLA微调有显著帮助。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00414 2026-04-23 cs.AI cs.CL 81%

Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training

认知内核-Pro:深度研究代理及代理基础模型训练的框架

Tianqing Fang, Zhisong Zhang, Xiaoyang Wang, Rui Wang, Can Qin, Yuxuan Wan, Jun-Yu Ma, Ce Zhang, Jiaqi Chen, Xiyun Li, Yonglin Wang, Jingchen Ni, Tianshi Zheng, Chun Chen, Wenhao Yu, Zhenwen Liang, Hongming Zhang, Haitao Mi, Dong Yu

机构 * Tencent AI Lab(腾讯人工智能实验室)

专题命中 预训练与数据 :foundation model(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出Cognitive Kernel-Pro框架,旨在通过开放源代码和免费资源促进高级AI代理的开发与评估,重点研究高质量训练数据的构建及代理测试时的反思与投票策略,实现了在GAIA上的领先性能。

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12596 2026-04-15 cs.LG cs.AI 81%

KumoRFM-2: Scaling Foundation Models for Relational Learning

KumoRFM-2:关系学习的规模化基础模型

Valter Hudovernik, Federico López, Vid Kocijan, Akihiro Nitta, Jan Eric Lenssen, Jure Leskovec, Matthias Fey

专题命中 预训练与数据 :foundation model(title,abstract);分类 cs.AI、cs.LG

AI总结 KumoRFM-2是一种新型关系学习基础模型,支持上下文学习和微调,可处理多种预测任务。相比传统表格模型,其原生处理关系数据,无需手动展平表格或生成目标变量,且保持时间一致性。通过实验表明,KumoRFM-2在8%以内超越监督方法,且在冷启动和噪声数据下表现稳定。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08977 2026-04-15 cs.LG cs.AI stat.ML 81%

Simulation as Supervision: Mechanistic Pretraining for Scientific Discovery

模拟作为监督:用于科学发现的机理预训练

Carson Dudley, Reiden Magdaleno, Christopher Harding, Marisa Eisenberg

专题命中 预训练与数据 :pretraining(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出SGNN框架,通过机理模拟作为训练数据提升科学推断的鲁棒性,展示了其在多个学科中的预测优势和对模型不规范的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08649 2026-04-13 cs.LG cs.CE cs.CL cs.IR q-fin.CP 81%

PRAGMA: Revolut Foundation Model

PRAGMA:Revolut 基础模型

Maxim Ostroukhov, Ruslan Mikhailov, Vladimir Iashin, Artem Sokolov, Andrei Akshonov, Vitaly Protasov, Dmitrii Beloborodov, Vince Mullin, Roman Yokunda Enzmann, Georgios Kolovos, Jason Renders, Pavel Nesterov, Anton Repushko

专题命中 预训练与数据 :foundation model(title,abstract);分类 cs.CL、cs.LG

AI总结 PRAGMA是一种多源银行事件序列的基础模型,通过预训练Transformer架构实现金融记录的自监督学习,适用于信用评分、欺诈检测等任务,提供通用金融应用的表示层。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04999 2026-04-08 cs.LG cs.AI 81%

PRIME: Prototype-Driven Multimodal Pretraining for Cancer Prognosis with Missing Modalities

PRIME:面向癌症预后分析的原型驱动多模态预训练方法,适用于缺失模态

Kai Yu, Shuang Zhou, Yiran Song, Zaifu Zhan, Jie Peng, Kaixiong Zhou, Tianlong Chen, Feng Xie, Meng Wang, Huazhu Fu, Mingquan Lin, Rui Zhang

机构 * University of Minnesota(明尼苏达大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) North Carolina State University(北卡罗来纳州立大学) National University of Singapore(新加坡国立大学) Institute of High Performance Computing, Agency for Science, Technology and Research(高性能计算研究所,新加坡科技研究局)

专题命中 预训练与数据 :pretraining(title,abstract);分类 cs.AI、cs.LG

AI总结 PRIME通过原型记忆银行实现缺失模态下的多模态自监督预训练,提升在碎片化临床数据中的预测性能,实现结构对齐和鲁棒性提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03180 2026-04-06 cs.LG cs.CL cs.IR cs.SI 81%

PRISM: LLM-Guided Semantic Clustering for High-Precision Topics

PRISM:基于LLM的语义聚类用于高精度主题

Connor Douglas, Utkucan Balci, Joseph Aylett-Bullock

机构 * New York University(纽约大学) Binghamton University(宾汉姆顿大学)

专题命中 预训练与数据 :LLM(title,abstract);分类 cs.CL、cs.LG

AI总结 PRISM结合LLM的丰富表示与低成本可解释的潜在语义聚类方法,通过稀疏LLM标签微调句子编码模型,提升主题分离度,适用于多语料库的高精度主题发现。

Comments To appear in Proceedings of the ACM Web Conference 2026 (WWW 26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21658 2026-03-24 cs.CL cs.LG 81%

A Comparative Analysis of LLM Memorization at Statistical and Internal Levels: Cross-Model Commonalities and Model-Specific Signatures

大语言模型在统计和内部层面的记忆比较分析:跨模型共性与模型特有特征

Bowen Chen, Namgi Han, Yusuke Miyao

机构 * Department of Computer Science, The University of Tokyo(东京大学计算机科学系) Research and Development Center for Large Language Models, National Institute of Informatics(信息学研究院大语言模型研究与开发中心)

专题命中 预训练与数据 :LLM(title,abstract);分类 cs.CL、cs.LG

AI总结 本文通过分析多个模型系列的记忆行为,揭示了记忆率与模型规模的对数线性关系及序列压缩,同时发现不同模型在内部层面存在独特的解码过程和重要头部分布特征。

Comments 8 pages of main content, in conference submission, other contents are references and extra appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19028 2026-03-20 cs.CV cs.AI cs.LG 81%

SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models

SEM:用于视觉-语言模型后验去偏的稀疏嵌入调制

Quentin Guimard, Federico Bartsch, Simone Caldarella, Rahaf Aljundi, Elisa Ricci, Massimiliano Mancini

机构 * University of Trento(特伦托大学) Toyota Motor Europe(丰田欧洲公司) Fondazione Bruno Kessler(布鲁诺·凯瑟尔基金会)

专题命中 预训练与数据 :language model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出SEM框架,通过稀疏自编码器空间调制嵌入,实现视觉-语言模型的后验去偏,提升检索和零样本分类的公平性。

Comments CVPR Findings 2026. Project website: https://sparse-embedding-modulation.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11894 2026-03-17 cs.CV cs.AI cs.LG 81%

3D-LFM: Lifting Foundation Model

3D-LFM:提升基础模型

Mosam Dabhi, Laszlo A. Jeni, Simon Lucey

机构 * Carnegie Mellon University(卡内基梅隆大学) The University of Adelaide(阿德莱德大学)

专题命中 预训练与数据 :foundation model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出3D-LFM,利用Transformer的排列等变性处理不同数量的3D点,提升3D结构和相机的重建能力,实现跨领域的高泛化性能。

Comments Visit the project page at https://3dlfm.github.io for links to additional media, code, and videos. The site also features a custom GPT tailored to address queries related to 3D-LFM. Accepted at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11950 2026-03-13 cs.AI cs.LG 81%

Learning Transferable Sensor Models via Language-Informed Pretraining

通过语言引导预训练学习可迁移的传感器模型

Yuliang Chen, Arvind Pillai, Yu Yvonne Wu, Tess Z. Griffin, Lisa Marsch, Michael V. Heinz, Nicholas C. Jacobson, Andrew Campbell

专题命中 预训练与数据 :pretraining(title);language model(abstract);分类 cs.AI、cs.LG

AI总结 SLIP通过语言引导预训练学习可迁移的传感器模型,整合对比对齐与传感器条件标注,支持不同时间分辨率和输入长度,实现跨领域任务的高效性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02879 2026-03-10 cs.LG cs.AI 81%

CauKer: Classification Time Series Foundation Models Can Be Pretrained on Synthetic Data

CauKer:分类时间序列基础模型可以在合成数据上进行预训练

Shifeng Xie, Vasilii Feofanov, Ambroise Odonnat, Lei Zan, Marius Alonso, Jianfeng Zhang, Themis Palpanas, Lujia Pan, Keli Zhang, Ievgen Redko

机构 * Université Paris Cité(巴黎大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 预训练与数据 :foundation model(title,abstract);分类 cs.AI、cs.LG

AI总结 CauKer通过生成因果一致的合成时间序列数据,实现对分类时间序列基础模型的样本高效预训练。

Comments This manuscript combines material from the ICML 2025 TSFM Workshop paper and the ICLR 2026 Main Track paper

Journal ref ICLR 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06950 2026-03-10 q-bio.GN cs.AI cs.LG 81%

How Private Are DNA Embeddings? Inverting Foundation Model Representations of Genomic Sequences

DNA嵌入的隐私性如何?逆向基础模型对基因组序列的表示

Sofiane Ouaari, Jules Kreuer, Nico Pfeifer

机构 * Methods in Medical Informatics, Department of Computer Science, University of Tuebingen, Germany(图宾根大学医学信息学方法系,计算机科学系) Institute for Bioinformatics and Medical Informatics (IBMI), University of Tuebingen, Germany(图宾根大学生物信息学与医学信息学研究所)

专题命中 预训练与数据 :foundation model(title,abstract);分类 cs.AI、cs.LG

AI总结 研究评估了DNA基础模型在逆向攻击下的隐私安全性,发现嵌入相似性与序列相似性密切相关,DNABERT-2的分词方法更具隐私保护性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06565 2026-03-09 cs.AI cs.LG 81%

Boosting deep Reinforcement Learning using pretraining with Logical Options

通过逻辑选项预训练提升深度强化学习

Zihan Ye, Phil Chau, Raban Emunds, Jannis Blüml, Cedric Derstroff, Quentin Delfosse, Oleg Arenz, Kristian Kersting

专题命中 预训练与数据 :pretraining(title,abstract);分类 cs.AI、cs.LG

AI总结 通过逻辑选项预训练提升深度强化学习,实现目标导向行为与长周期决策能力的提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02406 2026-03-09 cs.LG cs.AI 81%

Rigidity-Aware Geometric Pretraining for Protein Design and Conformational Ensembles

面向刚性的几何预训练用于蛋白质设计与构象集合

Zhanghan Ni, Yanjing Li, Zeju Qiu, Bernhard Schölkopf, Hongyu Guo, Weiyang Liu, Shengchao Liu

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Washington(华盛顿大学) MPI for Intelligent Systems, Tübingen(智能系统马克斯·普朗克研究所,图宾根) National Research Council of Canada(加拿大国家研究理事会) University of Ottawa(渥太华大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 预训练与数据 :pretraining(title,abstract);分类 cs.AI、cs.LG

AI总结 RigidSSL通过刚性感知的几何预训练,提升蛋白质设计和构象集合的生成能力。

Comments The Fourteenth International Conference on Learning Representations; Code available at: https://github.com/ZhanghanNi/RigidSSL.git

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18300 2026-03-09 cs.CV cs.AI cs.LG cs.RO 81%

FALCON: Future-Aware Learning with Contextual Object-Centric Pretraining for UAV Action Recognition

FALCON:面向未来的学习与基于上下文的对象中心预训练用于无人机动作识别

Ruiqi Xian, Xiyang Wu, Tianrui Guan, Xijun Wang, Boqing Gong, Dinesh Manocha

专题命中 预训练与数据 :pretraining(title,abstract);分类 cs.AI、cs.LG

AI总结 FALCON通过结合对象感知的遮蔽自动编码和对象中心双时间线未来重建,提升无人机动作识别的准确率和推理速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08393 2026-03-06 cs.LG cs.AI 81%

Controlled LLM Training on Spectral Sphere

在谱球上受控的大规模模型训练

Tian Xie, Haoming Luo, Haoyu Tang, Yiwen Hu, Jason Klein Liu, Qingnan Ren, Yang Wang, Wayne Xin Zhao, Rui Yan, Bing Su, Chong Luo, Baining Guo

机构 * Microsoft Research Asia(微软亚洲研究院) Renmin University(中国人民大学) Wuhan University(武汉大学) IQuest Research(IQuest研究)

专题命中 预训练与数据 :LLM(title);pretraining(abstract);分类 cs.AI、cs.LG

AI总结 本文提出谱球优化器(SSO),通过严格约束权重和更新的谱特性,实现与最大更新参数化(μP)的完全对齐,从而在大规模模型训练中提升稳定性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03583 2026-03-05 cs.CL cs.LG 81%

ByteFlow: Language Modeling through Adaptive Byte Compression without a Tokenizer

ByteFlow:通过自适应字节压缩实现语言建模而不使用分词器

Chunyuan Deng, Sanket Lokegaonkar, Colin Lockard, Besnik Fetahu, Nasser Zalmout, Xian Li

机构 * Rice University(里士德大学) Amazon Science(亚马逊科学)

专题命中 预训练与数据 :language model(title,abstract);分类 cs.CL、cs.LG

AI总结 ByteFlow通过自适应字节压缩实现无分词器的语言建模,提升模型性能与适应性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03862 2026-03-03 cs.CL cs.AI 81%

Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design Decisions

并非仅仅的规模法则:迈向更深入理解语言模型设计决策对下游影响的探索

Emmy Liu, Amanda Bertsch, Lintang Sutawika, Lindia Tjuatja, Patrick Fernandes, Lara Marinov, Michael Chen, Shreya Singhal, Carolin Lawrence, Aditi Raghunathan, Kiril Gashteovski, Graham Neubig

机构 * Carnegie Mellon University(卡内基梅隆大学) Language Technologies Institute(语言技术研究所) Instituto Superior Técnico (Lisbon ELLIS Unit)(里斯本ELLIS单位(理工学院)) Instituto de Telecomunicações(电信研究所) NEC Laboratories Europe, Germany(德国NEC欧洲实验室) CAIR, Ss. Cyril and Methodius University of Skopje, North Macedonia(北马其顿斯·西里尔和方法ius大学)

专题命中 预训练与数据 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本研究通过分析92个开源预训练模型,发现结合模型大小和训练token数量以外的特征,可提升对下游性能预测能力,揭示了数据组成和架构决策对模型性能的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23784 2026-03-02 cs.LG cs.AI q-fin.CP q-fin.TR 81%

TradeFM: A Generative Foundation Model for Trade-flow and Market Microstructure

TradeFM: 一种用于交易流和市场微观结构的生成基础模型

Maxime Kawawa-Beaudan, Srijan Sood, Kassiani Papasotiriou, Daniel Borrajo, Manuela Veloso

专题命中 预训练与数据 :foundation model(title,abstract);分类 cs.AI、cs.LG

AI总结 TradeFM是一种基于生成Transformer的市场微观结构基础模型,通过学习大规模交易数据实现跨资产泛化,有效捕捉市场特征并提升合成数据生成与交易策略研究能力。

Comments 29 pages, 17 figures, 6 tables. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22037 2026-02-26 cs.CL cs.LG 81%

ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality

ATLAS:适应性迁移缩放定律用于多语言预训练、微调和解码的多语言诅咒

Shayne Longpre, Sneha Kudugunta, Niklas Muennighoff, I-Hung Hsu, Isaac Caswell, Alex Pentland, Sercan Arik, Chen-Yu Lee, Sayna Ebrahimi

机构 * MIT(麻省理工学院) University of Washington(华盛顿大学) Stanford University(斯坦福大学) Google Cloud AI(谷歌云人工智能) Google DeepMind(谷歌DeepMind)

专题命中 预训练与数据 :pretraining(title,abstract);分类 cs.CL、cs.LG

AI总结 ATLAS提出了一种适应性迁移缩放定律,用于多语言预训练、微调和解码,解决了多语言学习中的性能瓶颈和计算优化问题。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16784 2026-02-20 cs.LG cs.CL stat.ME 81%

Omitted Variable Bias in Language Models Under Distribution Shift

语言模型中分布偏移下的遗漏变量偏差

Victoria Lin, Louis-Philippe Morency, Eli Ben-Michael

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 预训练与数据 :language model(title,abstract);分类 cs.CL、cs.LG

AI总结 本文研究了语言模型在分布偏移下的遗漏变量偏差问题,提出框架用于评估和优化模型在分布偏移下的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16537 2026-02-19 cs.CV cs.AI cs.CL cs.RO 81%

RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

RoboSpatial: 教授机器人2D和3D视觉-语言模型空间理解能力

Chan Hee Song, Valts Blukis, Jonathan Tremblay, Stephen Tyree, Yu Su, Stan Birchfield

机构 * The Ohio State University(俄亥俄州立大学) NVIDIA(英伟达)

专题命中 预训练与数据 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 RoboSpatial通过构建大规模空间理解数据集,提升机器人2D和3D视觉-语言模型的空间推理能力。

Comments CVPR 2025 (Oral); Project Website: https://chanh.ee/RoboSpatial

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15014 2026-02-17 cs.LG cs.CL 81%

Scaling Beyond Masked Diffusion Language Models

超越掩码扩散语言模型的扩展

Subham Sekhar Sahoo, Jean-Marie Lemercier, Zhihan Yang, Justin Deschenaux, Jingyu Liu, John Thickstun, Ante Jukic

机构 * Department of Computer Science, Cornell Tech, NYC, USA(康奈尔科技学院计算机科学系) Department of Computer Science, Cornell University, Ithaca, USA(康奈尔大学计算机科学系) School of Computer and Communication Sciences, EPFL Lausanne, Switzerland(洛桑联邦理工学院计算机与通信科学学院) Department of Computer Science, University of Chicago, Illinois(芝加哥大学计算机科学系) NVIDIA, Santa Clara, USA(英伟达)

专题命中 预训练与数据 :language model(title,abstract);分类 cs.CL、cs.LG

AI总结 本文研究了统一状态和插值离散扩散方法的扩展定律,发现掩码扩散在FLOPs效率上可提升12%,并展示了统一状态扩散在GSM8K任务中的优越表现

Comments code: https://github.com/s-sahoo/scaling-dllms

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02387 2026-02-11 cs.LG cs.AI 81%

BiSSL: Enhancing the Alignment Between Self-Supervised Pretraining and Downstream Fine-Tuning via Bilevel Optimization

BiSSL: 通过双层优化增强自监督预训练与下游微调之间的对齐

Gustav Wagner Zakarias, Lars Kai Hansen, Zheng-Hua Tan

机构 * Aalborg University(奥胡斯大学) Technical University of Denmark(技术大学) Pioneer Centre for AI(先锋人工智能中心)

专题命中 预训练与数据 :pretraining(title,abstract);分类 cs.AI、cs.LG

AI总结 BiSSL通过双层优化增强自监督预训练与下游微调之间的对齐,提升模型在下游任务中的性能。

Journal ref Transactions on Machine Learning Research (TMLR), (02/2026)

详情

展开后加载摘要…

URL PDF HTML 收藏