arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 138627 信号源:cs.CL, cs.AI, cs.LG

1. 预训练与数据 12379 篇

2512.24097 2026-01-01 cs.CV cs.AI cs.CL cs.MM 84%

Factorized Learning for Temporally Grounded Video-Language Models

分解学习用于时间感知的视频-语言模型

Wenzheng Zeng, Difei Gao, Mike Zheng Shou, Hwee Tou Ng

机构 * National University of Singapore(新加坡国立大学)

专题命中 预训练与数据 :language model(title,abstract);preference optimization(abstract);分类 cs.CL、cs.AI

AI总结 本文提出D$^2$VLM框架,通过分解学习方法提升视频-语言模型在时间定位和文本响应任务中的性能,引入证据标记和FPO算法以优化学习过程。

Comments ICCV 2025 paper. This arXiv version updates Figure 1 to include the concurrent work Qwen2.5-VL to ensure consistency with Table 1

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18387 2025-12-30 cs.AI cs.LG 84%

Scaling Capability in Token Space: An Analysis of Large Vision Language Model

令牌空间中的扩展能力:对大视觉语言模型的分析

Tenghui Li, Guoxu Zhou, Xuyang Zhao, Qibin Zhao

机构 * School of Automation, Guangdong University of Technology(广东工业大学自动化学院) Key Laboratory of Intelligent Detection and the Internet of Things in Manufacturing, Ministry of Education(教育部智能制造智能检测与物联网重点实验室) Guangdong Provincial Key Laboratory of Intelligent Systems and Optimization Integration(广东省智能系统与优化集成重点实验室) Medical Science Data-driven Mathematics Team, RIKEN Center for Interdisciplinary Theoretical and Mathematical Sciences(RIKEN跨学科理论与数学科学中心医学科学数据驱动数学团队) Medical Data Mathematical Reasoning Special Team, RIKEN Center for Integrative Medical Sciences(RIKEN整合医学科学中心医学数据数学推理特别团队) Department of Artificial Intelligence Medicine, Chiba University(千叶大学人工智能医学系) Tensor Learning Team, RIKEN Center for Advanced Intelligence Project(RIKEN高级人工智能项目中心张量学习团队)

专题命中 预训练与数据 :language model(title,abstract);large language model(abstract);分类 cs.AI、cs.LG

AI总结 本研究通过理论分析和实证验证,揭示了视觉语言模型在视觉令牌数量上的扩展规律,发现不同数量的视觉令牌对应不同的扩展模式,并提出了扩展指数与视觉令牌表示相关结构的关系。

Journal ref Journal of Machine Learning Research, volume 26, number 253, page 1--61, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23066 2025-12-29 cs.CL cs.LG 84%

Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank

不要关注,PLANT它:通过学习排序进行预训练注意力

Debjyoti Saha Roy, Byron C. Wallace, Javed A. Aslam

机构 * Khoury College of Computer Sciences, Northeastern University(东北大学克劳利计算机科学学院)

专题命中 预训练与数据 :pretraining(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 PLANT通过学习排序模型预训练注意力,提升多标签文本分类性能,尤其在少样本和稀有标签任务中表现突出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18360 2025-12-23 cs.CL cs.AI 84%

LLM Agents Implement an NLG System from Scratch: Building Interpretable Rule-Based RDF-to-Text Generators

LLM Agents 实现从零构建 NLG 系统:构建可解释的基于规则的 RDF-to-Text 生成器

Mateusz Lango, Ondřej Dušek

机构 * Charles University, Faculty of Mathematics and Physics(查理大学数学与物理系)

专题命中 预训练与数据 :LLM(title,abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种基于 LLM agent 的神经符号框架,通过协作交互生成可解释的 RDF-to-Text 生成器,减少幻觉并提升生成效率。

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15248 2025-12-16 cs.CL cs.AI 84%

Synthetic bootstrapped pretraining

合成增强预训练

Zitong Yang, Aonan Zhang, Hong Liu, Tatsunori Hashimoto, Emmanuel Candès, Chong Wang, Ruoming Pang

机构 * Apple(苹果公司) Stanford University(斯坦福大学)

专题命中 预训练与数据 :pretraining(title,abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 SBP通过合成大量新语料库提升语言模型性能,实现60%的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01951 2025-12-12 cs.AI cs.CL cs.CY 84%

The LLM Wears Prada: Analysing Gender Bias and Stereotypes through Online Shopping Data

大语言模型穿Prada:通过在线购物数据分析性别偏见和刻板印象

Massimiliano Luca, Ciro Beneduce, Bruno Lepri, Jacopo Staiano

机构 * University of Trento(特伦托大学)

专题命中 预训练与数据 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过在线购物数据分析大语言模型中的性别偏见,发现模型在性别预测中受刻板印象影响,且减少偏见指令只能降低预测确定性,无法消除刻板印象模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15272 2025-12-03 cs.CV cs.AI cs.CL 84%

CT-GLIP: 3D Grounded Language-Image Pretraining with CT Scans and Radiology Reports for Full-Body Scenarios

CT-GLIP:基于CT扫描和放射报告的3D grounded语言-图像预训练,用于全身场景

Jingyang Lin, Yingda Xia, Jianpeng Zhang, Ke Yan, Kai Cao, Le Lu, Jiebo Luo, Ling Zhang

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) University of Rochester(罗切斯特大学) Hupan Lab, \postcode 10587, \state Hangzhou, \country China(华普实验室,浙江省杭州市,中国) Department of Radiology, Shanghai Institution of Pancreatic Disease, \state Shanghai, \country China(胰腺疾病研究所放射科,上海市,中国)

专题命中 预训练与数据 :pretraining(title,abstract);foundation model(abstract);分类 cs.CL、cs.AI

AI总结 CT-GLIP通过构建细粒度CT报告对,提升3D grounded语言-图像预训练,实现更精确的跨模态对齐,从而在零样本任务中提升器官识别和肿瘤检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15390 2025-12-01 cs.CL cs.LG 84%

Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training

利用词汇频率不平衡性在语言模型预训练中

Woojin Chung, Jeonghoon Kim

机构 * KAIST(韩国科学技术院)

专题命中 预训练与数据 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.LG

AI总结 本研究通过扩大词汇表降低分词文本复杂性,揭示了语言模型预训练中词汇规模与性能的关系。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21861 2025-12-01 cs.LG cs.AI 84%

Towards a Foundation Model for Partial Differential Equations Across Physics Domains

面向跨物理领域的偏微分方程基础模型

Eduardo Soares, Emilio Vital Brazil, Victor Shirasuna, Breno W. S. R. de Carvalho, Cristiano Malossi

专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract);分类 cs.AI、cs.LG

AI总结 PDE-FM是一种跨物理领域预训练的基础模型,通过统一空间、谱和时间推理,实现复杂物理动态的高效建模与跨领域泛化。

Comments Accepted to the AAAI 2026 AI2ASE Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18411 2025-11-25 cs.CL cs.AI 84%

SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data

SmolKalam:大规模质量过滤翻译用于高质量阿拉伯语后训练数据

Sultan Alrashed, Chadi Helwe, Francesco Orabona

机构 * King Abdullah University of Science and Technology (KAUST)(卡斯特大学)

专题命中 预训练与数据 :post-training(title,abstract);pretraining(abstract);分类 cs.CL、cs.AI

AI总结 SmolKalam通过多模型集成和质量过滤技术,提升大规模阿拉伯语后训练数据的翻译质量。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16577 2025-11-21 cs.CL cs.AI 84%

Integrating Symbolic Natural Language Understanding and Language Models for Word Sense Disambiguation

整合符号自然语言理解与语言模型用于词义消歧

Kexin Zhao, Ken Forbus

专题命中 预训练与数据 :language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI

AI总结 本文提出利用统计语言模型和符号NLU系统整合的方法,实现无需人工标注数据的词义消歧。

Comments 16 pages

Journal ref Proceedings of the Twelfth Annual Conference on Advances in Cognitive Systems ACS-2025 (333-348)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12768 2025-11-18 cs.CL cs.AI 84%

Evidence of Phase Transitions in Small Transformer-Based Language Models

Noah Hong, Tao Hong

机构 * Lynbrook High School(林brook高中) Keysight Technologies(Keysight技术公司)

专题命中 预训练与数据 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11572 2025-11-18 cs.GL cs.CL cs.LG 84%

LLM Architecture, Scaling Laws, and Economics: A Quick Summary

William H. Press

专题命中 预训练与数据 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03251 2025-11-06 cs.LG cs.AI cs.SI 84%

GMoPE:A Prompt-Expert Mixture Framework for Graph Foundation Models

Zhibin Wang, Zhixing Zhang, Shuqi Wang, Xuanting Xie, Zhao Kang

专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18065 2025-11-05 cs.CV cs.AI cs.CL cs.RO 84%

Unseen from Seen: Rewriting Observation-Instruction Using Foundation Models for Augmenting Vision-Language Navigation

Ziming Wei, Bingqian Lin, Yunshuang Nie, Jiaqi Chen, Shikui Ma, Hang Xu, Xiaodan Liang

机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Shanghai Jiao Tong University(上海交通大学) The University of Hong Kong(香港大学) Hunan Artificial Intelligence and Robotics Institute Company Ltd.(湖南人工智能与机器人研究院有限公司) Huawei Noah’s Ark Lab(华为诺亚实验室) Peng Cheng Laboratory(鹏城实验室)

专题命中 预训练与数据 :foundation model(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments Accepted by IEEE Transactions on Neural Networks and Learning Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04340 2025-11-04 cs.CL cs.AI 84%

Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time

Daniel Tan, Anders Woodruff, Niels Warncke, Arun Jose, Maxime Riché, David Demitri Africa, Mia Taylor

机构 * University College London(伦敦大学学院) Center on Long-Term Risk(长期风险中心) McGill University(麦吉尔大学) UK AI Security Institute(英国人工智能安全研究所)

专题命中 预训练与数据 :prompting(title,abstract);language model(abstract);分类 cs.CL、cs.AI

Comments 40 pages, 22 figures. Under review at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00346 2025-11-04 cs.CR cs.AI cs.LG 84%

Exploiting Latent Space Discontinuities for Building Universal LLM Jailbreaks and Data Extraction Attacks

Kayua Oleques Paim, Rodrigo Brandao Mansilha, Diego Kreutz, Muriel Figueredo Franco, Weverton Cordeiro

专题命中 预训练与数据 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 5 figures, 4 tables, Published at the Brazilian Symposium on Cybersecurity (SBSeg 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09846 2025-11-04 cs.LG cs.AI math.OC stat.ML 84%

Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training

Minhak Song, Beomhan Baek, Kwangjun Ahn, Chulhee Yun

机构 * KAIST(韩国科学技术院) SNU & KAIST InnoCORE LLM(首尔国立大学及韩国科学技术院InnoCORE LLM) Microsoft Research(微软研究院)

专题命中 预训练与数据 :language model(title,abstract);pretraining(abstract);分类 cs.AI、cs.LG

Comments Published at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25771 2025-10-30 cs.CL cs.AI 84%

Gaperon: A Peppered English-French Generative Language Model Suite

Nathan Godey, Wissam Antoun, Rian Touchent, Rachel Bawden, Éric de la Clergerie, Benoît Sagot, Djamé Seddah

专题命中 预训练与数据 :language model(title,abstract);pretraining(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24139 2025-10-29 cs.CL cs.AI 84%

Beyond Line-Level Filtering for the Pretraining Corpora of LLMs

Chanwoo Park, Suyoung Park, Yelim Ahn, Jongmin Kim, Jongyeon Park, Jaejin Lee

专题命中 预训练与数据 :pretraining(title);language model(abstract);small language model(abstract);分类 cs.CL、cs.AI

Comments submitted to ACL ARR Rolling Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20860 2025-10-27 eess.AS cs.CL cs.LG 84%

Data-Centric Lessons To Improve Speech-Language Pretraining

Vishaal Udandarao, Zhiyun Lu, Xuankai Chang, Yongqiang Wang, Violet Z. Yao, Albin Madapally Jose, Fartash Faghri, Josh Gardner, Chung-Cheng Chiu

机构 * Apple(苹果公司) University of Cambridge(剑桥大学) University of Tübingen(图宾根大学)

专题命中 预训练与数据 :pretraining(title,abstract);language model(abstract);分类 cs.CL、cs.LG

Comments Tech Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18910 2025-10-23 cs.LG cs.AI 84%

Large Connectome Model: An fMRI Foundation Model of Brain Connectomes Empowered by Brain-Environment Interaction in Multitask Learning Landscape

Ziquan Wei, Tingting Dan, Guorong Wu

专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract);分类 cs.AI、cs.LG

Comments 12 pages 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08445 2025-10-21 cs.LG cs.AI 84%

Synthetic Series-Symbol Data Generation for Time Series Foundation Models

Wenxuan Wang, Kai Wu, Yujian Betterest Li, Dan Wang, Xiaoyu Zhang

机构 * School of Telecommunications Engineering Xidian University(电子科技大学电信工程学院) School of Artificial Intelligence Xidian University(电子科技大学人工智能学院) School of Cyber Engineering Xidian University(电子科技大学网络安全工程学院)

专题命中 预训练与数据 :foundation model(title,abstract);pretraining(abstract);分类 cs.AI、cs.LG

Comments 64 pages, 25 figures, 35 tables, NeurIPS 2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09426 2025-10-14 cs.CV cs.AI cs.CL 84%

BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning

Shengao Wang, Arjun Chandra, Aoming Liu, Venkatesh Saligrama, Boqing Gong

机构 * Boston University(波士顿大学)

专题命中 预训练与数据 :pretraining(title,abstract);language model(abstract);分类 cs.CL、cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10009 2025-10-14 cs.CL cs.AI cs.IR 84%

Beyond the limitation of a single query: Train your LLM for query expansion with Reinforcement Learning

Shu Zhao, Tan Yu, Anbang Xu

机构 * NVIDIA(英伟达) Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 预训练与数据 :LLM(title,abstract);post-training(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09846 2025-10-14 cs.LG cs.AI 84%

CALM: A Causal Analysis Language Model for Tabular Data in Complex Systems with Local Scores, Conditional Independence Tests, and Relation Attributes

Zhenjiang Fan, Zengyi Qin, Yuanning Zheng, Bo Xiong, Summer Han

专题命中 预训练与数据 :language model(title,abstract);large language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04268 2025-10-07 cs.CL cs.AI 84%

LongTail-Swap: benchmarking language models' abilities on rare words

Robin Algayres, Charles-Éric Saint-James, Mahi Luthra, Jiayi Shen, Dongyan Lin, Youssef Benchekroun, Rashel Moritz, Juan Pino, Emmanuel Dupoux

机构 * Meta AI

专题命中 预训练与数据 :language model(title,abstract);pretraining(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19551 2025-10-07 cs.CL cs.AI 84%

Scaling Laws of Synthetic Data for Language Models

Zeyu Qin, Qingxiu Dong, Xingxing Zhang, Li Dong, Xiaolong Huang, Ziyi Yang, Mahmoud Khademi, Dongdong Zhang, Hany Hassan Awadalla, Yi R. Fung, Weizhu Chen, Minhao Cheng, Furu Wei

机构 * Microsoft(微软公司) HKUST(香港科技大学) Peking University(北京大学) Penn State University(宾夕法尼亚州立大学)

专题命中 预训练与数据 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01571 2025-10-03 cs.LG cs.AI q-bio.BM 84%

From Supervision to Exploration: What Does Protein Language Model Learn During Reinforcement Learning?

Hanqun Cao, Hongrui Zhang, Junde Xu, Zhou Zhang, Lingdong Shen, Minghao Sun, Ge Liu, Jinbo Xu, Wu-Jun Li, Jinren Ni, Cesar de la Fuente-Nunez, Tianfan Fu, Yejin Choi, Pheng-Ann Heng, Fang Wu

机构 * The Chinese University of Hong Kong(香港中文大学) Peking University(北京大学) Stanford University(斯坦福大学) Nanjing University(南京大学) University of Pennsylvania(宾夕法尼亚大学) National University of Singapore(新加坡国立大学) University of Illinois Urbana-Champaign(伊利诺伊大学香槟分校) Toyota Technological Institute at Chicago(芝加哥丰田技术研究所)

专题命中 预训练与数据 :language model(title,abstract);pretraining(abstract);分类 cs.AI、cs.LG

Comments 24 pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25498 2025-10-01 cs.CL cs.AI 84%

Not Wrong, But Untrue: LLM Overconfidence in Document-Based Queries

Nick Hagar, Wilma Agustianto, Nicholas Diakopoulos

机构 * Northwestern University(西北大学) University of Minnesota(明尼苏达大学)

专题命中 预训练与数据 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments Accepted to Computation + Journalism Symposium 2025

详情

展开后加载摘要…

URL PDF HTML 收藏