arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 5804 信号源:cs.CL, cs.AI, cs.LG

1. 其他推理 5804 篇

2608.03036 2026-08-05 cs.SE cs.AI cs.LG 新提交 62%

LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs

野外环境下的大语言模型服务:框架、方法与系统设计的实证研究

Forough Majidi, Mohammad Mehdi Morovati, Foutse Khomh, Heng Li

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本研究通过实证分析开源软件系统中5种LLM服务框架的应用情况,明确了框架使用特点、常用服务方法及应用场景,为相关人员提供了实践见解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26760 2026-08-05 cs.CL cs.LG 版本更新 62%

Metis: Memory Foundation Model

Metis:记忆基础模型

Zeyu Zhang, Ziliang Guo, Yihang Sun, Xichong Zhang, Xixuan Hao, Zehao Lin, Yang Zhang, Xiaoyan Zhao, Tong Shen, Bo Tang, Zhi-Qin John Xu, Junchi Yan, Haofen Wang, Xu Chen, Feiyu Xiong, Zhiyu Li, Tat-Seng Chua

机构 * MemTensor (Shanghai) Technology Co., Ltd.(墨芯(上海)科技有限公司) Renmin University of China(中国人民大学) National University of Singapore(新加坡国立大学) Shanghai Jiao Tong University(上海交通大学) Tongji University(同济大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 该研究提出首个记忆基础模型原型Metis,赋予基础模型原生记忆能力,通过新架构与优化目标实现,经实验验证其具备原生记忆能力并发布相关资源。

Comments 46 pages, 11 figures, 16 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01585 2026-08-04 cs.CL cs.LG 新提交 62%

Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer

AI模型的语义对齐:概念坍缩、检查点动态与跨语言迁移

Tyler Ashoff, Jordan Rodu

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 该研究针对语言模型基准测试的难点,提出用拓扑方法将模型高维嵌入空间与可解释基线严格比较,以实现多模态对齐,追踪模型适应并测试跨语言短语理解。

Comments Code available at github.com/tylerashoff/persiscope (PyPI: persiscope)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27553 2026-08-04 cs.AI cs.CL econ.GN q-fin.EC 版本更新 62%

AI and Its Impact on Creativity and Diversity: An Empirical Study of LLM-Generated Product Ideas

使用大型语言模型进行创新中的创意生成

Christian Terwiesch, Lennart Meincke, Karan Girotra, Ethan Mollick, Gideon Nave, Karl T. Ulrich

专题命中 其他推理 :chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 该研究对比人类与GPT-4生成的大学生适用低价新产品创意,发现AI生成创意平均购买意向更高、跻身前10%的概率是人类的7倍,但新颖性较低,少样本提示效果略优于零样本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06160 2026-08-04 cs.CL cs.AI 版本更新 62%

LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis

LongCrafter:通过证据图引导的指令合成实现多样化的长上下文理解

Chenhao Yuan, Yinhao Xu, Shuwen Xu, Xizhi Yang, Jiaxiang Liu, Chenxi Zhou, Shaoping Huang, Haolin Ren, Pengfei Cao, Jun Zhao, Kang Liu

机构 * University of Chinese Academy of Sciences(中国科学院大学) The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所认知与决策智能复杂系统重点实验室)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 针对现有长上下文理解方法的局限,提出LongCrafter框架,结合分层任务分类法与证据管道,生成多样化长上下文SFT数据,微调后的模型在多个数据集上表现优异,能有效缓解“中间迷失”问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30363 2026-08-04 q-fin.CP cs.AI cs.LG q-fin.ST 版本更新 62%

Enhancing Regime Shift Detection Using Unstructured Data: A Study on the Treasury Market

利用非结构化数据增强制度转换检测:国债市场研究

Mingxuan Yi, Vidal Mehra, Jing Chen, John Cartlidge

机构 * School of Engineering Mathematics and Technology, University of Bristol, UK(布里斯托大学工程数学与技术学院) Propellant Digital B.V., Amsterdam, Netherlands(荷兰阿姆斯特丹Propellant Digital公司) School of Mathematics, Cardiff University, UK(卡迪夫大学数学学院)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 提出一种结合大语言模型推理与统计检验的文本增强型制度转换检测框架,在国债市场数据上实现F1=0.82,优于纯数据驱动方法。

Comments 9 pages, 4 figures. Selected for Long Oral presentation at the International Symposium on Large Language Models for Financial Services (FinLLM@IJCAI 2026), Bremen, Germany, 15 August 2026 (non-archival). Code available at: https://github.com/mingxuan-yi/regime_shift

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15607 2026-08-04 cs.CL cs.LG 版本更新 62%

Syntax Without Semantics: Teaching Large Language Models to Code in an Unseen Language

无语义的语法:教大语言模型在未见过的语言中编程

Vinayshekhar Bannihatti Kumar, Disha Makhija, Manoj Ghuhan Arivazhagan, Rashmi Gangadharaiah

机构 * AWS AI Labs(AWS人工智能实验室)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 研究探讨大语言模型在未见过的语言中生成代码的能力,发现微调仅能教授语法而无法转移语义能力,揭示了推理与语言实现之间的鸿沟。

Comments Accepted at COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28982 2026-08-03 cs.CL cs.AI 新提交 62%

PARALLEL: A Prefrontal-Aligned Reinforcement inspired Approach for Language-Model Learning under Explicit Limits

PARALLEL:显式约束下语言模型学习的前额叶对齐强化启发式方法

Namkyung Yoon, Sanghong Kim, Hwangnam Kim

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 PARALLEL是一种前额叶对齐的强化启发式语言模型学习方法,通过分配样本相关更新强度提升适配效率,在多任务上保留高性能且适配轨迹更稳定,支持高效稳定的部署后流式适配。

Comments 8 pages, 3 figures, and 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19321 2026-07-30 cs.AI cs.CR cs.LG 版本更新 62%

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

研究竞技场:评估自动化人工智能研发中的破坏行为与监控

Lena Libon, Ben Rank, Jehyeok Yeon, David Schmotz, Jeremy Qin, Daniel Donnelly, Derck Prinzhorn, Maksym Andriushchenko

专题命中 其他推理 :chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 研究针对自动化人工智能研发,用研究竞技场框架评估人工智能控制,通过四项长期任务及两类隐藏附带任务,评估前沿代理破坏与监控能力,发现训练数据中破坏行为难捕捉,发布框架用于评估自动化人工智能研发中的破坏与控制。

Comments 50 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05174 2026-07-30 cs.CL cs.AI 62%

Improving Heart-Focused Medical Question Answering in LLMs via Variance-Aware Rubric Rewards with GRPO

通过基于方差感知的评分规则奖励与GRPO改进LLMs中心脏医学问答

Arash Ahmadi, Parisa Masnadi Khiabani, Sarah Sharif, Charles Nicholson, David Ebert, Mike Banad

机构 * School of Electrical and Computer Engineering, University of Oklahoma, Norman, OK, USA(电气与计算机工程学院,俄克拉荷马大学,诺曼,OK,USA) Intelligent Neuromorphic and Quantum Understanding for Innovative Research and Engineering (INQUIRE) Laboratory, University of Oklahoma, Norman, OK, USA(创新研究与工程智能神经形态与量子理解实验室,俄克拉荷马大学,诺曼,OK,USA) Khiabani Data Science and Analytics Institute, University of Oklahoma, Norman, OK, USA(Khiabani数据科学与分析研究所,俄克拉荷马大学,诺曼,OK,USA) Data Institute for Societal Challenges (DISC), University of Oklahoma, Norman, OK, USA(社会挑战数据研究所(DISC),俄克拉荷马大学,诺曼,OK,USA) School of Industrial and Systems Engineering, University of Oklahoma, Norman, OK, USA(工业与系统工程学院,俄克拉荷马大学,诺曼,OK,USA) Office of Responsible Artificial Intelligence (ORAI), University of Arizona, Tucson, AZ, USA(负责任人工智能办公室(ORAI),亚利桑那大学,图森,AZ,USA)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 提出一种方差感知奖励框架,结合GRPO和RaR-Medicine的评分规则,通过连续分析奖励函数替代离散聚合,提升LLMs在心脏医学问答上的准确率和F1分数。

Comments 27 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12250 2026-07-30 cs.AI cs.CL cs.GT cs.MA 版本更新 62%

How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm

记忆如何影响基于LLM的社会粒子群中的集体和合作行为

Taisei Hishiki, Takaya Arita, Reiji Suzuki

机构 * Graduate School of Informatics(信息学研究生院)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨了大型语言模型代理的内部对齐特性如何影响记忆对其集体和合作动态的影响。通过将社会粒子群模型中的规则代理替换为具有大五人格特质和不同记忆长度的LLM代理,发现记忆长度是决定集体行为的关键参数,且不同模型对记忆的解读影响合作结果。

Comments 11 pages, 4 figures and 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25135 2026-07-29 cs.AI cs.LG 新提交 62%

ScalableRAG: High-Quality RAG at Zero Ingestion Cost

ScalableRAG:零摄入成本下的高质量检索增强生成

Hilaf Hasson, Aditya Chakravarty, Jayant Thomas, Krishna Gogineni

机构 * Cohesity(科赫思)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 研究提出零摄入成本的ScalableRAG及有限摄入的改进版本,通过维护工作区实现即时聚合推理,在六个语料库测试中表现出色,平均准确率大幅超越基线,还通过固定LLM调用次数等进一步提升大规模时的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22586 2026-07-28 cs.AI cs.CL 新提交 62%

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models

MM-ShiftKV:用于多模态大语言模型的解码感知预填充阶段键值选择

Jinsong Shu, Chenyang Wu, Zhongle Xie, Baokun Wang, Lidan Shou

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团) The State Key Laboratory of Blockchain and Data Security, Zhejiang University(浙江大学区块链与数据安全国家重点实验室) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术产业开发区(滨江)区块链与数据安全研究院)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 研究多模态大语言模型中KV缓存问题,提出MM-ShiftKV方法,通过构建方差扩展查询代理近似解码时查询行为,基于聚合注意力质量估计KV重要性,在严格缓存预算下性能优于现有方法。

Comments 19 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09239 2026-07-28 cs.CL cs.LG 版本更新 62%

Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs

重复标记计数揭示了表示与输出之间的脱节

Sohan Venkatesh

机构 * Manipal Institute of Technology Bengaluru(班加罗尔Manipal理工学院)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 研究发现大语言模型在重复标记计数任务中表现不佳,但通过线性探针发现错误源于多层感知机块的固定错误覆盖,而非表示层的缺陷。

Comments Code is available at https://github.com/sohv/counting-failure

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03423 2026-07-28 cs.CL cs.AI 62%

Training-Free Adaptation of New-Generation LLMs using Legacy Clinical Models

无需训练适应新一代语言模型使用遗留临床模型

Sasha Ronaghi, Chloe Stanwyck, Asad Aali, Amir Ronaghi, Miguel Fuentes, Tina Hernandez-Boussard, Emily Alsentzer

机构 * Stanford University(斯坦福大学) MemorialCare(纪念医疗中心)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出CAPT方法,通过模型融合实现无需训练的临床模型与通用领域模型适应,提升临床任务性能,适用于计算资源有限的医疗机构。

Journal ref Proceedings of the 7th Conference on Health, Inference, and Learning, PMLR 333:354-388, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01335 2026-07-28 cs.CL cs.AI 版本更新 62%

LEDOM: Reverse Language Model

LEDOM:反向语言模型

Xunjian Yin, Sitao Cheng, Yuxi Xie, Xinyu Hu, Li Lin, Xinyi Wang, Liangming Pan, William Yang Wang, Xiaojun Wan

机构 * Peking University(北京大学) University of California, Santa Barbara(加州大学圣芭芭拉分校) University of Arizona(亚利桑那大学) National University of Singapore(新加坡国立大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 LEDOM是一种反向自回归语言模型,通过反向训练获得独特能力,如演绎推理和解决反转诅咒,并在多个基准测试中提升了性能。

Comments Work in progress; Models can be found at: https://huggingface.co/Corning/Reverse-Model-7B-348B/tree/main

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02770 2026-07-27 cs.CL cs.AI 版本更新 62%

Gemma 4 Technical Report

Gemma 4技术报告

Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, Victor Cărbune, Michelle Casbon, Mayank Chaturvedi, Aditya Chawla, Victor Cotruta, Alice Coucke, Phil Culliton, Robert Dadashi, Lucas Dixon, Mohamed Elhawaty, Utku Evci, Clément Farabet, Johan Ferret, Filippo Galgani, Sertan Girgin, Jean-Bastien Grill, Maarten Grootendorst, Jiaxian Guo, Cassidy Hardin, Yanzhang He, Steven M. Hernandez, Omri Homburger, Léonard Hussenot, Juyeong Ji, Armand Joulin, Aishwarya Kamath, Parnian Kassraie, Olivier Lacombe, Preethi Lahoti, Gaël Liu, Gus Martins, Luciano Martins, Tatiana Matejovicova, Ramona Merhej, Nikola Momchev, Sneha Mondal, Ryan Mullins, Sindhu Raghuram Panyam, Shreya Pathak, Sarah Perrin, André Susano Pinto, Etienne Pot, Angéline Pouget, Alexandre Ramé, Sabela Ramos, Douglas Reid, David Rim, Morgane Rivière, Karsten Roth, Louis Rouillard, Omar Sanseviero, Pier Giuseppe Sessa, Shane Settle, Danila Sinopalnikov, Sara Smoot, Piotr Stanczyk, Andreas Steiner, Lawrence Stewart, Ilya Tolstikhin, Michael Tschannen, Anton Tsitsulin, Nino Vieillard, Renjie Wu, Pingmei Xu, Haichuan Yang, Edouard Yvinec, Biao Zhang, Li Zhang, Joe Zou, Nicolas Aagnes, Abdelrahman Abdelhamed, Jakub Adamek, Shivani Agrawal, Shubham Agrawal, Ibrahim Alabdulmohsin, Jean Baptiste Alayrac, Uri Alon, Chandramouli Amarnath, Ankesh Anand, Chrysovalantis Anastasiou, Setareh Ariafar, François-Xavier Aubet, Kyriakos Axiotis, Federico Barbero, Joelle Barral, Alexei Bendebury, Urs Bergmann, Stanley Bileschi, Kat Black, Mathieu Blondel, Sebastian Borgeaud, Arthur Bražinskas, Ryan Burnell, Robert Busa-Fekete, Mu Cai, Daniele Calandriello, Glenn Cameron, Charlotte Caucheteux, Rahma Chaabouni, Garima Chadha, Jetha Chan, Blake Jianhang Chen, Jesse Chen, Lin Chen, Xu Chen, Derek Cheng, Tzu-hsiang Chien, Nikolai Chinaev, Yi Chou, Zhaohui Chu, Benjamin Coleman, Pooja Consul, Sam Conway-Rahman, Scott Crowell, Dylan Cutler, Vivek Dani, Samira Daruki, Anil Das, Daniel Deutsch, Nishanth Dikkala, Li Ding, Qiuhan Ding, Shenil Dodhia, Konstantin Donhauser, Tulsee Doshi, Anca Dragan, Alex Druinsky, Sahil Dua, Zoltan Egyed, Danielle Eisenbud, Daniel Eppens, Cindy Fan, Bahare Fatemi, Yassir Fathullah, Vlad Feinberg, Milen Ferev, Sebastian Flennerhag, Takumi Fujimoto, João Gabriel Oliveira, Isaac Galatzer-Levy, João Gante, Simon Geisler, Soham Ghosal, Antonious M. Girgis, Tamara von Glehn, Alec Go, Alhaad Gokhale, Alex Grills, Yiming Gu, Mayank Gupta, Pramod Gupta, Guru Guruganesh, Raia Hadsell, Hamza Harkous, Jitendra Harlalka, Demis Hassabis, Anja Hauth, Joe Heyward, Arian Hosseini, Chih-Yang Hsia, I-Hung Hsu, Xiaopeng Huang, Yangsibo Huang, Kevin Hui, Adrian Hutter, Te I, Fotis Iliopoulos, Advait Jain, Ganesh Jawahar, Ziwei Ji, Qilin Jin, Melvin Johnson, Kandarp Joshi, Arun Kandoor, Wang-Cheng Kang, Koray Kavukcuoglu, Mehran Kazemi, Kathleen Kenealy, Amr Khalifa, Phoebe Kirk, Ivan Korotkov, Suraj Kothawade, Vitaly Kovalev, Neel Kovelamudi, Adam Kraft, Ravin Kumar, Vivek Kumar, Harish Kuppam, Justin Lannin, Chen-Yu Lee, Seungji Lee, Dmitry Lepikhin, Alon Levkovitch, Dongdong Li, Qiujia Li, Valentin Liévin, Ethan Lin, Ziqian Lin, Casper Liu, Tianlin Liu, Tianqi Liu, Xin Liu, Ivan Lobov, Mayank Lunayach, Min Ma, Gagan Madan, Andrii Maksai, Eric Malmi, Michal Matuszak, Daniel McDuff, Gaurav Menghani, Maciej Mikuła, Daniil Mirylenka, Karolis Misiunas, Vedant Misra, Andreea Mitran, Kareem Mohamed, Maksim Mukha, Eric Noland, James O'Donnell, Brendan O'Donoghue, Kate Olszewska, Bernett Orlando, Wanqiong Pan, Rina Panigrahy, Unnati Parekh, Nicolas Perez-Nieves, Chunjong Park, Eric Paskie, Liqian Peng, Bryce Petrini, Slav Petrov, Jonas Pfeiffer, Bilal Piot, Martyna Plomecka, Siim Poder, Octavio Ponce, Arijit Pramanik, David Racz, Anish Rajan, Michelle Ramanovich, Anand Rao, Marvin Ritter, Vitor Rodrigues, Evan Rosen, Mikołaj Rybiński, Noveen Sachdeva, Michaël E. Sander, Rohit Sathyanarayana, Sagar Savla, Samuel Schmidgall, Tal Schuster, George Scrivener, Benoit Seguin, Andrew Sellergren, Aliaksei Severyn, Izhak Shafran, Dhruv Shah, Bobak Shahriari, Yuan Shangguan, Ashish Shenoy, Pradeep Shenoy, Rakesh Shivanna, Pauline Sho, Lucas Spangher, Wojciech Stokowiec, Tim Strother, Yao Su, Yinghao Sun, Mukund Sundararajan, Andrea Tacchetti, Mor Hazan Taege, Pouya Tafti, Jean Tarbouriech, Chetan Tekur, Shantanu Thakoor, Rahul Thapa, Madeleine Traverse, Lenart Treven, Tao Tu, Chien Te Tung, Çağlar Ünlü, Petar Veličković, Malini Pooni Venkat, Sagar Gubbi Venkatesh, Vidya Venkiteswaran, Francesco Visin, Alex Vitvitskyi, Kiran Vodrahalli, Weiyi Wang, Xin Wang, Tris Warkentin, Jan Wassenberg, John Wieting, Cindy Wu, Lechao Xiao, Hao Xu, Yuhui Xu, Fuzhao Xue, Arun Yadav, Jun Yan, Antoine Yang, Lin Yang, Ming-Hsuan Yang, Ziyu Ying, Jae Hyeon Yoo, Morteza Zadimoghaddam, Sajjad Zafar, Fred Zhang, Jiageng Zhang, Jianyi Zhang, Xiaofan Zhang, Chao Zhao, David Zhou, Chen Zou

机构 * Google DeepMind(谷歌DeepMind)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 介绍新一代Gemma 4开源多模态语言模型,通过密集和专家混合架构提升计算效率与推理能力,提出统一无编码器架构,集成思考模式,改进多方面性能,在多基准测试中有显著提升。

Comments 17 pages, 2 figures, technical report, updated

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25774 2026-07-27 cs.CL cs.AI 62%

CGU-ILALab at FoodBench-QA 2026: Comparing Traditional and LLM-based Approaches for Recipe Nutrient Estimation

CGU-ILALab 在 FoodBench-QA 2026 中:比较传统方法与基于 LLM 的方法在食谱营养估算中的表现

Wei-Chun Chen, Yu-Xuan Chen, I-Fang Chung, Ying-Jia Lin

机构 * CGU-ILALab

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文比较了传统方法与基于 LLM 的方法在食谱营养估算中的性能,发现 LLM 能有效处理模糊术语和非标准单位,但计算效率较低。

Comments Accepted by the Third Workshop on Patient-oriented Language Processing (CL4Health) at LREC 2026

Journal ref https://lrec.elra.info/lrec2026-ws-cl4health-31

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21291 2026-07-24 cs.CL cs.LG 新提交 62%

Adaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMs

自适应深度稀疏框架:基于相似性驱动的预训练语言模型资源分配

Yidu Wu, Xiang Wang, Kejie Zhao, Zhangchi Wang, Qinghai Guo, Xiaoying Tang

机构 * Southern University of Science and Technology(南方科技大学) Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 研究针对预训练语言模型推理成本高的问题,提出自适应深度稀疏框架AdaDSF,基于层输入输出隐藏状态余弦相似性分配令牌保留率,用轻量级路由器选择信息令牌,在多任务中大幅降推理FLOPs且精度退化小。

Comments Accepted by ICIC 2026. 12 pages, 2 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20557 2026-07-24 cs.LG cs.AI 新提交 62%

Monkey King Bang: A Unified Scientific Multimodal Foundation Model

孙悟空爆炸:一个统一的科学多模态基础模型

Hesen Chen, Xinyu Su, Xiaomeng Yang, Yuetan Lin, Zixiong Yang, Junyi An, Fenglei Cao, Yifeng Jiao, Yunqi Zhang, Yuan Cheng, Zhiyu Tan, Hao Li, Libo Wu, Yuan Qi

机构 * Shanghai Academy of AI for Science(上海人工智能研究院)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 研究针对科学多领域推理需求,引入统一科学多模态模型MKB,围绕共享Transformer主干及定制组件构建,涵盖多科学分支。通过两阶段训练,在多方面实验表现出色,证明该范式可行,为跨领域科学多模态探索提供基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20449 2026-07-24 cs.CL cs.AI 新提交 62%

The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs

模型中的讲述者:大语言模型中的叙事模式继承、升级动态和对齐治理

Adam Rigby, Raz Saremi, Azadeh Sohrabinejad, Mehdi Rahimi

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 探讨大语言模型训练中是否吸收人类写作叙事模式并致输出漂移,通过文献综述和跨论文分析发现三个关键模式,包括复制训练数据模式、出现潜在特征及微调带来意外变化,指出叙事漂移是未监测的升级途径,需专门监测工具。

Comments 2 figures, 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09504 2026-07-24 cs.CR cs.AI cs.LG 版本更新 62%

AI Security Policy Should Assess Systems, Not Only Models

位置:AI安全政策应针对系统,而非模型

Michael A. Riegler, Inga Strümke

机构 * AI Safety Department(人工智能安全部门) SimulaMet and OsloMet(SimulaMet和奥斯陆计量)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出swarm-attack框架,展示通过协调轻量级LLM代理发现系统漏洞和绕过安全机制的可能性,证明在低成本硬件上可实现零成本安全测试。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13792 2026-07-24 cs.AI cs.CL 版本更新 62%

StackingNet: Collective Inference Across Independent AI Foundation Models

StackingNet:跨独立AI基础模型的集体推理

Siyang Li, Chenhao Liu, Dongrui Wu, Zhigang Zeng, Lieyun Ding

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 StackingNet通过元集成框架整合独立模型输出,提升准确率并减少误差,无需内部参数或训练数据,有效识别和修剪退化模型,在语言理解、视觉属性估计和学术论文评分中优于单一模型和经典集成方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19824 2026-07-23 cs.AI cs.CL 新提交 62%

Rewarding Better Thinking for LLM Preference Alignment

奖励更好的思维以实现大语言模型偏好对齐

Xubo Liu, Wenya Guo, Ruxue Yan, Xinying Qian, Ying Zhang

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 研究大语言模型偏好对齐问题,提出思维清单奖励(TCR)方法,将偏好对转化为思维清单评估推理轨迹,引入EMA残差公式减少与结果监督重叠,实验证明该方法能提升对齐性能。

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19674 2026-07-23 cs.CR cs.AI cs.LG 新提交 62%

FedLSG: LLM-Enhanced Semantic Calibration for Federated Graph Backdoor Defense

FedLSG:用于联邦图后门防御的大语言模型增强语义校准

Chenyu Zhou, Yabin Peng, Wei Huang, Kunlin Li, Shuaishuai Zhang, Xinyuan Miao

机构 * Southeast University(东南大学) Purple Mountain Laboratories(紫金山实验室) Institute of AI for Industries(人工智能产业研究院) Chinese Academy of Sciences(中国科学院)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 针对联邦图神经网络易受后门攻击问题,提出FedLSG框架,将大语言模型集成防御,通过图与行为到文本的转换及轻量级师生架构,在不损图完整性时显著提升对后门攻击的抵抗力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00821 2026-07-23 cs.AI cs.CL cs.IR 版本更新 62%

Fidelity Before Structure: Verbatim Chunks Beat Lossy Artifact Extraction in Long-Conversation LLM Memory

逐字块胜过提取的人工制品:长LLM对话中记忆表征的控制消融研究

Tao An

机构 * Hawaii Pacific University(夏威夷太平洋大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 通过控制消融实验,发现逐字对话块在长对话记忆检索中比LLM提取的结构化人工制品(事实、决策等)准确率高15.9-22.0点,原因是提取过程丢失了逐字细节,而结构化记忆应作为逐字文本的补充而非替代。

Comments v4: title and abstract aligned with ARR August 2026 submission; six confound controls; 34 pages, 6 figures. Code: https://github.com/tao-hpu/cog-canvas

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19331 2026-07-22 cs.LG cs.AI 新提交 62%

ISO: An RLVR-Native Optimization Stack

ISO:一种原生 RLVR 的优化栈

Hanqing Zhu, Wenyan Cong, Zhizhou Sha, Sagnik Mukherjee, Xinyuan Song, David González-Martínez, Xiaoxia Wu, Yuandong Tian, Shiwei Liu, David Z. Pan, Zhangyang "Atlas" Wang

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) UIUC(伊利诺伊大学香槟分校) Emory University(埃默里大学) Together AI(联合人工智能公司) Recursive Superintelligence Inc(递归超级智能公司) ELLIS Institute Tübingen(图宾根埃利斯研究所)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 研究 RLVR 中缺失的优化层,提出等谱优化(ISO)框架,包括离线的 ISO-Merger 和在线的 ISO-Optimizer。通过光谱继承,在推理和编码任务中,用更少训练步骤提高准确率,为 RLVR 优化层问题提供了具体解决方案。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18570 2026-07-22 cs.CL cs.AI 新提交 62%

For What Reason? Interpreting Models' Encoding of Causation and Antithesis

出于什么原因?解释模型对因果关系和对立关系的编码

Abhidip Bhattacharyya, Shira Wein

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) University of South Florida(南佛罗里达大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 研究指令微调的Transformer模型对英语话语关系的编码,聚焦因果和对立关系。通过下一个token预测任务及可解释性技术,发现早期层在序列中间做预测,中层接近末尾确定决策,部分层有答案偏好,揭示话语推理的不对称表示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18567 2026-07-22 cs.AI cs.CR cs.LG 新提交 62%

Attacking Graph Foundation Models Through Their Shared Representation

通过共享表示攻击图基础模型

Pankaj Kumar, Subhankar Mishra

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 研究通过攻击图基础模型的共享表示来探索其脆弱性,采用表示空间扰动和输入空间攻击等方法,针对六个公共模型进行实验,发现不同模型的脆弱性表现,如OpenGraph的频谱分词器有特定脆弱性,还分析了脆弱性与解码器读取表示的关系。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03691 2026-07-22 cs.SE cs.AI cs.LG 版本更新 62%

Don't Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent Quality

不要责怪大语言模型:脚手架演化如何塑造编码代理质量

Oussama Ben Sghaier, Hao Li, Bram Adams, Ahmed E. Hassan

机构 * Queen’s University Canada(加拿大女王大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 研究编码代理中脚手架演化对其质量的影响。通过固定模型、仅改变脚手架,对Qwen Code CLI的35个连续版本进行评估,追踪质量波动与开发模式及架构组件的关系。

详情

展开后加载摘要…

URL PDF HTML 收藏