arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-05-18 至 2026-05-18 共收录 320 信号源:cs.CL, cs.AI, cs.LG

1. 其他LLM 21 篇

2604.21251 2026-05-18 cs.LG cs.AI 89%

CAP: Controllable Alignment Prompting for Unlearning in LLMs

CAP:用于大语言模型中去学习的可控对齐提示

Zhaokun Wang, Jinyu Guo, Jingwen Pu, Hongli Pu, Meng Yang, Xunlei Chen, Jie Ou, Wenyi Li, Guangchun Luo, Wenhong Tian

机构 * School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件学院)

专题命中 其他LLM :prompting(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文提出CAP框架,通过强化学习将去学习过程转化为可学习的提示优化,实现可控的去学习,无需更新模型参数,解决了现有方法的计算成本高、遗忘边界不可控等问题。

Comments Accpeted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15918 2026-05-18 q-bio.QM 89%

The Impact of Heatwaves on Population Health: A Large Language Model-Enhanced Agent-Based Simulation

热浪对人口健康的影响:一种增强型大语言模型的群体模拟

Yuanhao Liu, Yuanfei Liu, Tian Lu, Hengyang Zhang, Zuowei Wang, Ying Dai

专题命中 其他LLM :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 本文通过增强型大语言模型进行群体模拟,研究热浪对人口健康的影响,发现心理社会因素在社区韧性中的作用,并提出针对脆弱群体的干预措施。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04126 2026-05-18 eess.SP 89%

Semantic Pilot Design for Data-Aided Channel Estimation Using a Large Language Model

基于大型语言模型的语义前导符设计用于数据辅助信道估计

Sojeong Park, Hyun Jong Yang

专题命中 其他LLM :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 本文提出利用大型语言模型设计语义前导符,用于文本包含数据传输中的数据辅助信道估计。通过比对初始解码文本与语言模型校正版本,识别可靠解码符号,提升信道估计性能。

Journal ref ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15842 2026-05-18 physics.soc-ph cs.SI 88%

Reconstructing temporal multi-relational firm networks at scale using large language models. The case of the semiconductor industry

利用大语言模型重建大规模时间多关系企业网络:半导体行业案例

Seyda Köse, Christian Diem, Elma Dervic, Klaus Friesenbichler, Georg Heiler, Jan Hurt, Hernan Picatto, Peter Klimek

专题命中 其他LLM :large language model(title,abstract);language model(title,abstract)

AI总结 本文利用大语言模型和开放网络数据重建半导体行业企业网络,识别供应链、合作关系和所有权链接,揭示2022年芯片短缺期间的网络变化及AI供应链瓶颈企业的中心性变化。

Comments 32 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14200 2026-05-18 cs.CL cs.AI cs.LG 82%

A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement

可扩展的多语言模型协作系统:基于检索的选择与探索-利用驱动增强

Shengji Tang, Jianjian Cao, Weihao Lin, Jiale Hong, Bo Zhang, Shuyue Hu, Lei Bai, Tao Chen, Wanli Ouyang, Peng Ye

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 其他LLM :LLM(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出SMCS系统,通过检索优先选择模块和探索-利用驱动后验增强模块,有效协调多个开源语言模型,实验显示其在多个任务中优于闭源模型,且在不同数据集上超越开源模型的平均最佳结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16245 2026-05-18 cs.CY cs.AI cs.CL cs.LG cs.SI 80%

AI-Mediated Communication Can Steer Collective Opinion

AI介导的交流可以引导集体意见

Stratis Tsirtsis, Kai Rawal, Chris Russell, Brent Mittelstadt, Sandra Wachter

机构 * Hasso Plattner Institute(哈索普兰特纳研究所) Oxford Internet Institute, University of Oxford(牛津互联网研究所,牛津大学) Weizenbaum Institute(魏泽纳姆研究所)

专题命中 其他LLM :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文研究AI在人类间交流中对集体意见形成的影响,通过实证和理论分析展示AI引入的方向性偏见如何通过网络放大并改变集体观点,探讨平台如何控制此类偏见。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15676 2026-05-18 cs.CL 79%

Dynamic Chunking for Diffusion Language Models

扩散语言模型的动态分块

Yichen Zhu, Xiaoming Shi, Peng Zhao, Weiyu Chen, Debing Zhang, James Kwok

机构 * CSE, HKUST(香港科技大学计算机科学与工程系) Xiaohongshu Inc.(小红书公司) Alibaba group(阿里巴巴集团) CityUHK(城市大学香港校区)

专题命中 其他LLM :language model(title,abstract);分类 cs.CL

AI总结 本文提出动态分块扩散模型,通过内容定义语义分块替代固定位置分块,提升序列结构利用效率,在参数规模达1.5B的下游任务中表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15562 2026-05-18 cs.CL 79%

GiLT: Augmenting Transformer Language Models with Dependency Graphs

GiLT:通过依赖图增强Transformer语言模型

Tianyu Huang, Yida Zhao, Chuyan Zhou, Kewei Tu

机构 * School of Information Science and Technology, ShanghaiTech University(信息科学与技术学院,上海科技大学) Shanghai Engineering Research Center of Intelligent Vision and Imaging(智能视觉与成像上海工程研究中心)

专题命中 其他LLM :language model(title,abstract);分类 cs.CL

AI总结 GiLT通过依赖图增强Transformer语言模型,提升语法泛化能力,同时保持竞争力的困惑度,且能通过微调提升下游任务表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15440 2026-05-18 cs.CL 79%

Why are language models less surprised than humans? Testing the Parse Multiplicity Mismatch Hypothesis

为何语言模型比人类更不惊讶?测试解析多重性不匹配假说

William Timkey, Brian Dillon, Tal Linzen

机构 * Department of Linguistics, New York University(纽约大学语言学系) Center for Data Science, New York University(纽约大学数据科学中心) Department of Linguistics, University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校语言学系)

专题命中 其他LLM :language model(title,abstract);分类 cs.CL

AI总结 研究探讨语言模型在处理句子时的 surprisal 与人类处理的差异,通过调整解析数量来测试解析多重性不匹配假说对句法歧义的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15613 2026-05-18 cs.CL 77%

Toward LLMs Beyond English-Centric Development

迈向超越英语中心化发展的语言模型

Sho Takase, Ukyo Honda

机构 * CyberAgent

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究发现语言模型对英语存在显著偏见,持续预训练并非优于从头训练的低成本方案,未来需加强多语言投入。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02453 2026-05-18 cs.LG cs.AI cs.CL 75%

How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models

如何训练你的导师:通过导师模型引导黑盒大语言模型

Parth Asawa, Alan Zhu, Abigail O'Neill, Matei Zaharia, Alexandros G. Dimakis, Joseph E. Gonzalez

机构 * University of California, Berkeley(加州大学伯克利分校) Bespoke Labs(Bespoke实验室)

专题命中 其他LLM :language model(abstract);prompting(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出Advisor Models,通过训练小型开放权重模型生成动态个性化建议,提升黑盒前沿模型性能,实验显示在多个任务中效果显著,且具有良好的迁移性和鲁棒性。

Comments International Conference on Machine Learning (ICML) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26733 2026-05-18 cs.AI cs.LG 73%

FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards

FutureWorld: 一个用于预测代理的实时强化学习环境,具有现实世界结果奖励

Zhixin Han, Yanzhi Zhang, Chuyang Wei, Maohang Gao, Xiawei Yue, Kefei Chen, Yu Zhuang, Haoxiang Guan, Jiyan He, Jian Li, Yitong Duan, Yu Shi, Mengting Hu, Shuxin Zheng

机构 * College of Software, Nankai University(南开大学软件学院) Academy of Mathematics and Systems Science, Chinese Academy of Sciences(中国科学院数学与系统科学研究院) School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) IIIS, Tsinghua University(清华大学智能系统与信息工程研究院) Zhongguancun Academy, Beijing, China(北京中关村学院)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出FutureWorld,一个实时强化学习环境,通过闭环预测、结果实现与参数更新,提升预测准确性与校准能力。

Comments The code will be released in the near future. The experiments are currently ongoing

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15763 2026-05-18 cs.CL cs.AI 73%

CompactQE: Interpretable Translation Quality Estimation via Small Open-Weight LLMs

CompactQE: 通过小规模开源大语言模型实现可解释的翻译质量估计

Kamil Guttmann, Zofia Fraś, Artur Nowakowski, Krzysztof Jassem

机构 * Laniqo Faculty of Mathematics and Computer Science, Adam Mickiewicz University(亚当·密茨凯维奇大学数学与计算机科学学院)

专题命中 其他LLM :LLM(abstract_cn);prompting(abstract);分类 cs.CL、cs.AI

AI总结 本文提出CompactQE,利用小规模开源大语言模型实现翻译质量估计,生成质量评分、错误标注、修正建议和完整润色,其性能优于传统指标和人类标注。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05966 2026-05-18 cs.CL 70%

FinReporting: An Agentic Workflow for Localized Reporting of Cross-Jurisdiction Financial Disclosures

FinReporting: 一种用于跨司法管辖区财务披露本地化报告的代理工作流

Fan Zhang, Mingzi Song, Rania Elbadry, Yankai Chen, Shaobo Wang, Yixi Zhou, Xunwen Zheng, Yueru He, Yuyang Dai, Georgi Georgiev, Ayesha Gull, Muhammad Usman Safder, Fan Wu, Liyuan Meng, Fengxian Ji, Junning Zhao, Xueqing Peng, Jimin Huang, Yu Chen, Xue, Liu, Preslav Nakov, Zhuohan Xie

机构 * MBZUAI The University of Tokyo(东京大学) Meiji Gakuin University(明治大学) McGill University(麦吉尔大学) Kyoto University(京都大学) Columbia University(哥伦比亚大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出FinReporting,一种代理工作流,用于跨司法管辖区的财务披露本地化报告。该系统构建了涵盖损益表、资产负债表和现金流量表的统一本体,将报告分解为可审计的阶段,并通过约束验证器提升一致性和可靠性。

Comments Accepted at ACL 2026 Demo Track. 9 pages, including figures and tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22440 2026-05-18 cs.CY cs.LG cs.MA econ.GN q-fin.EC 70%

From Model Design to Organizational Design: Complexity Redistribution and Trade-Offs in Generative AI

从模型设计到组织设计:生成AI中的复杂性再分配与权衡

Sharique Hasan, Alexander Oettl, Sampsa Samila

机构 * Duke University(杜克大学) Georgia Institute of Technology(佐治亚理工学院) IESE Business School(IESE商学院)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出GAS框架,分析大语言模型如何重塑组织与竞争策略,揭示生成AI中通用性、准确性与简洁性之间的权衡及复杂性再分配对管理挑战的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02832 2026-05-18 cs.CR cs.AI 70%

FlipAttack: Jailbreak LLMs via Flipping

FlipAttack:通过翻转对大型语言模型进行劫持

Yue Liu, Xiaoxin He, Miao Xiong, Jinlan Fu, Shumin Deng, Yingwei Ma, Jiaheng Zhang, Bryan Hooi

机构 * Engineering Programme, NUS Graduate School, National University of Singapore(国立新加坡大学整合科学与工程计划) Institute of Data Science (IDS), National University of Singapore(国立新加坡大学数据科学研究所) Department of Computer Science, School of Computing, National University of Singapore(国立新加坡大学计算机科学系)

专题命中 其他LLM :LLM(summary_cn);分类 cs.AI

AI总结 FlipAttack通过在提示左侧添加噪声,利用语言模型自左向右理解文本的特性,实现对黑盒LLM的高效劫持,实验显示其在多个模型上均表现出高成功率。

Comments 43 pages, 31 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07720 2026-05-18 cs.RO 67%

Empowering Robot Teleoperation: Exploring the Synergies Between Devices and Manipulator Controllers in a Comparative Study

赋予机器人远程操控能力:在比较研究中探索设备与机械臂控制器之间的协同效应

Yuxuan Zhao, Yuanchen Tang, Jindi Zhang, Hongyu Yu

机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, Guangdong, China(香港中文大学(深圳)科学与工程学院) Department of Mechanical and Aerospace Engineering, Hong Kong University of Science and Technology(香港科学与技术大学机械与航空航天工程系) Shenzhen Institute of Artificial Intelligence and Robotics for Society (AIRS), Shenzhen, Guangdong, China(深圳人工智能与机器人社会研究院(AIRS))

专题命中 其他LLM :large language model(abstract);language model(abstract)

AI总结 研究通过比较不同设备与控制器策略在机械臂任务数据收集中的效果,探讨设备与控制器之间的协同关系对实际任务的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12994 2026-05-18 physics.optics physics.ao-ph physics.chem-ph 67%

Trapping absorbing and non-absorbing aqueous particle using a universal 4-arm Laguerre-Gaussian mode light trap

利用通用的四臂拉盖尔-贝塞尔光陷阱捕捉吸收和非吸收的水性粒子

Krispin M. Dettlaff, James Wenger, Grégory David, Ruth Signorell

专题命中 其他LLM :SLM(abstract,abstract_cn)

AI总结 本文提出一种无需机械调整的通用光陷阱,通过四束可调的拉盖尔-贝塞尔或基础高斯光束实现对吸收和非吸收粒子的连续捕捉,用于研究大气中粒子的老化过程及水滴的光化学反应。

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25384 2026-05-18 cs.CL 57%

Wiki Dumps to Training Corpora: South Slavic Case

维基数据转训练语料:南斯拉夫语法

Mihailo Škorić, Cosimo Palma

机构 * University of Belgrade, Faculty of Mining and Geology(贝尔格莱德大学,采矿与地质学系) University of Pisa, Department of Computer Sciences(比萨大学,计算机科学系)

专题命中 其他LLM :language model(abstract);分类 cs.CL

AI总结 本文提出将维基数据转化为七种南斯拉夫语言高质量语料的流程,通过文本提取清洗和冗余过滤提升语料质量,为语言模型训练和跨语言比较提供可靠资源。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05030 2026-05-18 cs.CV 50%

LUIVITON: Learned Universal Interoperable VIrtual Try-ON

LUIVITON: 学习通用互操作虚拟试衣

Cong Cao, Xianhang Cheng, Jingyuan Liu, Yujian Zheng, Zhenhui Lin, Ren Li, Meriem Chkir, Hao Li

机构 * The University of Tokyo(东京大学)

专题命中 其他LLM :foundation model(abstract)

AI总结 本文提出一种全自动虚拟试衣系统,通过SMPL作为中间代理,解决服装与人体之间的对应问题,实现复杂多层服装在多样化人体上的合理 draped 效果,且支持快速定制。

详情

展开后加载摘要…

URL PDF HTML 收藏