arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-07-31 至 2026-07-31 共收录 349 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 81 篇

2601.07853 2026-07-31 cs.CR cs.AI 版本更新 70%

FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments

FinVault:在执行 grounded 环境中评估金融代理安全性的基准测试

Zhi Yang, Runguo Li, Qiqi Qiang, Jiashun Wang, Fangqi Lou, Mengping Li, Dongpo Cheng, Rui Xu, Heng Lian, Shuo Zhang, Xiaolong Liang, Xiaoming Huang, Zheng Wei, Zhaowei Liu, Xin Guo, Huacan Wang, Ronghao Chen, Liwen Zhang

机构 * SUFE(上海财经大学) CUHKSZ(香港中文大学) SUIBE(上海对外经贸大学) FDU(福建大学) XDU(西安电子科技大学) BUPT(北京邮电大学) Tencent(腾讯) UCAS(中国科学院大学) PKU(北京大学) QuantaAlpha(量子Alpha)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 FinVault 是首个针对金融代理的执行 grounded 安全基准测试,通过31个监管案例驱动的沙盒场景和963个测试用例,评估金融代理在现实金融环境中的安全性和防御有效性。

Comments After further review of the current submission, we have identified potential legal and intellectual property concerns associated with keeping the manuscript publicly available as a preprint. In particular, there are ongoing considerations regarding institutional affiliation information, intellectual property ownership, and related compliance matters

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28133 2026-07-31 econ.GN q-fin.EC 新提交 67%

AI Sycophancy and Decisions

AI 奉承与决策

John Conlon, Peter Schwardmann

专题命中 评测与基准 :LLM(abstract,abstract_cn)

AI总结 该研究通过1500名参与者的30项决策实验,发现奉承式AI建议平均使选择去极化,虽奉承程度会削弱该效应,且市场力量不会引发更大极化效应,缓解了相关担忧。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27893 2026-07-31 cs.SE 新提交 67%

A comparative analysis of automated techniques for security bug report identification

安全漏洞报告自动识别技术的对比分析

Muhammad Laiq

专题命中 评测与基准 :large language model(abstract);language model(abstract)

AI总结 本文对比分析了逻辑回归、SetFit等多种安全漏洞报告自动识别技术,发现SetFit整体性能最优,GPT-5.2表现较差,迁移学习对不同数据量项目的性能影响存在差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27806 2026-07-31 cs.CV 新提交 67%

LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

LoMeVQA:纵向医学视觉问答综合基准

Zhilin Wu, Zhangkai Ni, Chengmei Yang, Longzhen Yang, Yihang Liu, Ying Wen, Lianghua He

机构 * Tongji University(同济大学) East China Normal University(华东师范大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract)

AI总结 该研究提出纵向医学视觉问答基准LoMeVQA,发现现有多模态大语言模型在该任务上时间推理能力不足,推出MedLong-8B实现最优性能,并开展相关分析。

Comments 23 pages, 17 figures, 7 tables. Code and data: https://github.com/pepperbubble/LoMeVQA

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27378 2026-07-31 cs.CV cs.MM 新提交 67%

PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology

PanDent:面向牙科放射学中全面的牙级结构-语言一致性

Xiaohan Li, Xinyu Liu, Chang Liu, Sum Wing Au Yeung, Jun Liu, Yixuan Yuan, Hui Chen

机构 * Faculty of Dentistry, The University of Hong Kong(香港大学牙医学院) Imperial College London(帝国理工学院) University of Science and Technology of China(中国科学技术大学) Department of Data and Systems Engineering, The University of Hong Kong(香港大学数据与系统工程系) Department of Electronic Engineering, The Chinese University of Hong Kong(香港中文大学电子工程系)

专题命中 评测与基准 :large language model(abstract);language model(abstract)

AI总结 本研究推出PanDent牙科OPG基准,经实验发现现有MLLM生成的牙科报告流畅但临床一致性差,在PanDent上微调可提升其结构-语言一致性,该基准可用于评估MLLM的牙级临床推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28210 2026-07-31 physics.ed-ph 新提交 67%

AI-based scoring systematically underestimates conceptual understanding of linguistically weak students' explanations in physics

基于人工智能的评分系统低估了语言薄弱学生在物理解释中的概念理解能力

Markus S. Feser, Paul L. Tschisgale

专题命中 评测与基准 :large language model(abstract);language model(abstract)

AI总结 该研究发现,所有AI评分方法均存在低估语言薄弱学生物理解释中概念理解的语言偏见,类似物理教师的相关偏见,多语言学习者受影响最大。

Comments Shared lead authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09306 2026-07-31 cs.CL cs.AI cs.HC cs.LG 版本更新 67%

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

创造力、诚实和设计遗忘在小型双曲语言模型中出现

Kwan Soo Shin

机构 * PolymathMinds Lab(多智思维实验室) POSTECH(浦项科技大学) aSSIST University(aSSIST大学) Korean Educational Development Institute(韩国教育开发院)

专题命中 评测与基准 :language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究探讨小型双曲语言模型如何成为可信陪伴式AI,通过行为审计器检测合规差距,用线性读出检测相关问题,创意框架播种器更优,记忆操作系统实现设计遗忘,为可信陪伴式AI提供了小模型路径。

Comments Substantially revised and narrowed version with a new title and estimand-centred analysis. Comparisons are now reported at three output resolutions, and the reproducibility package has been rebuilt. The author list was changed with the approval of all authors listed on v1-v2; previous versions remain publicly available. 17 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28274 2026-07-31 cs.CL cs.LG 新提交 62%

MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek

MORFES:现代希腊语屈折生成能力基准测试集

Ioakeim Perros, Cleopatra Papadopoulou, Ayoub Kirouane, Christos Petrocheilos

机构 * Sophea AI

专题命中 评测与基准 :language model(abstract);分类 cs.CL、cs.LG

AI总结 该研究针对现代希腊语屈折生成能力缺乏专用基准的问题,构建了含500个条目的MORFES基准,评估多款开源模型,其中自研的Sophea-Genesis-1在屈折形态任务中表现领先。

Comments 12 pages, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27393 2026-07-31 cs.CL cs.AI 新提交 62%

AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes

AHA-Memes:用于理解阿拉伯语表情包中仇恨内容的细粒度多模态基准

Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Abul Hasnat, Md. Rafiul Biswas, Wajdi Zaghouani, Firoj Alam

机构 * Qatar Computing Research Institute(卡塔尔计算研究所) Hamad Bin Khalifa University(哈马德·本·哈利法大学) Northwestern University in Qatar(卡塔尔西北大学)

专题命中 评测与基准 :language model(abstract);分类 cs.CL、cs.AI

AI总结 该研究推出首个带细粒度多标签标注的阿拉伯语仇恨表情包基准AHA-Memes,构建含5000张人工标注及6.6万张银标的数据集,对多类模型基准测试并发布资源以推动相关研究。

Comments 26 pages, 14 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27217 2026-07-31 stat.AP cs.LG 新提交 57%

Foundation-Model Earth Representations Enable Regional-Scale Forest Aboveground Biomass Monitoring Across the Northeastern United States

基于基础模型的地球表征实现美国东北部区域尺度森林地上生物量监测

Shashika Lamahewage, Chandi Witharana

专题命中 评测与基准 :foundation model(abstract);分类 cs.LG

AI总结 该研究利用AlphaEarth基础模型生成的Google卫星嵌入,结合LiDAR数据与森林清查数据构建模型,实现美国东北部区域森林地上生物量的高精度监测,为规模化碳评估提供了新途径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28496 2026-07-31 cs.CL 新提交 57%

Beyond Sentiment: Structured Information Extraction from Financial News

超越情感:从金融新闻中进行结构化信息提取

Daohan Zhu, Sitong Ge, Ruofei Wang, Honggu Chen, Yubo Hou, Tao Wan, Zengchang Qin

机构 * School of ASEE, Beihang University(北京航空航天大学ASEE学院) School of BME, Beihang University(北京航空航天大学生物医学工程学院) CAIR and CECS, VinUniversity(VinUniversity CAIR与CECS机构)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL

AI总结 该研究针对金融情感分析仅压缩新闻为单一极性评分的局限,提出用LLaMA-3.1-70B提取金融新闻的结构化语义特征,结合情感与结构化特征可提升股票预测性能,为多维金融NLP开辟新方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28187 2026-07-31 cs.AI cs.CR cs.SE 新提交 57%

Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation

旧技巧,新模型:简单图像变换如何破解基于现代AI的内容审核

Marco Alecci, Francesco Marchiori, Iyiola Emmanuel Olatunji, Tegawendé F. Bissyandé, Jacques Klein

专题命中 评测与基准 :foundation model(abstract);分类 cs.AI

AI总结 该研究通过黑盒评估发现,商用多模态图像审核API可被简单图像变换绕过,且不同服务稳健性差异大,表明这类API无法单独作为可靠安全过滤器,需结合分层审核管道部署。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27826 2026-07-31 cs.AI cs.CV 新提交 57%

Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding

手语问答:手语理解的新任务、基准与基线模型

Shiwei Gan, Lichen Wang, Xiao Liu, Yafeng Yin, Kuizhuang Liu, Sanglu Lu, Lei Xie

专题命中 评测与基准 :language model(abstract);分类 cs.AI

AI总结 该研究提出手语问答新任务,构建基于PHOENIX14T与CSL-Daily的SignQA基准,设计带特定模块的基线模型,实验显示其在各问题类别上优于代表性视觉-语言模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27739 2026-07-31 cs.IR cs.CL cs.HC 新提交 57%

Measuring Alignment With Reader Highlights Net of Position and Length

通过读者标注的位置和长度网络测量对齐度

Kazuki Nakayashiki, Keisuke Watanabe

专题命中 评测与基准 :language model(abstract);分类 cs.CL

AI总结 该研究针对上下文压缩评估的混淆问题,提出匹配位置长度的方法,发现语言模型重要性排名对齐度优于人类读者外的基准,且前期某结论无法在当前语料库复现。

Comments 15 pages, 7 tables. Analysis code and de-identified artifacts included as ancillary files; five of six scripts reproduce the paper's numbers from the shipped artifacts alone. Reports claims from our own prior work that this corpus does not reproduce, and lists twelve claims withdrawn during internal adversarial review in Appendix A

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27597 2026-07-31 cs.RO cs.AI cs.SY eess.SY 新提交 57%

A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response

面向视觉语言赋能的无人机分类与灾害响应的系统工程框架

Swapnil Saha, Bhuvan Rajanasiriyur Jagadeesha, Karishma Patnaik, Neelakshi Majumdar

机构 * University of Arkansas(阿肯色大学) University of Michigan-Dearborn(密歇根大学迪尔伯恩分校)

专题命中 评测与基准 :language model(abstract);分类 cs.AI

AI总结 该研究提出将视觉语言模型(VLMs)作为协同智能体嵌入人-无人机回路的系统工程框架,经评估可降低人因负担、提升AI信任度,推进了高风险灾害响应的人-自主系统协同。

Comments 10 pages, 8 figures. Author accepted manuscript of AIAA Paper 2026-4010, published in the AIAA AVIATION 2026 Forum

Journal ref AIAA AVIATION 2026 Forum, AIAA Paper 2026-4010, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27263 2026-07-31 cs.LG physics.data-an stat.ME 新提交 57%

DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series

DoTime:用于干预和反事实时间序列的合成基准生成器

Dennis Thumm, Billy Tim Anthony, Ying Chen

机构 * National University of Singapore(新加坡国立大学)

专题命中 评测与基准 :foundation model(abstract);分类 cs.LG

AI总结 DoTime是用于干预和反事实时间序列的合成基准生成器,具备现有工具没有的多项功能,附带评估套件与基线,经测试其干预训练的因果模型相比同容量观测模型有方向准确性优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21419 2026-07-31 cs.AI 版本更新 57%

PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning

PATS:用于智能体强化学习的策略感知训练框架

Yipeng Shi, Zhipeng Ma, Yue Wang, Qitai Tan, Yang Li, Peng Chen, Zhengzhou Zhu

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

AI总结 研究针对长期语言模型智能体强化学习中弱策略问题,提出以策略为中心的训练范式Pats,将技能作为动态训练框架,通过转换展开组为证据卡、特定任务评估调整上下文等改进策略,在多个任务上取得良好效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15297 2026-07-31 cs.CL 版本更新 57%

AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports

AfriEconQA:基于世界银行报告的非洲经济分析基准数据集

Edward Ajayi, Mustapha Alaba, David Stephen

机构 * Carnegie Mellon University Africa(卡内基梅隆大学非洲分校)

专题命中 评测与基准 :LLM(abstract_cn);分类 cs.CL

AI总结 AfriEconQA是一个基于世界银行报告的非洲经济分析基准数据集,旨在测试信息检索和RAG系统在处理复杂经济查询时的性能。

Comments Dataset Explorer: https://afrieconqa.pages.dev/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25991 2026-07-31 cs.AI cs.CV 版本更新 57%

Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline

面向社交媒体中统一的多模态虚假信息检测:基准数据集与基线模型

Haiyang Li, Yaxiong Wang, Shengeng Tang, Yuchen Zhang, Lianwei Wu, Lechao Cheng, Liu Liu, Chaofeng Dong, Zhun Zhong

机构 * School of Computer Science and Information Engineering, Hefei University of Technology(计算机科学与信息工程学院,合肥工业大学) School of Computer Science and Technology, Northwestern Polytechnical University(计算机科学与技术学院,西北工业大学)

专题命中 评测与基准 :language model(abstract);分类 cs.AI

AI总结 该研究构建了含9.8万样本的OmniFake基准数据集,提出UMFDet框架,实现对人工与AI生成两类多模态虚假内容的统一检测,性能优于专用基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28532 2026-07-31 cs.CV 新提交 50%

MarkushGlyph and OCSRGlyph: Improved Chemical Structure Recognition

MarkushGlyph与OCSRGlyph:改进的化学结构识别

Alex Andonian, Samuel G Rodriques, Andrew D White, Siddharth M Narayanan

专题命中 评测与基准 :language model(abstract)

AI总结 本研究将化学结构识别视为图像到文本转换任务,提出OCSRGlyph与MarkushGlyph模型,前者优化OCSR性能,后者针对Markush结构采用整体视觉语言建模,还提出新指标解决Markush结构翻译的准确性判定问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27011 2026-07-31 eess.AS 版本更新 50%

Qwen-Audio-3.0-Gen-Preview Technical Report

Qwen-Audio-3.0-Gen-Preview 技术报告

Junyu Dai, Xiaoyue Duan, Xinyue Fan, Yihan Feng, Jingbei Li, Xiangang Li, Yunjia Li, Lejun Min, Yufei Shi, Xingchen Song, Yiran Wang, Cheng Wen, Menglin Wu, Bajian Xiang, Huaicheng Zhang, Han Zhao, Ruichen Zheng

专题命中 评测与基准 :language model(abstract)

AI总结 针对现有音频系统难以组织多类型音频形成长时场景的问题,提出 Qwen-Audio-3.0-Gen-Preview 框架,采用 DiT 与共享 VAE 实现统一多域音频生成,在多个基准上展现出性能优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26518 2026-07-31 cs.CV 版本更新 50%

EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

EgoSafe:用于视觉安全理解的第一人称移动采集基准

Yuyun Chen, Tianao Li, TianQuan Feng, Cen Chen, Huiping Zhuang, Hao Peng, Ziqian Zeng

机构 * South China University of Technology(华南理工大学) Beihang University(北京航空航天大学)

专题命中 评测与基准 :language model(abstract)

AI总结 该研究推出第一人称移动采集的视觉安全理解基准 EgoSafe-Bench,含12000个样本,用分层推理评估(HRE)协议测试,发现现有大视觉语言模型存在感知-推理解耦问题,为逻辑鲁棒视频理解系统提供评估框架

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25946 2026-07-31 cs.SE 版本更新 50%

A Low-Cost Human-in-the-Loop Investigation of Toxicity on GitHub at Scale

大规模低成本的 GitHub 毒性人工参与调查

Rahat Rizvi Rahman, Mia Mohammad Imran, Kostadin Damevski

专题命中 评测与基准 :LLM(abstract)

AI总结 研究 GitHub 毒性标注问题,提出人工参与注释方法,通过小型本地语言模型预测及随机森林验证器筛选,降低标注成本,应用此流程处理大量对话,评估先前研究发现并给出新见解。

Comments Accepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21076 2026-07-31 cs.CV 版本更新 50%

PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation

PoseMaster: 一个统一的3D原生框架用于风格化姿态生成

Hongyu Yan, Kunming Luo, Weiyu Li, Kaiyi Zhang, Yixun Liang, Jingwei Huang, Chunchao Guo, Ping Tan

机构 * Hong Kong University of Science and Technology(香港科技大学) Tencent Hunyuan(腾讯文脉)

专题命中 评测与基准 :foundation model(abstract)

AI总结 本文提出PoseMaster框架,通过统一姿态风格化与3D生成,提升3D姿态风格化的精度和多样性,采用3D骨架直接指导生成,增强模型在姿态风格化中的表现。

Comments Accepted by CVPR 2026, Code: https://github.com/hanryyan/PoseMaster

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21007 2026-07-31 cs.CV 版本更新 50%

MetaRank: Task-Aware Metric Selection for Model Transferability Estimation

MetaRank: 为模型可迁移性估计的任务感知度量选择

Yuhang Liu, Wenjie Zhao, Xin Wang, Yunhui Guo

机构 * Fudan University(复旦大学) University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 评测与基准 :language model(abstract)

AI总结 MetaRank通过元学习框架实现任务感知的MTE度量选择,提升模型迁移性估计的准确性。

Comments 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10117 2026-07-31 cs.RO 50%

Do Visual-Language Grid Maps Capture Latent Semantics?

Matti Pekkanen, Tsvetomila Mihaylova, Francesco Verdoja, Ville Kyrki

机构 * School of Electrical Engineering, Aalto University(奥卢大学电气工程学院)

专题命中 评测与基准 :language model(abstract)

Comments IROS 2025

Journal ref 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Hangzhou, China, 2025, pp. 4059-4066

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 效率与部署 59 篇

2607.27704 2026-07-31 cs.AR cs.LG 新提交 92%

LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference

LightRot:用于精确低比特大语言模型推理的轻量级旋转方案与架构

Sangjin Kim, Yuseon Choi, Jungjun Oh, Byeongcheol Kim, Hoi-Jun Yoo

专题命中 效率与部署 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.LG

AI总结 本研究提出LightRot轻量级旋转方案与硬件加速器,通过算法创新结合28nm工艺实现4比特推理27.4 TOPS/W能效,适配LLaMA系列模型,为低比特LLM推理提供新范式。

Comments 13 pages, journal version. Published in IEEE Journal on Emerging and Selected Topics in Circuits and Systems (JETCAS), vol. 15, no. 2, pp. 231-243, 2025, DOI: 10.1109/JETCAS.2025.3558300

Journal ref IEEE Journal on Emerging and Selected Topics in Circuits and Systems, vol. 15, no. 2, pp. 231-243, June 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27918 2026-07-31 math.OC 新提交 92%

OptGraph: Large Language Models Enhanced Evolutionary Optimization Via Graph Retrieval-Augmented Generation

OptGraph:基于图检索增强生成的大语言模型增强进化优化

Xianchao Xiu, Jianhao Li, Huangyue Chen, Wanquan Liu

专题命中 效率与部署 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract)

AI总结 针对现有LLM自动化进化优化的局限,提出引入GraphRAG的OptGraph工作流,通过类型化图复用经验等方法,在基准数据集上较现有框架平均精确准确率提升8.9%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24162 2026-07-31 cs.CR cs.AI 版本更新 92%

Defusing the Trigger: Tail-Risk-Informed Attention Rebalancing for LLM Backdoor Mitigation

解除触发器:通过尾风险内在几何平滑实现的插件式防御方法用于后门LLM

Kaisheng Fan, Yishu Gao, Xunzhu Tang, Tegawendé F. Bissyandé, Weizhe Zhang

机构 * School of Cyber Science and Technology(网络安全科学与技术学院) Harbin Institute of Technology(哈尔滨工业大学) Peng Cheng Laboratory(鹏城实验室) University of Luxembourg(卢森堡大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出TIGS方法,通过在推理时进行内容感知的尾风险筛查和内在几何平滑,有效抑制后门攻击,同时保持推理稳定性和语义一致性,适用于多种架构的LLM。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27698 2026-07-31 cs.AI cs.NE 新提交 90%

Guiding Large Language Models with Genetic Programming-Evolved Heuristic Knowledge for Dynamic Multi-Mode Project Scheduling

用遗传编程演化的启发式知识指导大语言模型解决动态多模式项目调度问题

Yuan Tian, Yi Mei, Mengjie Zhang

机构 * Centre for Data Science and Artificial Intelligence(数据科学与人工智能中心) School of Engineering and Computer Science(工程与计算机科学学院) Victoria University of Wellington(惠灵顿维多利亚大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI

AI总结 本研究将遗传编程演化的调度规则知识反向转移,通过四种机制指导大语言模型决策,提升了调度性能、决策稳定性并降低 token 消耗。

详情

展开后加载摘要…

URL PDF HTML 收藏