arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2026-04-14 至 2026-04-14 共收录 233 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 61 篇

2601.14698 2026-04-14 cs.CL 57%

ClaimDB: A Fact Verification Benchmark over Large Structured Data

ClaimDB: 一个基于大规模结构化数据的事实验证基准

Michael Theologitis, Preetam Prabhu Srikar Dammu, Chirag Shah, Dan Suciu

机构 * University of Washington(华盛顿大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 本文提出ClaimDB基准,通过百万级记录和多表组合验证事实,发现传统阅读方法失效,需转向可执行程序推理,实验显示超过一半模型准确率低于55%。

Comments ACL 2026 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04932 2026-04-14 cs.CL 57%

GenProve: Learning to Generate Text with Fine-Grained Provenance

GenProve:学习生成文本的细粒度溯源

Jingxuan Wei, Xingyue Wang, Yanghaoyu Liao, Jie Dong, Yuchen Liu, Caijun Jia, Bihui Yu, Junnan Zhu

机构 * Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(计算能力网络与信息安全教育部重点实验室,山东省计算中心(国家超级计算济南中心),齐鲁工业大学(山东省科学院)) Shenyang Institute of Computing Technology, Chinese Academy of Sciences(中国科学院沈阳计算技术研究所) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统实验室) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 本文提出GenProve框架,通过结合监督微调与组相对策略优化,提升生成文本的准确性和细粒度溯源能力,优于14个强LLM。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10297 2026-04-14 cs.CV cs.AI 57%

FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data

FashionMV:基于多视角时尚数据的产品级复合图像检索

Peng Yuan, Bingyin Mei, Hui Zhang

机构 * Tsinghua University(清华大学)

专题命中 推理评测 :chain-of-thought(abstract);分类 cs.AI

AI总结 本文提出FashionMV,首个大规模多视角时尚数据集,用于产品级复合图像检索。通过构建包含127K产品、472K多视角图像和220K三元组的数据集,结合ProCIR框架,采用多阶段对话、标题对齐和思维链引导等机制,提升了检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09952 2026-04-14 cs.LG 57%

SLM Finetuning for Natural Language to Domain Specific Code Generation in Production

为生产环境中的自然语言到领域特定代码生成进行SLM微调

Renjini R. Nair, Damian K. Kowalczyk, Marco Gaudesi, Chhaya Methani

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

AI总结 本文研究了通过微调小型语言模型提升自然语言到领域代码生成的性能与效率,展示了在生产环境中优于大模型的延迟和成本效益。

Comments 11 pages (including appendix), 5 tables, 1 figure. Submitted to arXiv as a preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09889 2026-04-14 cs.AI 57%

In-situ process monitoring for defect detection in wire-arc additive manufacturing: an agentic AI approach

现场过程监控用于焊接电弧增材制造中的缺陷检测:一种基于代理的AI方法

Pallock Halder, Satyajit Mojumder

机构 * School of Mechanical and Materials Engineering, Washington State University(华盛顿州立大学机械与材料工程学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 本文提出基于代理的AI框架,用于现场过程监控以检测焊接电弧增材制造中的缺陷。通过开发处理和监控代理,结合真实数据和分类工具,实现高准确率的缺陷检测。

Comments 42 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09644 2026-04-14 cs.CY cs.AI 57%

Detecting Corporate AI-Washing via Cross-Modal Semantic Inconsistency Learning

通过跨模态语义不一致性学习检测企业AI洗钱

Zhanjie Wen, Jingqiao Guo

机构 * School of Economics and Trade, Guangdong University of Finance(广东金融学院经济贸易学院) Department of Computer Science, Faculty of Science, Hong Kong Baptist University(香港浸会大学理学院计算机科学系)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 本文提出AWASH框架,通过跨模态主张-证据推理检测企业AI洗钱,利用AW-Bench基准测试,实现高准确率的AI能力识别。

Comments 28 pages, 6 figures, Journal Submission (Finance/Accounting & Computer Science Interdiscipline), 6 tables, 40 references, trimodal benchmark (88,412 firm-quarter observations) and end-to-end multimodal detection framework for corporate AI-washing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22887 2026-04-14 cs.CL 57%

Infusing Theory of Mind into Socially Intelligent LLM Agents

将心智理论融入社交智能的大语言模型代理

EunJeong Hwang, Yuwei Yin, Giuseppe Carenini, Peter West, Vered Shwartz

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute for AI(向量人工智能研究所)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 本文提出ToMAgent,通过结合心智理论与对话前瞻训练,提升对话效果和目标达成能力,实验表明其在社交交互评估基准中优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09249 2026-04-14 cs.CV cs.IR 50%

FashionStylist: An Expert Knowledge-enhanced Multimodal Dataset for Fashion Understanding

FashionStylist: 一个增强专家知识的多模态数据集用于时尚理解

Kaidong Feng, Zhuoxuan Huang, Huizhong Guo, Yuting Jin, Xinyu Chen, Yue Liang, Yifei Gai, Li Zhou, Yunshan Ma, Zhu Sun

机构 * Yanshan University(燕山大学) Central South University(中南大学) Zhejiang University(浙江大学) Southwest University(西南大学) Singapore Management University(新加坡管理大学) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 推理评测 :reasoning(abstract)

AI总结 本文提出FashionStylist,一个专家标注的多任务时尚理解基准,支持服装到物品定位、服装补全和评估任务,提升多模态时尚系统的语义理解能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07522 2026-04-14 cs.SE 50%

Byam: Fixing Breaking Dependency Updates with Large Language Models

Byam:利用大型语言模型修复破坏性依赖更新

Frank Reyes, May Mahmoud, Federico Bono, Sarah Nadi, Benoit Baudry, Martin Monperrus

专题命中 推理评测 :reasoning(abstract)

AI总结 本文利用大型语言模型自动修复因依赖更新导致的客户端代码问题,通过BUMP数据集验证,o3-mini模型在修复构建和编译错误方面表现最佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10456 2026-04-14 cs.CV 50%

A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation

一个用于指令驱动电影视频编译的基准和多智能体系统

Peixuan Zhang, Chang Zhou, Ziyuan Zhang, Hualuo Liu, Chunjie Zhang, Jingqi Liu, Xiaohui Zhou, Xi Chen, Shuchen Weng, Si Li, Boxin Shi

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) AI Technology Center, Online Video Business Unit, Tencent PCG(腾讯PCG在线视频事业部AI技术中心) Tsinghua University(清华大学) Beijing Academy of Artificial Intelligence(北京智源人工智能研究院) State Key Lab of Multimedia Info. Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室) Nat’l Eng. Research Ctr. of Visual Technology, School of Computer Science, Peking University(北京大学计算机学院国家视觉技术工程研究中心) School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)

专题命中 推理评测 :planning(abstract)

AI总结 本文提出CineBench基准和CineAgents系统,解决电影视频编译中的上下文崩溃和时间碎片化问题,通过脚本逆向工程和迭代叙事规划生成更连贯的编译结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05510 2026-04-14 cs.CV 50%

Benchmarking Vision-Language Models under Contradictory Virtual Content Attacks in Augmented Reality

在增强现实中评估视觉-语言模型在矛盾虚拟内容攻击下的基准测试

Yanming Xiu, Zhengyuan Jiang, Neil Zhenqiang Gong, Maria Gorlatova

机构 * Duke University(杜克大学)

专题命中 推理评测 :reasoning(abstract)

AI总结 本文提出ContrAR基准,评估视觉-语言模型在增强现实中的鲁棒性,通过312个真实AR视频验证,测试11种VLMs,发现现有模型在检测对抗性内容操纵方面仍有改进空间。

Comments CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10039 2026-04-14 cs.CV 50%

Counting to Four is still a Chore for VLMs

为VLMs计数仍是一项苦差事

Duy Le Dinh Anh, Patrick Amadeus Irawan, Tuan Van Vo

机构 * MBZUAI(穆罕默德·本·扎耶德人工智能大学)

专题命中 推理评测 :reasoning(abstract)

AI总结 本文通过行为和机制分析研究VLMs的计数行为,提出COUNTINGTRICKS评估套件,揭示模型在不同布局和对抗提示下的漏洞,并验证Modality Attention Share方法能缓解计数失败问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09721 2026-04-14 cs.IR cs.MM cs.SD 50%

Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering

Jamendo-MT-QA:多轨比较音乐问答基准测试

Junyoung Koh, Jaeyun Lee, Soo Yong Kim, Gyu Hyeong Choi, Jung In Koh, Jordan Phillips, Yeonjin Lee, Min Song

机构 * Yonsei University(延世大学) KRAFTON University of Oxford(牛津大学) George Mason University(乔治梅森大学) Sungkyul University(圣洁大学) MODULABS MAAP Onoma AI

专题命中 推理评测 :reasoning(abstract)

AI总结 本文提出Jamendo-MT-QA基准测试,用于评估多轨比较音乐问答能力,基于Jamendo-QA数据集构建36519个比较问答项,采用LLM辅助生成和过滤问题,并通过自动指标和LLM-as-a-Judge评估模型性能。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他推理 40 篇

2602.10042 2026-04-14 cs.CV cs.AI 85%

Fake-HR1: Rethinking Reasoning of Vision Language Model for Synthetic Image Detection

Fake-HR1: 重新思考视觉语言模型在合成图像检测中的推理方式

Changjiang Jiang, Xinkuan Sha, Fengchang Yu, Jingjing Liu, Jian Liu, Mingqi Fang, Chenfeng Zhang, Wei Lu

机构 * Wuhan University(武汉大学) AntGroup(蚂蚁集团) Zhejiang University(浙江大学)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.AI

AI总结 Fake-HR1通过两阶段训练框架实现自适应推理,提升合成图像检测性能与效率,优于现有LLMs。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16796 2026-04-14 cs.CV 82%

RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought

RealSR-R1: 用强化学习解决真实世界图像超分辨率问题

Junbo Qiao, Miaomiao Cai, Wei Li, Xudong Huang, Jie Hu, Xinghao Chen, Shaohui Lin, Hongkai Xiong

机构 * School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) University of Science and Technology of China(中国科学技术大学) Shanghai Jiaotong University(上海交通大学)

专题命中 其他推理 :chain-of-thought(title);reasoning(abstract);CoT(abstract)

AI总结 本文提出RealSR-R1,通过视觉-语言链式推理框架和Group Relative Policy Optimization方法,提升真实世界图像超分辨率的生成质量与内容理解能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11791 2026-04-14 cs.LG cs.AI 81%

A Mechanistic Analysis of Looped Reasoning Language Models

循环推理语言模型的机理分析

Hugh Blayney, Álvaro Arroyo, Johan Obando-Ceron, Pablo Samuel Castro, Aaron Courville, Michael M. Bronstein, Xiaowen Dong

机构 * University of Oxford(牛津大学) Mila – Quebec AI Institute(魁北克AI研究所)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 本文分析了循环推理语言模型中潜在状态的机理,揭示了循环块在潜在空间中的稳定轨迹及注意力头行为的稳定特性。

Comments 39 pages, 63 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10900 2026-04-14 cs.AI cs.LG 81%

CASK: Core-Aware Selective KV Compression for Reasoning Traces

CASK:面向推理轨迹的核心感知选择性KV压缩

Buseong Kim, Heejun Gwon

机构 * d’strict Korea(d’strict 韩国)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 CASK通过核心保护与选择性擦除策略优化推理轨迹的KV缓存,提升内存效率和推理稳定性,在AIME24和AIME25测试中优于TriAttention。

Comments 25 pages, 8 figures, 3 main tables, appendices included

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10455 2026-04-14 cs.CL 79%

EviCare: Enhancing Diagnosis Prediction with Deep Model-Guided Evidence for In-Context Reasoning

EviCare: 通过深度模型引导的证据增强诊断预测

Hengyu Zhang, Xuyun Zhang, Pengxiang Zhan, Linhao Luo, Hang Lv, Yanchao Tan, Shirui Pan, Carl Yang

机构 * Macquarie University(麦考瑞大学) Fuzhou University(福州大学) Monash University(莫纳什大学) Griffith University(格里菲斯大学) Emory University(埃默里大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL

AI总结 EviCare通过整合深度模型指导的证据选择、证据优先级排序和关系证据构建,提升基于LLM的诊断预测性能,在MIMIC-III和MIMIC-IV数据集上实现20.65%的平均提升,尤其在新诊断预测上提升显著。

Comments Accepted by KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04672 2026-04-14 cs.CV cs.CL 79%

Agri-R1: Agricultural Reasoning for Disease Diagnosis via Automated-Synthesis and Reinforcement Learning

Agri-R1:通过自动合成与强化学习进行农业疾病诊断的推理增强

Wentao Zhang, Mingkun Xu, Qi Zhang, Shangyang Li, Derek F. Wong, Lifei Wang, Yanchao Yang, Lina Lu, Tao Fang

机构 * Shandong University of Technology(山东理工大学) Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) Faculty of Data Science, City University of Macau(澳门城市大学数据科学学院) School of Physical Science and Technology, Beijing University of Posts and Telecommunications(北京邮电大学物理科学与技术学院) NLP2CT Lab, Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系自然语言处理与中葡机器翻译实验室) Institute of International Language Services Studies, Macau Millennium College(澳门千禧学院国际语言服务研究所)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL

AI总结 Agri-R1通过自动合成与强化学习提升农业疾病诊断,利用仅19%的数据生成高质量推理数据,采用改进的奖励函数提升模型在疾病识别和农业知识问答上的性能。

Comments This paper is submitted for review to the 2026 ACM MM Conference. The corresponding authors are Tao Fang and Lina Lu, where Tao Fang is the senior Corresponding Author (Last Author) and the principal supervisor of this work, having led the research design, guided the methodology, and overseen the entire project

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09917 2026-04-14 cs.MA cs.GT 78%

Toward Explanatory Equilibrium: Verifiable Reasoning as a Coordination Mechanism under Asymmetric Information

迈向解释均衡:在信息不对称下可验证推理作为协调机制

Feliks Bańka, Jarosław A. Chudziak

专题命中 其他推理 :reasoning(title,abstract)

AI总结 本文提出解释均衡概念,研究在信息不对称环境下,通过结构化推理 artifacts 和有限验证机制实现安全协调,证明结构化推理能有效降低不良批准率。

Comments 18 pages, 4 figures. Accepted for presentation at EXTRAAMAS 2026 (AAMAS 2026 workshop); to appear in post-proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10701 2026-04-14 cs.LG cs.AI cs.CL 75%

Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning

带回价值模型:生成式批评者在大语言模型强化学习中的价值建模

Zikang Shan, Han Zhong, Liwei Wang, Li Zhao

机构 * Peking University(北京大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出生成式批评者GenAC,通过链式推理改进价值建模,提升RL性能。

Comments 16 pages including appendix, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10539 2026-04-14 cs.LG cs.AI 73%

IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs

IceCache: 用于长序列LLM的内存高效KV缓存管理

Yuzhen Mao, Qitong Wang, Martin Ester, Ke Li

机构 * Simon Fraser University(西蒙菲莎大学) Harvard University(哈佛大学)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 本文提出IceCache,通过语义令牌聚类与PagedAttention结合,提升长序列LLM的内存效率与性能,实验表明其在256个令牌预算下保持99%的准确性,且在延迟和精度上优于其他方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10161 2026-04-14 cs.SD 67%

From Speech to Profile: A Protocol-Driven LLM Agent for Psychological Profile Generation

从语音到档案:一种基于协议的LLM代理用于心理档案生成

Xingjian Yang, Yudong Yang, Zhixing Guo, Yongjie Zhou, Nan Yan, Lan Wang

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) Key Laboratory of Biomedical Imaging Science and System, Chinese Academy of Sciences(中国科学院生物医学成像科学与系统重点实验室)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract)

AI总结 本文提出StreamProfile框架,通过增量处理咨询语音,提取证据并生成可追溯的心理档案,有效避免幻觉和长上下文遗忘问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00487 2026-04-14 cs.MA cs.GT cs.SY eess.SY 67%

Competition and Cooperation of LLM Agents in Games

博弈中的大语言模型代理竞争与合作

Jiayi Yao, Cong Chen, Baosen Zhang

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract)

AI总结 研究大语言模型代理在资源分配和Cournot竞争博弈中的行为,发现多轮提示和非零和上下文促进合作,基于公平性推理提出分析框架。

Comments Submitted to CDC'2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15593 2026-04-14 cs.CL cs.AI cs.LG 67%

Parallelism and Generation Order in Masked Diffusion Language Models: Limits Today, Potential Tomorrow

掩码扩散语言模型中的并行性与生成顺序:今日的限制,明日的潜力

Yangyang Zhong, Yanmei Gu, Zhengqing Zang, Xiaomeng Li, Yuqi Ding, Xibei Jia, Yuting Shen, Zhenzhong Lan, Liwang Zhu, Weiping Liu, Junlin Zhou, Haisheng Liu, Zhong Xin Yu, Pengxin Luo, Donglian Qi, Yunfeng Yan, Junbo Zhao

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团) Shanghai Jiao Tong University(上海交通大学) University of Chinese Academy of Social Sciences(中国社会科学院大学) Westlake University(西湖大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究评估了八种主流MDLM在58个基准上的表现,发现其并行性和生成顺序受任务领域和正确性影响显著,且在需要逆向信息的任务中表现更优,提出生成后编辑范式以提升效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11805 2026-04-14 cs.LG cs.AI cs.CV cs.RO 62%

Solving Physics Olympiad via Reinforcement Learning on Physics Simulators

通过物理模拟器强化学习解决物理竞赛

Mihir Prabhudesai, Aryan Satpathy, Yangmin Li, Zheyang Qin, Nikash Bhardwaj, Amir Zadeh, Chuan Li, Katerina Fragkiadaki, Deepak Pathak

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文通过物理模拟器生成合成数据,利用强化学习训练LLM,实现零样本迁移到现实物理竞赛,提升模型物理推理能力。

Comments Project Webpage - https://sim2reason.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03402 2026-04-14 cs.AI cs.LG 62%

Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility

风险感知注入:为安全校准视觉-语言模型而不牺牲实用性

Mengxuan Wang, Yuxin Chen, Gang Xu, Tao He, Hongjie Jiang, Ming Li

机构 * Shien-Ming Wu School of Intelligent Engineering, South China University of Technology(华南理工大学吴贤铭智能工程学院) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室(深圳)) Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系) University of Electronic Science and Technology of China(电子科技大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出RAI框架,通过增强不安全信号恢复视觉-语言模型的安全识别能力,同时保持语义完整性,实验显示有效降低攻击成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11217 2026-04-14 cs.CL cs.AI 62%

Domain-Specific Data Generation Framework for RAG Adaptation

面向RAG适应的领域特定数据生成框架

Chris Xing Tian, Weihao Xie, Zhen Chen, Zhengyuan Yi, Hui Liu, Haoliang Li, Shiqi Wang, Siwei Ma

机构 * Peng Cheng Laboratory(鹏城实验室) City University of Hong Kong(香港城市大学) Peking University(北京大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出RAGen框架,通过识别文档中的关键概念生成领域相关的问答对,支持多种RAG适应策略,提升领域适应效果。

Comments To appear in ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10985 2026-04-14 cs.AI cs.CL cs.CV 62%

Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models

回到牛棚与LLAMAs:在微调视觉语言模型中进化预训练LLM骨干

Sameera Horawalavithana, Lauren Phillips, Ian Stewart, Sai Munikoti, Karl Pazdernik

机构 * Pacific Northwest National Laboratory(太平洋西北国家实验室)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨了在微调视觉语言模型时,不同预训练LLM骨干对下游任务性能的影响,发现新版本LLM并非总能提升性能,且表现依赖具体任务。

Comments Preprint and under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10800 2026-04-14 cs.SE cs.AI cs.CR cs.LG cs.PL 62%

Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis

在修复前验证:面向可信跨语言代码分析的代理执行 grounding

Jugal Gajjar

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出统一的跨语言漏洞生命周期框架,通过LLM驱动的三个阶段:混合结构-语义检测、执行 grounding 的代理验证和验证-aware 的迭代修复,确保修复前有执行验证。框架利用uAST和图神经网络融合,实现高准确率的漏洞检测与跨语言修复。

Comments 20 pages (13 main + 7 appendices), 9 figures, 10 tables. Submitted to NeurIPS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏