arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 1309 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 1309 篇

2605.28346 2026-07-24 cs.CL 版本更新 79%

When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs

当话语压力冲突时:视觉-语言模型输出中的信息结构

Marcell Fekete, Johannes Bjerva, Tamás Káldi

机构 * Department of Computer Science, Aalborg University(奥尔堡大学计算机科学系) Department of Psycholinguistics and Neurolinguistics, ELTE Research Centre for Linguistics(ELTE语言研究中心心理学语言学与神经语言学系) ELTE Bárczi Gusztáv Faculty of Special Needs Education(ELTE巴尔茨吉斯塔夫特殊教育学院)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

AI总结 研究视觉-语言模型在视觉问答中是否区分话语旧主题和新焦点,发现模型虽产生信息结构相关结构但过度正则化,倾向于窄响应模板,类似模式崩溃。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14480 2026-07-22 cs.CL 版本更新 79%

Lower-Resource, Higher Scores: Language Bias in LLM Evaluators

语言模型评估器在不同语言间存在偏差

Ej Zhou, Lucas Resck, Zheng Hui, Anna Korhonen

机构 * University of Cambridge(剑桥大学) Language Technology Lab(语言技术实验室)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.CL

AI总结 研究发现语言模型评估器在多语言环境中存在偏差,不同语言评分差异显著,与语言资源水平相关,成对准确率无法检测到这些偏差,还探究了资源少的语言得分高的原因,揭示了语言层面的结构错位。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12985 2026-07-22 cs.AI 版本更新 79%

SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents

SENTINEL:面向基础模型基于具身代理的安全性评估多级形式框架

Simon Sinong Zhan, Philip Wang, Yao Liu, Yiyan Peng, Zinan Wang, Qineng Wang, Zhian Ruan, Xiangyu Shi, Xinyu Cao, Frank Yang, Zhenyang Ni, Kangrui Wang, Ruohan Zhang, Huajie Shao, Manling Li, Qi Zhu

机构 * Northwestern University, USA University of Southern California, USA College of William \& Mary

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.AI

AI总结 SENTINEL通过形式时序逻辑多级验证框架,系统性评估基础模型具身代理在模拟环境中的物理安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22259 2026-07-21 cs.LG 版本更新 79%

Tabular Foundation Models Can Do Survival Analysis

表格基础模型也能进行生存分析

Da In Kim, Wei Siang Lai, Kelly W. Zhang

机构 * Department of Computing, Imperial College London(帝国理工学院计算机系) Department of Mathematics, Imperial College London(帝国理工学院数学系)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.LG

AI总结 本文提出了一种基于分类的框架,通过将生存分析转化为二分类问题,使表格基础模型能够无需显式训练即可进行生存分析,并在多个数据集上验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18613 2026-07-21 q-bio.NC cs.CL cs.CV 版本更新 79%

The Illusion-Illusion: Vision Language Models See Illusions Where There Are None

错觉-错觉:视觉语言模型在不存在错觉的地方看到错觉

Tomer Ullman

机构 * Harvard University(哈佛大学)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

AI总结 研究通过给视觉语言模型呈现不应引发处理错误的‘错觉-错觉’,发现许多模型会误将其视为错觉,揭示了模型存在基本处理错误,此失败是文献中更广泛失败的一部分。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10310 2026-07-20 cs.CL 版本更新 79%

PolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment

PolyInterview:一个基于大语言模型的沉浸式模拟面试平台,具备全面的多模态评估

Zhiyuan Wen, Jiannong Cao, Kelly Chan, Zijian Wang, Chen Chen, Xiaoyun Liu, Jianing Yin, Zhuo Li

机构 * The Hong Kong Polytechnic University(香港理工大学) Chongqing University of Posts and Telecommunications(重庆邮电大学)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.CL

AI总结 PolyInterview平台利用大语言模型,基于职位描述和简历为求职者生成定制面试问题,通过数字人类面试官进行多轮口语面试,全面评估回答内容、语音表达和非语言行为,提供结构化反馈,助力求职者更好地准备面试。

Comments 10 pages, 7 figures, and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05268 2026-07-17 cs.CV cs.LG 版本更新 79%

Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models

几何在发挥作用吗?对双曲视觉-语言模型层次性的工作点审计

Jaeyoung Kim, Eunseok Kim, Dongsuk Jang

机构 * MADI

专题命中 评测与基准 :language model(title,abstract);分类 cs.LG

AI总结 本研究针对双曲视觉-语言模型是否利用其几何特性的问题,提出多组诊断指标审计三类主流模型,发现其实际未激活双曲径向/锥机制,层次性表现与几何特性无关。

Comments 48 pages, 5 figures, Under review at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29759 2026-07-17 cs.CV cs.AI 版本更新 79%

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios

TSHA:用于可信安全危害评估场景的视觉语言模型基准

Qiucheng Yu, Ruijie Xu, Mingang Chen, Jianfeng Dong, Xin Tan

机构 * City University of Hong Kong(香港城市大学) East China Normal University(华东师范大学) University of Western Australia(西澳大学) Shanghai AI Laboratory(上海人工智能实验室) Shanghai Development Center of Computer Software Technology(上海计算机软件技术开发中心) Zhejiang Gongshang University(浙江工商大学)

专题命中 评测与基准 :language model(title,abstract);分类 cs.AI

AI总结 本文提出TSHA基准,通过81809个精心挑选的训练样本和1707个挑战性测试样本,评估视觉语言模型在复杂家庭安全场景中的鲁棒性和泛化能力,发现现有模型在安全危害评估中存在显著不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15892 2026-07-15 cs.CV cs.AI 版本更新 79%

Egocentric Bias in Vision-Language Models

视觉语言模型中的自我中心偏差

Maijunxian Wang, Yijiang Li, Bingyang Wang, Tianwei Zhao, Ran Ji, Qingying Gao, Emmy Liu, Hokin Deng, Dezhi Luo

机构 * Cognitive Science Program, University of California, Berkeley(加州大学伯克利分校认知科学项目) Department of Electrical and Computer Engineering, University of California San Diego(加州大学圣地亚哥分校电气与计算机工程系) School of Computer Science, Georgia Institute of Technology & Emory University(佐治亚理工学院计算机科学学院及埃默里大学) Department of Computer Science, Johns Hopkins University(约翰霍普金斯大学计算机科学系) Department of Cognitive Science, University of California San Diego(加州大学圣地亚哥分校认知科学系) Equal Advising Department of Computer Science & Wilmer Eye Institute, Johns Hopkins University(约翰霍普金斯大学计算机科学系及威尔默眼科研究所) Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所) Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) Weinberg Institute for Cognitive Science, University of Michigan(密歇根大学韦恩伯格认知科学研究所)

专题命中 评测与基准 :language model(title,abstract);分类 cs.AI

AI总结 本文提出FlipSet基准,揭示视觉语言模型在第二级视觉视角推理中存在系统性自我中心偏差,表明模型在整合社会认知与空间操作方面存在根本性不足。

Comments Accepted at CogSci 2026 (Best Undergraduate Student Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07405 2026-07-14 cs.AI cs.CR 版本更新 79%

Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents

少些推理,多些验证:确定性门控在使用工具的语言模型智能体中恢复了一种无声的违反策略失败模式

Vikas Reddy, Sumanth Reddy Challaram, Abhishek Basu

机构 * Indian Institute of Technology Kharagpur(印度理工学院卡拉格布尔分校) Massachusetts Institute of Technology(麻省理工学院)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.AI

AI总结 研究使用工具的语言模型智能体违反策略问题,提出用确定性预执行门控干预,在τ²基准航空公司领域评估,该方法能提高成功率,防止无声违反策略写入,虽不保证任务成功,但有可靠性成果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12612 2026-07-10 cs.IR cs.AI 版本更新 79%

Self-EvolveRec: Self-Evolving Recommender Systems with LLM-based Directional Feedback

Self-EvolveRec:基于大语言模型定向反馈的自进化推荐系统

Sein Kim, Sangwu Park, Hongseok Kang, Wonjoong Kim, Jimin Seo, Yeonjun In, Kanghoon Yoon, Hyunsik Jeon, Chanyoung Park

机构 * Microsoft(微软)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.AI

AI总结 研究针对传统推荐系统设计方法局限,提出Self-EvolveRec框架,集成用户模拟器与模型诊断工具建立定向反馈循环,并引入协同进化策略。实验证明该框架在推荐性能和用户满意度上显著优于现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17248 2026-07-07 eess.AS cs.CL cs.SD 版本更新 79%

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

VIBE:通过真实世界语音进行大型音频-语言模型生成偏见评估的语音诱导开放式偏见评估

Yi-Cheng Lin, Yusuke Hirota, Sung-Feng Huang, Hung-yi Lee

机构 * Graduate Institute of Communication Engineering, National Taiwan University, Taiwan(台湾大学通信工程研究所) NVIDIA, Taiwan(台湾NVIDIA) Artificial Intelligence Center of Research Excellence, National Taiwan University, Taiwan(台湾大学人工智能卓越研究中心)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

AI总结 VIBE通过真实世界语音的开放式任务评估大型音频-语言模型的生成偏见,揭示了性别线索比口音线索更易引发分布偏移,表明当前模型复现了社会刻板印象。

Comments Submitted to SLT 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08467 2026-07-07 cs.CV cs.LG eess.IV 版本更新 79%

Zero-Shot Distracted Driver Detection via Vision Language Models with Double Decoupling

基于双解耦视觉语言模型的零样本分心驾驶员检测

Takamichi Miyata, Sumiko Miyata, Andrew Morris

机构 * Chiba Institute of Technology(千叶工业大学) Institute of Science Tokyo(东京科学大学) Loughborough University(拉夫堡大学)

专题命中 评测与基准 :language model(title,abstract);分类 cs.LG

AI总结 提出一种主体解耦框架,通过提取驾驶员外观嵌入并从图像嵌入中移除其影响,再结合度量投影正交化文本嵌入,实现零样本分心驾驶检测,显著提升实际场景性能。

Comments Accepted to IEEE 15th International Symposium on Communication Systems, Networks and Digital Signal Processing (CSNDSP 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06643 2026-07-07 cs.SE cs.CL 版本更新 79%

Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models

你的基准仍然有用吗?代码语言模型的动态基准测试

Batu Guan, Xiao Wu, Yuanyuan Yuan, Shaohua Li

机构 * The Chinese University of Hong Kong(香港中文大学) Huazhong University of Science and Technology(华中科技大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

AI总结 研究模型评估中如何在模型训练可能已见过基准的情况下保持其有用性,提出动态基准测试框架,通过语义保留突变变换输入构建新基准,评估发现模型性能变差、排名变化及能抵抗数据污染问题。

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05150 2026-07-03 cs.CV cs.AI 版本更新 79%

Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment

面向生物标志物评估的病理基础模型中的细胞级可解释性

Jingsong Liu, Han Li, Zhengyang Xu, Franz-Leonard Klaus, Fabian Stögbauer, Shihui Zu, Weiwei Zhou, Atsuko Kasajima, Felix Schicktanz, Alexander Muckenhuber, Julius Shakhtour, Jiale Yu, Tiannan Zheng, Xun Ma, Maggie Wang, Christian Grashei, Bao Li, Guiyang Jiang, Hongming Xu, Shaohua Kevin Zhou, Nassir Navab, Peter J. Schüffler

机构 * Institute of Pathology, Technical University of Munich(慕尼黑技术大学病理学研究所) School of Computation, Information and Technology, Technical University of Munich(慕尼黑技术大学计算、信息与技术学院) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Computer Aided Medical Procedures (CAMP), Technical University of Munich(慕尼黑技术大学计算机辅助医疗程序中心) School of Biomedical Engineering, Faculty of Medicine, Dalian University of Technology(大连理工大学医学院生物医学工程学院) Affiliated Hospital of Chifeng University(赤峰大学附属医院) Center for Medical Imaging, Robotics, and Analytic Computing & Learning (MIRACLE), Suzhou Institute for Advanced Research, USTC, Suzhou, China(苏州先进研究院医学影像、机器人与分析计算与学习中心) Department of Biomedical Informatics, Harvard Medical School(哈佛医学院生物医学信息学系) The First Hospital and the College of Basic Medical Sciences of China Medical University(中国医科大学第一医院及基础医学科学学院) Munich Data Science Institute (MDSI)(慕尼黑数据科学研究所)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.AI

AI总结 提出Hireca病理基础模型和CytoMap可解释性模块,在10项生物标志物任务中多数领先,提供细胞级证据定位,实现透明可审查的生物标志物评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28418 2026-07-02 cs.LG 版本更新 79%

Explaining Tabular Foundation Model Differences Through Meta-Features

重新审视元特征以解释表格数据上的模型差异

Markus Herre, Andrej Tschalzev, Sascha Marton, Christian Bartelt

机构 * Clausthal University of Technology, Clausthal-Zellerfeld, Germany(Clausthal技术大学,Clausthal-Zellerfeld,德国) University of Mannheim, Mannheim, Germany(曼海姆大学,曼海姆,德国)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.LG

AI总结 研究通过严格统计检验和留一法分析,发现数据集元特征无法稳健解释表格数据上不同模型族(如神经网络与树模型、非基础模型与基础模型)之间的性能差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12887 2026-07-02 cs.IR cs.AI 版本更新 79%

EcoGEO: Trajectory-Aware Evidence Ecosystems for Web-Enabled LLM Search Agents

EcoGEO:基于轨迹的证据生态系统用于网络增强的大语言模型搜索代理

Hengwei Ye, Jiasheng Mao, Zhenhan Guan, Zheng Tian

机构 * ShanghaiTech University(上海科技大学)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.AI

AI总结 本文提出EcoGEO,一种基于轨迹的证据生态系统,用于研究网络增强的大语言模型搜索代理在多步骤浏览中的证据获取过程,通过协调导航页面与支持页面提升产品推荐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08237 2026-07-01 cs.MM cs.AI cs.CV cs.SD eess.AS 版本更新 79%

VGGSounder: Audio-Visual Evaluations for Foundation Models

VGGSounder:基础模型的音视频评估

Daniil Zverev, Thaddäus Wiedemer, Ameya Prabhu, Matthias Bethge, Wieland Brendel, A. Sophia Koepke

机构 * Technical University of Munich, MCML(慕尼黑技术大学,MCML) University of Tübingen(图宾根大学) Tübingen AI Center(图宾根人工智能中心) MPI for Intelligent Systems, ELLIS Institute(智能系统Max Planck研究所,ELLIS研究所)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.AI

AI总结 针对VGGSound数据集在音视频基础模型评估中的标签不完整、类别重叠和模态错位等问题,提出重新标注的多标签测试集VGGSounder,并引入模态混淆指标分析模型性能退化。

Comments Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05738 2026-06-25 cs.CL 版本更新 79%

MedLayBench-V: A Large-Scale Benchmark for Expert-Lay Semantic Alignment in Medical Vision Language Models

MedLayBench-V:面向医学视觉语言模型中专家与普通人语义对齐的大规模基准

Han Jang, Junhyeok Lee, Heeseong Eum, Kyu Sung Choi

机构 * Seoul National University(首尔国立大学) Seoul National University College of Medicine(首尔国立大学医学院) Department of Radiology, Seoul National University Hospital(首尔国立大学医院放射科) Healthcare AI Research Institute, Seoul National University Hospital(首尔国立大学医院健康人工智能研究所) The Advanced Imaging and Computational Neuroimaging (AICON) Laboratory(先进影像与计算神经影像实验室)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

AI总结 提出首个大规模多模态基准MedLayBench-V,通过结构化概念基础精炼管道实现专家-普通人语义对齐,用于训练和评估能弥合医患沟通鸿沟的医学视觉语言模型。

Comments Findings of ACL 2026. 9 pages, 5 figures, 11 tables, plus appendix

Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pages 18375-18394

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04584 2026-06-25 cs.CL cs.SD eess.AS 版本更新 79%

Robustness assessment of large audio language models in multiple-choice evaluation

大型音频语言模型在多项选择评估中的鲁棒性评估

Fernando López, Santosh Kesiraju, Jordi Luque

机构 * Scientific Research, Telefónica Innovación Digital, Spain(Telefónica Innovación Digital科研部,西班牙) Universidad Autónoma de Madrid, Spain(马德里自治大学,西班牙) Brno University of Technology, Czech Republic(布拉格技术大学,捷克)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

AI总结 研究大型音频语言模型在多项选择问答中对选项顺序、问题措辞的敏感性,并提出考虑细微变化的评估协议和指标。

Comments Accepted in Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17742 2026-06-16 eess.SP cs.AI cs.HC 版本更新 79%

EEG-FM-Bench: A Comprehensive Benchmark for the Systematic Evaluation and Diagnostic Analyses of EEG Foundation Models

EEG-FM-Bench:脑电图基础模型系统评估与诊断分析的综合基准

Wei Xiong, Jiangtong Li, Jie Li, Kun Zhu, Changjun Jiang

机构 * School of Computer Science and Technology, Tongji University, Shanghai, China(同济大学计算机科学与技术学院,上海,中国) Translational Research Center, Shanghai Yangzhi Rehabilitation Hospital (Shanghai Sunshine Rehabilitation Center), China(上海杨氏康复医院(上海阳光康复中心)转化研究中心,中国)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.AI

AI总结 提出EEG-FM-Bench统一基准,整合14个数据集和10种范式,通过多种微调策略和诊断分析揭示多任务学习可缓解过拟合、预训练效率受梯度冲突限制、模型规模非唯一决定因素等关键发现。

Comments 36 pages, 30 figures, Accepted by ICML2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17588 2026-06-16 cs.CV cs.CL 版本更新 79%

Dual-branch Prompting for Multimodal Machine Translation

双分支提示用于多模态机器翻译

Jie Wang, Zhendong Yang, Liansong Zong, Xiaobo Zhang, Dexian Wang, Ji Zhang

机构 * School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机与人工智能学院) School of Computer and Software Engineering, Xihua University(西华大学计算机与软件工程学院) School of Intelligent Medicine, Chengdu University of Traditional Chinese Medicine(成都中医药大学针灸推拿学院)

专题命中 评测与基准 :prompting(title,abstract);分类 cs.CL

AI总结 提出基于扩散模型的双分支提示框架D2P-MMT,利用重建图像过滤视觉噪声,通过分布对齐损失提升鲁棒翻译性能。

Comments This manuscript has been fully accepted and published by ACM Transactions on Multimedia Computing, Communications, and Applications (ACM TOMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21071 2026-08-17 cs.CL cs.AI 版本更新 79%

Fine-grained Claim-level RAG Benchmark for Law

细粒度声明级法律RAG基准

Souvick Das, Sallam Abualhaija, Domenico Bianculli

机构 * University of Luxembourg(卢森堡大学)

专题命中 评测与基准 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出一个支持法语和英语的细粒度声明级法律RAG数据集ClaimRAG-LAW,用于评估检索和生成性能,揭示法律领域RAG系统的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11616 2026-08-14 cs.AI cs.CV cs.LG 版本更新 79%

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

MBA:面向现实世界商业创意的多模态基准与智能体

Hojun Choi, Jaeyo Shin, Suin Lee, Hyunjung Shim

专题命中 评测与基准 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 该研究推出首个多模态商业创意基准MBA-Bench,提出MBA-b和MBA-k两种智能体,经实验其性能显著优于相关基准,为多模态商业创意智能体研究提供了重要支撑。

Comments Project page: https://hchoi256.github.io/projects/mba/ Code: https://github.com/hchoi256/MBA

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20295 2026-08-10 cs.LG cs.AI cs.CV 版本更新 79%

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription

通过封面判断书籍:调查多模态大语言模型用于多页手写文档转录

Benjamin Gutteridge, Matthew Thomas Jackson, Toni Kukurin, Xiaowen Dong

机构 * University of Oxford(牛津大学) QuantCo

专题命中 评测与基准 :LLM(abstract,abstract_cn);prompting(abstract);分类 cs.AI、cs.LG

AI总结 本文研究多模态大语言模型在多页手写文档转录中的应用,提出OCR+PAGE-1和OCR+PAGE-N策略,通过共享页面内容提升转录效果。

Comments 10 pages (36 including references and appendices), 11 figures, accepted at COLM 2026, earlier version accepted at AAAI 2025 Workshop on Document Understanding and Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14695 2026-08-07 cs.LG cs.CL 版本更新 79%

Persona-Pruner: Sculpting Lightweight Models for Role-Playing

Persona-Pruner: 为角色扮演雕琢轻量级模型

Jinsu Kim, Jihoon Tack, Noah Lee, Jongheon Jeong

机构 * Department of Artificial Intelligence, Korea University, Seoul, South Korea(韩国大学人工智能系) Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 评测与基准 :LLM(abstract,abstract_cn);language model(abstract);分类 cs.CL、cs.LG

AI总结 提出Persona-Pruner框架,通过从单个描述中隔离特定角色的子网络来剪枝语言模型,在保持角色扮演性能的同时大幅降低计算成本,性能下降比最强基线减少93.8%。

Comments 25 pages; ICML 2026; Code is available at https://github.com/jsu-kim/Persona-Pruner

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21692 2026-08-06 cs.CL cs.AI 版本更新 79%

Revisiting Generalization Across Difficulty Levels: It's Not So Easy

重新审视不同难度层级间的泛化:这并不容易

Yeganeh Kordi, Nihal V. Nayak, Max Zuo, Ilana Nguyen, Stephen H. Bach

机构 * Brown University(布朗大学) Harvard University(哈佛大学)

专题命中 评测与基准 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了LLMs在不同任务难度间泛化的能力,发现训练数据的难度对泛化效果影响有限,强调在训练和评估中需涵盖多种难度以避免风险。

Comments Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19827 2026-08-03 cs.CL cs.AI cs.IR 版本更新 79%

When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering

当迭代RAG优于理想证据:科学多跳问答中的诊断研究

Mahdi Astaraki, Mohammad Arshi Saloot, Ali Shiraee Kasmaee, Hamidreza Mahyar, Soheila Samiee

机构 * Faculty of Engineering, McMaster University, Canada(麦斯特大学工程学院,加拿大) BASF Canada Inc., Canada(巴斯夫加拿大公司,加拿大)

专题命中 评测与基准 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 通过化学多跳问答数据集,诊断发现迭代检索-推理循环在科学领域显著优于静态RAG上限,揭示了阶段式检索的优势与失败模式。

Comments 51 pages, 29 figures, Published in Transactions on Machine Learning Research (05/2026). OpenReview: https://openreview.net/forum?id=pa5TnBdyDP

Journal ref Transactions on Machine Learning Research (05/2026), ISSN 2835-8856

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01153 2026-07-30 cs.CL cs.AI cs.SE 版本更新 79%

Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control

面向AI安全评估的对抗语用学:指令冲突、嵌入命令与策略模糊性基准

Brett Reynolds

机构 * Humber Polytechnic(汉博理工学院) University of Toronto(多伦多大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出对抗语用学基准和标注协议,通过语言学控制的分类法评估模型在指令冲突、嵌入命令等场景下的行为,为安全评估提供实证和方法论工具。

Comments 32-page main paper plus 13-page supplement; 6 figures and 17 tables total; code and data artifact available at the linked repository

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05710 2026-07-30 physics.ao-ph cs.AI cs.LG 版本更新 79%

The Rise of AI in Weather and Climate Information and its Impact on Global Inequality

人工智能在天气和气候信息中的兴起及其对全球不平等的影响

Amirpasha Mozaffari, Amanda Duarte, Lina Teckentrup, Stefano Materia, Gina E. C. Charnley, Lluis Palma, Eulalia Baulenas Serra, Dragana Bojovic, Paula Checchia, Aude Carreric, Francisco Doblas-Reyes

机构 * Catalan Institution for Research and Advanced Studies (ICREA)(加泰罗尼亚研究与高级研究机构)

专题命中 评测与基准 :large language model(abstract);language model(abstract);foundation model(abstract);分类 cs.AI、cs.LG

AI总结 人工智能在天气和气候信息中的兴起加剧了全球不平等,需通过数据为中心的发展、气候数字基础设施和知识共生产来解决不平等问题。

详情

展开后加载摘要…

URL PDF HTML 收藏