arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 1309 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 1309 篇

2602.20064 2026-07-13 cs.PL cs.AI cs.CR 版本更新 81%

The LLMbda Calculus: AI Agents, Conversations, and Information Flow

LLMlambda 计算:人工智能代理、对话与信息流

Zac Garby, Andrew D. Gordon, David Sands

机构 * University of Nottingham, UK(诺丁汉大学) University of Edinburgh, UK(爱丁堡大学) Chalmers University of Technology(查尔姆斯理工大学) The University of Gothenburg, Sweden(哥德堡大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出了一种无类型的lambda计算模型,用于形式化描述和验证人工智能代理中对话和信息流的安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12857 2026-07-10 cs.CY cs.AI 版本更新 81%

Adaptive Generation of Bias-Eliciting Questions for LLMs

为大语言模型自适应生成偏差引发问题

Robin Staab, Jasper Dekoninck, Maximilian Baader, Martin Vechev

机构 * ETH Zurich, Switzerland(苏黎世联邦理工学院)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 针对大语言模型固有偏差问题,引入反事实框架,通过迭代问题变异生成开放式问题评估偏差,构建CAB基准,评估发现模型在某些场景有持续偏差,凸显公平性研究需求。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27395 2026-07-09 cs.CY cs.AI 版本更新 81%

Informing AI Policy Assessment using Large-Scale Simulation of Interventions

利用大规模干预模拟为AI政策评估提供信息

Julia Barnett, Kimon Kieslich, Natali Helberger, Nicholas Diakopoulos

机构 * Northwestern University USA University of Amsterdam, The Netherlands \& University of Hohenheim Germany University of Amsterdam The Netherlands Northwestern University University of Amsterdam, The Netherlands \& University of Hohenheim University of Amsterdam

专题命中 评测与基准 :LLM(summary_cn,abstract);分类 cs.AI

AI总结 提出一种结合参与式评估、专家成本评估和基于LLM的伤害缓解评估的方法,通过遗传算法模拟探索政策组合空间,以识别缓解特定AI危害的可行政策选项。

Comments This work is published in the proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT) 2026. 15 pages plus end matter and appendix

Journal ref In Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26). Association for Computing Machinery, New York, NY, USA, 3942-3978

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.03202 2026-07-07 cs.AI 版本更新 81%

Stop Automating Peer Review Without Rigorous Evaluation

停止在未经严格评估的情况下自动化同行评审

Joachim Baumann, Jiaxin Pei, Sanmi Koyejo, Dirk Hovy

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 该立场文通过对比人类与AI生成的ICLR 2026评审结果,指出当前AI评审存在视角趋同、分数易被论文改写操控的问题,主张需建立同行评审自动化科学,而非直接部署未严格评估的通用大模型。

Comments Accepted at ICML 2026 (Oral). Forty-third International Conference on Machine Learning Position Paper Track (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30133 2026-07-07 cs.CL 版本更新 81%

CorPipe at CRAC 2026: Empty Nodes and Cross-Lingual Transfer in Multilingual Coreference Resolution

CorPipe at CRAC 2026: 多语言共指消解中的空节点与跨语言迁移

Milan Straka

机构 * Charles University, Faculty of Mathematics and Physics Institute of Formal and Applied Linguistics(查理大学数学与物理系形式与应用语言学研究所)

专题命中 评测与基准 :LLM(summary_cn,abstract);分类 cs.CL

AI总结 本文提出CorPipe 26系统,通过单一模型联合预测空节点、提及和共指链接,在CRAC 2026多语言共指消解共享任务中超越所有其他系统,并在LLM赛道和不受限赛道分别领先2.8和9.5个百分点。

Comments Accepted to CODI-CRAC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30441 2026-07-03 cs.MA cs.AI 版本更新 81%

Translating Natural Language to Strategic Temporal Specifications via LLMs

通过大语言模型将自然语言翻译为策略时序规范

Marco Aruta, Francesco Improta, Vadim Malvone, Aniello Murano, Vladana Perlić

机构 * University of Naples Federico II(那不勒斯费德里科二世大学) LTCI, Télécom Paris, Institut Polytechnique de Paris(LTCI,巴黎电信,巴黎理工学院)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出利用大语言模型将自然语言描述的策略需求翻译为ATL/ATL*公式,构建专家验证数据集,通过领域内微调小模型达到与强少样本专有API基线相当的语义准确性(0.84 vs 0.86),并集成到模型检查器中支持非专家用户。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22487 2026-07-03 cs.CL 版本更新 81%

Polite on the Surface, Broken in Practice: A Curated Dataset for Fixing Generation and Register Failures in Low-Resource Bangla Text Generation

表面礼貌,实践错误:一个用于修复多语言孟加拉语生成中敬语失误的定制数据集

Md. Asaduzzaman Shuvo, Mahedi Hasan, Md. Tashin Parvez, Azizul Haque Noman, Md. Shafayet Hossain Ovi

机构 * United International University(国际大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出了一个定制数据集BLADE,用于改进多语言孟加拉语生成中敬语处理的准确性,通过系统微调和评估领先的开源架构,如DeepSeek-8B和LLaMA-3.2-3B,以提高结构忠实度和敬语对齐度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26144 2026-06-24 cs.SE cs.AI cs.CV 版本更新 81%

VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents

VISTA:面向视觉规格到网页应用编码智能体的端到端基准

JunJia Guo, Yuhang Yao, Jiawei, Zhou, Jingdi Chen

机构 * University of Arizona(亚利桑那大学) Zoom Stony Brook University(石溪大学)

专题命中 评测与基准 :LLM(summary_cn,abstract);分类 cs.AI

AI总结 提出VISTA基准,通过多维度输入条件和评估指标,衡量基于LLM的智能体从视觉规格生成功能完整、视觉一致的网页应用的能力。

Comments Project page: https://kaboider.github.io/VIS_APP/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13836 2026-06-18 cs.CL cs.CV cs.MM 版本更新 81%

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs

FutureOmni:从全模态上下文中评估多模态大语言模型的未来预测能力

Qian Chen, Jinlan Fu, Changsong Li, Min Zhang, See-Kiong Ng, Xipeng Qiu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳分校) National University of Singapore(新加坡国立大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 提出FutureOmni基准,评估多模态大模型从音视频线索预测未来的能力,发现现有模型在语音密集场景下表现差,并设计OFF训练策略提升性能。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21570 2026-06-12 cs.AI cs.RO 版本更新 81%

From Digital to Physical: Digital Agents as Autonomous Coaches for Physical Intelligence

从数字到物理:数字代理作为物理智能的自主教练

Zixing Lei, Genjia Liu, Yuanshuo Zhang, Qipeng Liu, Yuzhu Cai, Sixiang Chen, Jixian Wu, Yunhong Wang, Weixin Li, Chuan Wen, Bo Zhao, Shanghang Zhang, Wenzhao Lian, Siheng Chen

机构 * School of Artificial Intelligence, Shanghai Jiao Tong University, Shanghai, China(上海交通大学人工智能学院) Zhongguancun Academy, Beijing, China(中关村学院) School of Integrated Circuits, Shanghai Jiao Tong University, Shanghai, China(上海交通大学集成电路学院) School of Computer Science, Shanghai Jiao Tong University, Shanghai, China(上海交通大学计算机科学学院) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China(北京大学计算机科学学院多媒体信息处理国家重点实验室)

专题命中 评测与基准 :LLM(summary_cn,abstract);分类 cs.AI

AI总结 提出EmboCoach-Bench基准,评估LLM代理自主设计具身策略的能力,通过迭代调试和优化,代理在平均成功率上超越人工基线26.5%,并具备自我修正能力。

Comments 53 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23823 2026-06-12 cs.CL 版本更新 81%

RAGPPI: RAG Benchmark for Protein-Protein Interactions in Drug Discovery

RAGPPI:药物发现中蛋白质-蛋白质相互作用的RAG基准

Youngseung Jeon, Ziwen Li, Thomas Li, JiaSyuan Chang, Morteza Ziyadi, Xiang 'Anthony' Chen

机构 * University of California Los Angeles(加州大学洛杉矶分校) Palo Alto High School(帕洛阿尔托高中) Amazon AGI(亚马逊人工智能研究院)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 提出RAGPPI基准,包含4420个问答对,用于评估检索增强生成在药物发现中识别蛋白质-蛋白质相互作用生物学影响的能力。

Comments 17 pages, 4 figures, 8 tables

Journal ref Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05463 2026-06-10 cs.AI 版本更新 81%

PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage

PSEBench: 一个用于评估大语言模型在患者安全事件分类中的可控且可验证的基准

Keqi Han, Ryan Young, Annabel Strauss, Lindsey Hughes, Katharine M. Nesbitt, Nicole Schueler, Che Ngufor, Carl Yang, Yuan Xue, Zhijun Yin

机构 * Emory University(埃默里大学) Scale AI Mayo Clinic(梅奥诊所) Vanderbilt University Medical Center(范德比大学医学中心)

专题命中 评测与基准 :LLM(summary_cn,abstract);分类 cs.AI

AI总结 提出基于政策条款卡的结构化构建方法,通过锚点驱动实例化和闭环验证生成带真实标签的叙事,并创建包含5074个案例的基准PSEBench,评估15个代表性LLM在患者安全事件分类中的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07061 2026-06-10 cs.CL 版本更新 81%

Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages

重新审视印度语言机器翻译和摘要细粒度评估的度量可靠性

Amir Hossein Yari, Kalmit Kulkarni, Ahmad Raza Khan, Fajri Koto

机构 * Sharif University of Technology(谢里夫理工学院) Vellore Institute of Technology(韦洛雷理工学院) IIT Kharagpur(印度理工学院达卡分校) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 评测与基准 :LLM(summary_cn,abstract);分类 cs.CL

AI总结 针对印度语言评估不足的问题,提出ITEM基准,系统评估29种自动度量与人工判断的对齐,发现基于LLM的评估器表现最佳,并揭示了异常值影响、任务差异及扰动鲁棒性等关键发现。

Comments 18 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08304 2026-06-09 cs.CR cs.AI 版本更新 81%

Securing Retrieval-Augmented Generation: A Taxonomy of Attacks, Defenses, and Future Directions

保障检索增强生成:攻击、防御与未来方向的分类法

Yuming Xu, Mingtao Zhang, Zhuohan Ge, Haoyang Li, Nicole Hu, Yongqi Zhang, Zhiyuan Wen, Jason Chen Zhang, Qing Li, Lei Chen

机构 * The Hong Kong Polytechnic University(香港理工大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出SLOT分类法,从攻击面、防御层、目标(遵循CIA属性)和攻击目标四个维度系统化梳理检索增强生成(RAG)的安全风险与防御,并指出知识访问管道中的结构性错配,最后展望未来方向。

Comments We have curated a paper list on RAG security in https://github.com/TreeAI-Lab/Awesome-RAG-Security, and we warmly welcome authors who wish to have their new work included to contact us via email

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08724 2026-06-09 cs.LG 版本更新 81%

Exposing Hidden Biases in Text-to-Image Models via Automated Prompt Search

通过自动化提示搜索暴露文本到图像模型中的隐藏偏见

Manos Plitsis, Giorgos Bouritsas, Vassilis Katsouros, Yannis Panagakis

机构 * University of Edinburgh(爱丁堡大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出Bias-Guided Prompt Search框架,通过自动生成提示最大化图像偏见,揭示文本到图像模型中的隐藏偏见,提升公平性评估。

Comments ICML 2026. Code is here: https://github.com/manosplitsis/BGPS

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22859 2026-06-09 cs.SE cs.AI 版本更新 81%

MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering

MEnvAgent:可扩展的多语言环境构建用于可验证软件工程

Chuanzhe Guo, Jingjing Wu, Sijun He, Yang Chen, Zhaoqi Kuang, Shilong Fan, Bingjin Chen, Siqi Bao, Jing Liu, Hua Wu, Qingfu Zhu, Wanxiang Che, Haifeng Wang

机构 * Tsinghua University(清华大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出MEnvAgent框架,通过多智能体规划-执行-验证架构和环境复用机制,自动构建多语言可执行环境,生成可验证任务实例,在10种语言1000个任务上提升F2P率8.6%并降低时间成本43%。

Comments Accepted as a Spotlight Paper at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09634 2026-06-08 cs.CL 版本更新 81%

Creation of the Estonian Subjectivity Dataset: Assessing the Degree of Subjectivity on a Scale

爱沙尼亚主观性数据集的创建:评估主观性程度的一个量表

Karl Gustav Gailit, Kadri Muischnek, Kairit Sirts

机构 * University of Tartu(塔尔图大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文创建了爱沙尼亚语文档级主观性数据集,通过连续量表标注并分析标注一致性,初步实验使用大语言模型进行自动主观性分析,发现自动评分可行但不可完全替代人工。

Comments 9 pages, 5 figures, 3 appendixes, LREC 2026

Journal ref Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026) 8204-8216

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24519 2026-08-17 cs.LG cs.AI cs.NE 版本更新 81%

A Negative-Control Protocol for Clinical EEG Foundation-Model Benchmarks: Dataset Identity and External-Cohort Stress Testing

用于临床解码的脑电图基础模型压力测试:数据集标识和目标阴性对照

Marzieh Zare

机构 * Université Laval(拉瓦尔大学) NeuroGenis Inc.(NeuroGenis公司)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.AI、cs.LG

AI总结 研究预训练脑电图基础模型在临床解码中的表现,通过多种方法在多数据集多任务上对六个模型进行基准测试,发现模型结论依赖评估单元、数据集偏移等因素,如在韩国痴呆症等任务中不同模型和方法表现各异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09992 2026-08-17 cs.CL cs.AI 版本更新 81%

A Unified Assessment of the Poverty of the Stimulus Argument for Neural Language Models

对神经语言模型中贫乏刺激论的统一评估

Xiulin Yang, Arianna Bisazza, Nathan Schneider, Ethan Gotlieb Wilcox

机构 * Georgetown University(乔治城大学) University of Groningen(格罗宁根大学)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本文通过poshbench评估了神经语言模型在贫乏刺激论中的表现,发现其泛化能力弱于儿童,但引入认知动机归纳偏置后提升语法能力,挑战了内在句法是唯一泛化途径的观点。

Comments TACL under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25589 2026-08-12 cs.CV cs.AI cs.CL 版本更新 81%

Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

放射学视觉语言模型基准的法证可重复性审计:从预期协议到发布工件

Mateusz Kozłowski

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 对胸部X光视觉语言模型试点进行法证可重复性审计,追踪提示绑定等多方面情况,发现存在图像渲染、数据分割等问题,重建队列改变统计值,撤回原声明并指定机器可验证控制。

Comments Withdrawn by the author. On further review, the archived artifacts underlying this audit are too incomplete to support the reported statistics, and the paper's conclusions do not follow from the available evidence. The work is withdrawn in full; earlier versions should not be cited

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05550 2026-08-11 cs.SD cs.CL cs.LG 版本更新 81%

Assessing Factual Music Comprehension in Large Audio Language Models

评估大型音频语言模型中的事实音乐理解能力

Daniel Chenyu Lin, Michael Freeman, John Thickstun

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL、cs.LG

AI总结 针对现有MusicQA数据集无法衡量模型回答事实正确性的问题,提出基于可验证信息的评估协议,通过精确率、召回率和F1分数客观评估模型,并在三个数据集上定义六项事实检索任务,对九个最新LALM进行基准测试。

Comments 17 pages; ISMIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20959 2026-08-07 cs.LG cs.CL 版本更新 81%

Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models

正确知识,错误答案:开放权重语言模型中时间事实冲突的测试时引导

Elias Hossain, Sourav Saha, Tasfia Nuzhat Ornee, Sanjeda Sara Jennifer, Umesh Chandra Biswas, Shubhashis Roy Dipta, Rajib Rana, Niloofar Yousefi

机构 * University of Central Florida(中佛罗里达大学) Mississippi State University(密西西比州立大学)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL、cs.LG

AI总结 针对开放权重语言模型中参数化时间冲突问题,提出无需重训练或外部检索的三阶段测试时干预方法TAS,通过单层激活修补实现0.72-0.85的答案翻转率,并保持非冲突查询85-99%的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12790 2026-07-31 cs.AI cs.CL cs.MA 版本更新 81%

Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents

谁来评估评估者?用于自我改进的语言模型代理的协同进化评估指标和技能

Xing Zhang, Guanghui Wang, Yanwei Cui, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He

机构 * Amazon(亚马逊)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.CL、cs.AI

AI总结 研究自我进化智能体系统中评估指标缺失问题,提出指标可进化,通过“双棘轮”协同进化指标与技能循环,在多任务中保留提升效果,还阐述安全性源于锚定规则与外部审核,为无可靠自动验证器场景提供架构。

Comments Code: https://github.com/amazon-science/Self-Evolving-Agents-Double-Ratchet

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20780 2026-07-31 cs.AI cs.CL cs.CV 版本更新 81%

MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models

MedHallTune:用于缓解视觉-语言模型中医患幻觉的指令调优基准

Qiao Yan, Yuchen Yuan, Xiaowei Hu, Yihan Wang, Jiaqi Xu, Xiwen Wu, Jinpeng Li, Chi-Wing Fu, Pheng-Ann Heng

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本研究提出MedHallTune基准,用于评估和缓解医疗视觉-语言模型的幻觉,实验表明用该基准微调模型可提升幻觉管理能力与下游VQA任务零样本性能,助力开发更可信的医疗视觉-语言模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12796 2026-07-28 cs.CL cs.AI cs.CY 版本更新 81%

The One-Word Census: Answer-Choice Conformity Across 44 Language Models

单字普查:44种语言模型中的答案选择一致性

Tapan Parikh

机构 * Cornell Tech(康奈尔科技)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 研究44种语言模型在单字选择上的一致性,用31个单轮提示词刻画,通过答案选择惊讶度评分,发现一致性程度高但模型间有差异且有结构,还对比了与人类规范差异,所有数据公开。

Comments 19 pages, 5 figures. Data, prompts, and code: https://github.com/tap2k/modelun

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05515 2026-07-28 cs.AI cs.CL cs.CV 版本更新 81%

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants

乐高协同构建器:探索用于多模态乐高组装助手的细粒度视觉语言建模

Haochen Huang, Yue Su, Xin Sun, Moonisa Ahsan, Mohammad Aliannejadi, Irene Viola, Zhaochun Ren, Chuang Yu, Aneta Lisowska, Artem Belopolsky, Koen Hindriks, Pablo Cesar, Junxiao Wang, Jiahuan Pei

机构 * Vrije University of Amsterdam University of Amsterdam Centrum Wiskunde \& Informatic University of Amsterdam National Institute of Informatics Centrum Wiskunde \& Informatic Centrum Wiskunde \& Informatic Leiden University University College London Vrije University of Amsterdam Centrum Wiskunde \& Informatica Technische Universiteit Delft Guangzhou University Vrije University of Amsterdam Centrum Wiskunde \& Informatic

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 研究针对视觉语言模型理解多模态组装指令的挑战,提出乐高协同构建器基准,引入统一框架评估多个VLMs及推理模型,发现物体检测性能高但细粒度场景理解和状态检测有挑战,还发布相关资源助力未来研究。

Comments This version has been accepted by ICMI 2026 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11945 2026-07-22 cs.CL cs.LG 版本更新 81%

Belief-reality separation lives in routing over a shared value slot in language models

信念与现实的分离存在于语言模型中共享值槽的路由中

Oliver Steele, Jiangtao Wen, Yuxing Han

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL、cs.LG

AI总结 研究语言模型中信念与现实分离的位置,核心方法是通过通用值槽和查询位置路由器,表明其基于两个可分离机制处于两个位置,结果在三种架构及多个模型家族成立,深入研究了信念 - 现实轴且该格式在多语境共享。

Comments 21 pages, 6 figures, 6 tables. v2: cite the released Mental Spaces Corpus (dataset DOI); switch to ACL bibliography style; minor copy-editing. No change to results

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18121 2026-07-22 cs.CL cs.AI 版本更新 81%

Saving the legacy of Hero Ibash: Evaluating Four Language Models for Aminoacian

拯救伊巴什英雄的遗产:评估四种用于氨基语的语言模型

Yunze Xiao, Yiyang Pan

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 研究评估四种前沿语言模型在氨基语中的表现,审视其在文本生成等方面的适应性等,揭示低资源语言中模型性能见解,为自然语言处理未来发展奠基,推动语言模型在类似语言环境中应用。

Comments arXiv admin note: This submission has been withdrawn due to violation of arXiv policies for acceptable submissions

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11098 2026-07-16 cs.SE cs.AI cs.CL 版本更新 81%

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP

AgentCheck:用于基于MCP的大型语言模型智能体的重现-干预-缓解工作台

Aritra Mazumder, Nusrat jahan Lia

机构 * University of Utah(犹他大学) University of Dhaka(达卡大学)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.CL、cs.AI

AI总结 研究针对工具使用型大语言模型智能体在工具出现问题时开发者难以处理故障的情况,提出AgentCheck开源工作台,通过特定运行和重放方式形成重现-干预-确认循环,能对故障模式进行评估验证,提升智能体故障处理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22758 2026-07-08 cs.CL cs.CE cs.LG 版本更新 81%

MASCA: LLM based-Multi Agents System for Credit Assessment

MASCA:基于大语言模型的信用评估多智能体系统

Gautam Jajoo, Atharva Pandey, Pranjal A Chitale, Saksham Agarwal

机构 * Kairosity(凯罗斯蒂)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.CL、cs.LG

AI总结 研究信用评估问题,提出基于大语言模型的MASCA多智能体系统,采用分层架构,整合对比学习,从信号博弈论角度提供理论见解,进行偏差分析,实验证明该系统在金融应用尤其是信用评分中表现优于基线方法。

Comments Accepted at NeurIPS GenAI In Finance Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏