arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-05-20 至 2026-05-20 共收录 396 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 93 篇

2605.19752 2026-05-20 cs.LG 83%

MSAlign: Aligning Molecule and Mass Spectra Foundation Models for Metabolite Identification

MSAlign: 用于代谢物鉴定的分子和质谱基础模型对齐方法

Paul Krzakala, Gabriel Melo, Camille Lançon, Charlotte Laclau, Rémi Flamary, Etienne Thévenot, Florence d'Alché-Buc

机构 * LTCI, Télécom Paris & CMAP, Ecole Polytechnique, Institut Polytechnique de Paris(LTCI,巴黎电信学院及巴黎高等技术学院的联合机构,CMAP,巴黎高等理工学院,巴黎高等技术学院) LTCI, Télécom Paris, Institut Polytechnique de Paris(LTCI,巴黎电信学院,巴黎高等技术学院) CEA, INRAE, MetaboHUB, Université Paris-Saclay(CEA,国家核能研究中心,法国农业研究机构,代谢组学枢纽,巴黎萨克雷大学)

专题命中 评测与基准 :foundation model(title,abstract);language model(abstract);分类 cs.LG

AI总结 本研究提出MSAlign方法,通过多模态对齐技术对齐分子和质谱基础模型,以提高代谢物鉴定的准确性,并解决了数据分割策略中的分布偏移问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19193 2026-05-20 cs.LG 83%

Sequential Consensus for Multi-Agent LLM Debates: A Wald-SPRT compute governor with calibration-based failure detection

多智能体大语言模型辩论中的顺序共识:一种基于Wald-SPRT的计算控制器与基于校准的故障检测

Andrea Morandi

机构 * Cisco(思科)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.LG

AI总结 本文提出了一种基于Wald-SPRT的计算控制器,用于多智能体大语言模型辩论,通过校准来检测故障,从而在保证准确性的同时减少计算资源的使用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17340 2026-05-20 cs.LG 83%

Olivia: Harmonizing Time Series Foundation Models with Power Spectral Density

Olivia:通过功率谱密度和谐化时间序列基础模型

Jingru Fei, Kun Yi, Alex Xing Wang, Qingsong Wen, Xiangxiang Zhu, Wei Fan

机构 * Beijing Institute of Technology(北京理工大学) North China Institute of Computing Technology(华北计算技术研究所) State Information Center(国家信息中心) University of Auckland(奥克兰大学) Northwest Polytechnical University(西北工业大学) Victoria University of Wellington(威灵顿维多利亚大学)

专题命中 评测与基准 :foundation model(title,abstract);pretraining(abstract);分类 cs.LG

AI总结 本文提出Olivia,一种基于谐化机制的时间序列基础模型,通过在频域中使用功率谱密度来减少数据集间的不匹配并增强预训练效果,从而在零样本、少样本和全样本预测场景中取得最佳性能。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19966 2026-05-20 cs.LG cs.AI 82%

Detecting Fluent Optimization-Based Adversarial Prompts via Sequential Entropy Changes

通过顺序熵变化检测基于优化的对抗性提示

Mohammed Alshaalan, Miguel R. D. Rodrigues

机构 * Department of Electronic and Electrical Engineering, University College London, London, United Kingdom(电子与电气工程系,伦敦大学学院,伦敦,英国)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种基于在线变化点检测的对抗性后缀检测方法CPD,通过标准化用户令牌熵并应用单侧CUSUM统计量,提高了对优化基于对抗性提示的检测性能,同时在多个大型语言模型上实现了更高的F1分数和AUC性能。

Comments Accepted at ICML 2026; 20 pages, including 9 pages main text, references, and appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19351 2026-05-20 cs.MA cs.AI cs.CL 82%

PAVE: A Cognitive Architecture for Legitimate Violation in Generative Agent Societies

PAVE:生成代理社会中的合法违规认知架构

Ahmad Yehia, Abduallah Mohamed, Kun Qian, Tianyi Wang, Jiseop Byeon, Omar Hassanin, Christian Claudel

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Meta Reality Labs(Meta现实实验室) University of Calgary(卡尔加里大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出PAVE认知架构,通过四个模块处理生成代理在需要违规的场景中的推理问题,实现了合法违规、对权威的服从、有限的范围和恢复四个特性,同时提高了决策的结构化和可解释性。

Comments Preprint. 23 pages, 4 figures. Code and environment will be released upon publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19093 2026-05-20 cs.AI cs.LG 82%

Embedding by Elicitation: Dynamic Representations for Bayesian Optimization of System Prompts

通过 elicitation 进行嵌入:用于系统提示贝叶斯优化的动态表示

Zhiyuan Jerry Lin, Benjamin Letham, Samuel Dooley, Maximilian Balandat, Eytan Bakshy

机构 * Meta

专题命中 评测与基准 :LLM(summary_cn,abstract);分类 cs.AI、cs.LG

AI总结 本文研究了在仅有聚合反馈的情况下,如何通过动态表示进行系统提示的贝叶斯优化,提出了一种基于 elicitation 的嵌入方法 ReElicit,利用 LLM 构建可解释的特征空间,并通过概率高斯过程代理选择目标特征向量,最终实现系统提示的优化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19111 2026-05-20 cs.CV cs.AI 82%

FAGER: Factually Grounded Evaluation and Refinement of Text-to-Image Models

FAGER:基于事实的文本到图像模型评估与改进

Youngsun Lim, Cusuh Ham, Pin-Yu Chen, Deepti Ghadiyaram

机构 * Boston University(波士顿大学) Adobe(Adobe公司) IBM Research(IBM研究院)

专题命中 评测与基准 :LLM(summary_cn,abstract);分类 cs.AI;foundation model(comments)

AI总结 本文提出FAGER框架,用于评估和改进文本到图像模型的事实准确性,通过结合LLM生成事实和参考引导的视觉事实提取与验证,构建结构化事实评估标准,并通过VLM进行评估,验证FAGER在事实性测试中优于现有方法,并能无训练改进T2I输出。

Comments It was accepted for an oral presentation at the 2nd Workshop on the Evaluation of Generative Foundation Models (EVGENFM2026) at CVPR 2026. Total 8 pages (1 page for references). 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18413 2026-05-20 cs.CV 82%

Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models

基础的裂缝:一个挑战视觉基础模型的民用基础设施数据集

Nicola Farronato, Niccolo Avogaro, Thomas Frick, Mattia Rigotti, Rizwan Ullah Khan, Michele Magno, Konrad Schindler, Cristiano Malossi, Florian Scheidegger

机构 * IBM Research(IBM研究院) ETH Zürich(苏黎世联邦理工学院) University of Twente(特文特大学)

专题命中 评测与基准 :foundation model(title,abstract);language model(abstract)

AI总结 本文提出Cracks in the Foundation数据集,通过高分辨率图像挑战视觉基础模型在民用基础设施中的密集图像理解能力,揭示了现有模型在真实世界中的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18824 2026-05-20 cs.LG cs.AI cs.CL 82%

Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models

细粒度基准生成用于基础模型的全面评估

Mohammed Saidul Islam, Negin Baghbanzadeh, Farnaz Kohankhaki, Afshin Cheraghi, Ali Kore, Shayaan Mehdi, Elham Dolatabadi, Arash Afkanpour

机构 * Vector Institute(Vector研究院) York University(约克大学)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出了一种自动化基准生成框架,用于生成覆盖广泛、元数据丰富且抗污染的评估问题,从而提升基础模型的全面评估能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19194 2026-05-20 cs.CL 81%

MMoA: An AI-Agent framework with recurrence for Memoried Mixure-of-Agent

MMoA: 一个具有递归性的记忆混合代理框架

Rui Chu

机构 * Rui Chu(楚瑞)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出MMoA框架,通过引入LSTM门控机制,改进了传统混合代理方法在时间依赖性和上下文感知方面的不足,实现了更高效的多代理系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15768 2026-05-20 cs.AI cs.CY 81%

ALSO: Adversarial Online Strategy Optimization for Social Agents

ALSO: 用于社交代理的对抗在线策略优化

Xiang Li, Liping Yi, Mingze Kong, Min Zhang, Zhongxiang Dai, QingHua Hu

机构 * School of Artificial Intelligence, Tianjin University, Tianjin, China(天津大学人工智能学院,天津,中国) The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)) East China Normal University, Shanghai, China(华东师范大学,上海,中国)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出ALSO框架,通过将多轮交互建模为对抗性带薪问题,并引入轻量级神经代理来预测奖励,从而在动态环境中实现社交代理的鲁棒策略优化。

Comments Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14700 2026-05-20 cs.SI cs.CL cs.CY 81%

Context-Aware Detection and Victim-Centered Response Generation for Online Harassment in Private Messaging

基于上下文的在线骚扰检测与以受害者为中心的回应生成:私人信息交流中的在线骚扰

Pinxian Lu, Nimra Ishfaq, Emma Win, Morgan Rose, Sierra R Strickland, Candice L Biernesser, Jamie Zelazny, Munmun De Choudhury

机构 * Georgia Institute of Technology(佐治亚理工学院) The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Pittsburgh(匹兹堡大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文研究了大型语言模型如何支持私人信息交流中的在线骚扰检测与回应,通过构建一个包含80,053条Instagram私信的标注数据集,开发了上下文感知的级联分类流水线,并提出了一种以受害者为中心的回应框架,生成心理上合理的AI回应,经评估发现其在情感支持和缓和冲突方面显著优于原始回应。

Comments 16 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19600 2026-05-20 cs.RO 80%

FlyMirage: A Fully Automated Generation Pipeline for Diverse and Scalable UAV Flight Data via Generative World Model

FlyMirage: 一种用于生成多样化和可扩展的无人机飞行数据的完全自动化生成流程

Jinhan Li, Xijie Huang, Zhaoqi Wang, Yijin Wang, Weiqi Ge, Qiyi He, Mo Zhu, Fei Gao, Yuze Wu, Xin Zhou

机构 * State Key Laboratory of Industrial Control Technology, Zhejiang University, Hangzhou 310027, China(浙江大学工业控制技术状态重点实验室,杭州310027,中国) Differential Robotics, Hangzhou 311121, China(差分机器人,杭州311121,中国)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文提出FlyMirage,一种完全自动化的生成流程,通过生成世界模型生成大规模、多样化且逼真的无人机视觉-语言导航数据,支持下一代具身导航模型的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18806 2026-05-20 cs.IR 80%

Towards FairRAG: Preventing Representational Harm in Retrieval-Augmented Generation by Enforcing Fair Exposure at Retrieval Time

迈向公平RAG:通过在检索时间强制公平曝光来防止表示性伤害

Riddhi Tikoo

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文研究了如何通过在检索阶段强制公平曝光来减少检索增强生成中表示性伤害,提出了一种新的公平检索排名方法,并通过实验验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17370 2026-05-20 cs.AI 79%

CBT-Audio: Evaluating Audio Language Models for Patient-Side Distress Intensity Estimation in CBT Session Recordings

CBT-Audio: 评估音频语言模型以估计CBT会话录音中患者压力强度

Qixuan Hu, Shuchang Ye, Xumou Zhang, Anastasia Serafimovska, Anastasia Suraev, Amit Saha, Ping-hsiu Lin, Sydney Su, Usman Naseem, Adam G. Dunn, Jinman Kim

机构 * School of Computer Science, Faculty of Engineering, University of Sydney, Australia(悉尼大学工程学院计算机科学学院,澳大利亚) School of Psychology, Faculty of Science, University of Sydney, Australia(悉尼大学科学学院心理学学院,澳大利亚) School of Computing, Faculty of Science and Engineering, Macquarie University, Australia(麦考瑞大学科学与工程学院计算学院,澳大利亚) CHeBA (Centre for Healthy Brain Ageing), School of Clinical Medicine, Discipline of Psychiatry & Mental Health, The University of New South Wales, Australia(新南威尔士大学临床医学学院精神病与心理健康学科健康大脑年龄中心,澳大利亚) Sydney School of Public Health, Faculty of Medicine and Health, University of Sydney, Australia(悉尼大学医学与健康学院公共卫生学院,澳大利亚)

专题命中 评测与基准 :language model(title,abstract);分类 cs.AI

AI总结 本文提出CBT-Audio数据集,用于评估音频语言模型在估计CBT会话中患者压力强度方面的性能,通过结合音频和文本输入提升了压力强度估计的准确性。

Comments 9 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16593 2026-05-20 cs.CL 79%

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models

重新审视一个令人头疼的问题:一种用于语言模型的语义推理基准

Yang Liu, Hongming Li, Melissa Xiaohui Qin, Qiankun Liu, Chao Huang

机构 * University of Science and Technology Beijing(北京科技大学) State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

AI总结 本文提出SemanticQA基准,用于评估语言模型在语义短语处理任务中的表现,通过整合现有多词表达资源并重新组织为统一测试平台,涵盖通用词汇现象及三种细粒度类别,评估不同架构和规模的语言模型在提取、分类、解释及任务组合中的性能,揭示语义推理任务中模型性能的显著差异,为提升语言模型在非平凡语义短语上的理解能力提供见解。

Comments ACL 2026 (Oral), 24 pages, 22 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20086 2026-05-20 cs.NE cs.AI cs.LG 79%

What Do Evolutionary Coding Agents Evolve?

进化编码代理进化什么?

Nico Pelleriti, Sree Harsha Nelaturu, Zhanke Zhou, Zongze Li, Max Zimmer, Bo Han, Sebastian Pokutta

机构 * Zuse Institute Berlin(柏林Zuse研究所) Technical University of Berlin(柏林技术大学) Hong Kong Baptist University(香港 Baptist大学) RIKEN Center for Advanced Intelligence Project(RIKEN高级智能项目中心)

专题命中 评测与基准 :LLM(abstract,abstract_cn);prompting(abstract);分类 cs.AI、cs.LG

AI总结 本文研究了进化编码代理在数学发现和算法设计中通过任务特定反馈生成、修改和选择代码的过程,通过EvoTrace数据集和EvoReplay方法分析了进化过程中的机制,发现大部分得分提升来自少数几种编辑类型,并发现存在确定性的循环模式。

Comments 28 pages, 12 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05201 2026-05-20 eess.AS 78%

Exploring Speech Foundation Models for Speaker Diarization Across Lifespan

探索跨生命周期的语音基础模型用于说话人分离

Anfeng Xu, Tiantian Feng, Shrikanth Narayanan

专题命中 评测与基准 :foundation model(title,abstract)

AI总结 本文研究了语音基础模型在跨生命周期说话人分离中的鲁棒性,通过统一的端到端神经分离框架(EEND-VC)评估了儿童、成人和老年人的语音样本,比较了零样本跨年龄推理、联合多年龄训练和领域特定适应方法,发现成人训练模型在儿童和老年人数据上性能下降,而联合多年龄训练和针对性年龄组适应提升了分离性能,特别是使用Whisper编码器时。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25620 2026-05-20 cs.CL 77%

PICon: A Multi-Turn Interrogation Framework for Evaluating Persona Agent Consistency

PICon: 一种用于评估人设代理一致性的多轮询问框架

Minseo Kim, Sujeong Im, Junseong Choi, Junhee Lee, Chaeeun Shim, Hwajung Hong, Edward Choi

机构 * KAIST(韩国科学技术院)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出PICon框架,通过逻辑链式的多轮提问评估人设代理的一致性,发现即使之前被认为高度一致的系统在三个维度上也未能达到人类基准水平,揭示了矛盾和逃避回应。

Comments 20 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22258 2026-05-20 cs.CV cs.AI 77%

Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks

超越分类准确度:Neural-MedBench与更深层次推理基准的需求

Miao Jing, Mengting Jia, Junling Lin, Zhongxia Shen, Huan Gao, Mingkun Xu, Shangyang Li

机构 * School of Physics Science and Technology, Beijing University of Posts and Telecommunications(北京邮电大学物理科学与技术学院) Guangdong Institute of Intelligence Science and Technology(广东智能科学技术研究院) Beijing Chaoyang Hospital, Capital Medical University(北京朝阳医院) Sleep Medical Center, Huzhou Third Municipal Hospital, Affiliated Hospital of Wenzhou Medical University(湖州第三人民医院睡眠医学中心,温州医科大学附属医院) University of Macau(澳门大学) Renyixun Health Technology Co., Ltd(仁颐讯健康科技有限公司) Academy for Advanced Interdisciplinary Studies, Peking University(北京大学交叉学科研究院)

专题命中 评测与基准 :LLM(abstract,abstract_cn);language model(abstract);分类 cs.AI

AI总结 本文提出Neural-MedBench,一个专门用于测试多模态神经病学推理能力的基准,揭示现有医疗数据集过于强调分类准确度的问题,并通过系统评估发现模型推理失败而非感知误差主导性能下降,强调需要兼顾广度与深度的评估框架。

Comments 23 pages, 12 figures

Journal ref ICLR'2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18791 2026-05-20 eess.IV cs.CV cs.LG q-bio.OT 77%

SpecX: A Large-Scale Benchmark for Multi-Modal Spectroscopy and Cross-Paradigm Evaluation

SpecX:多模态光谱的大规模基准及跨范式评估

Chengrui Xiang, Tengfei Ma, Yujie Chen, Tong Wang, Haowen Chen, Xiangxiang Zeng

机构 * College of Computer Science and Technology, Hunan University(湖南大学计算机科学与技术学院)

专题命中 评测与基准 :language model(abstract);foundation model(abstract);pretraining(abstract);分类 cs.LG

AI总结 本文提出SpecX,一个用于多模态光谱的大规模基准,通过不同层级的数据集支持分子解析、光谱模拟和理解任务,揭示了专用光谱模型和多模态语言模型在光谱智能中的不同优势。

Comments 9 pages,1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20147 2026-05-20 cs.CV 75%

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset

PixVerve:通过大规模高质量数据集将原生超高清图像生成推至100MP

Haojun Chen, Haoyang He, Chengming Xu, Qingdong He, Junwei Zhu, Yabiao Wang, Zhucun Xue, Xianfang Zeng, Zhennan Chen, Xiaobin Hu, Hao Zhao, Yong Liu, Jiangning Zhang, Dacheng Tao

机构 * Zhejiang University(浙江大学) Fudan University(复旦大学) Nanjing University(南京大学) National University of Singapore(新加坡国立大学) Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);foundation model(abstract)

AI总结 本文提出PixVerve-95K数据集,通过精心设计的数据管道构建,包含95K张高分辨率图像和七维标注,用于推动超高清图像生成技术,通过三种训练方案将T2I基础模型扩展到100MP生成,并建立PixVerve-Bench评估协议。

Comments Project page is available at https://haojunchen663.github.io/projects/PixVerve/

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01411 2026-05-20 cs.SE 75%

CodePori: Large-Scale System for Autonomous Software Development Using Multi-Agent Technology

CodePori:利用多智能体技术实现大规模自主软件开发的系统

Zeeshan Rasheed, Muhammad Waseem, Kai-Kristian Kemell, Aakash Ahmad, Malik Abdul Sami, Mika Saari, Jussi Rasku, Pekka Abrahamsson

专题命中 评测与基准 :LLM(summary_cn,abstract)

AI总结 本文研究了基于大语言模型的多智能体系统在自主软件开发中的潜力与局限,通过开发CodePori系统和参与式评估,揭示了LLM多智能体系统的优势、挑战及改进方向。

Comments 18 pages, 8 figures, and 4 Table

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19341 2026-05-20 cs.CL cs.AI cs.LG stat.ML 75%

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models

HalluWorld: 一个用于通过参考世界模型控制幻觉的基准

Emmy Liu, Varun Gangal, Michael Yu, Zhuofu Tao, Karan Singh, Sachin Kumar, Steven Y. Feng

机构 * Carnegie Mellon University(卡内基梅隆大学) Patronus AI Independent Researcher(独立研究者) Stanford University(斯坦福大学) The Ohio State University(俄亥俄州立大学) DegenAI Labs(DegenAI实验室)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出HalluWorld基准,通过显式参考世界模型研究语言模型的幻觉问题,发现不同任务中幻觉表现不一致,表明幻觉源于多种失败模式而非单一能力。

Comments HalluWorld benchmark (code and data) at github.com/DegenAI-Labs/HalluWorld

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11024 2026-05-20 cs.CV cs.AI 74%

Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style

AI 是否能像艺术史家一样看?解析视觉语言模型如何识别艺术风格

Marvin Limpijankit, Milad Alshomary, Yassin Oulad Daoud, Amith Ananthram, Tim Trombley, Emily L. Spratt, Anna Filonenko, Hannah Pivo, Elias Stengel-Eskin, Mohit Bansal, Noam M. Elcott, Kathleen McKeown

机构 * Columbia University, Department of Computer Science(哥伦比亚大学计算机科学系) Columbia University, Department of Art History & Archaeology(哥伦比亚大学艺术史与考古系) University of Texas at Austin(德克萨斯大学奥斯汀分校) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 评测与基准 :language model(title);分类 cs.AI

AI总结 本文研究了视觉语言模型(VLMs)在识别艺术风格方面的机制,通过跨学科合作,分析VLMs如何预测艺术风格,并评估其与艺术史家判断艺术风格的标准的一致性。

Comments 20 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11796 2026-05-20 cs.CL cs.AI 73%

C-ReD: A Comprehensive Chinese Benchmark for AI-Generated Text Detection Derived from Real-World Prompts

C-ReD:一个源自真实世界提示的综合性中文AI生成文本检测基准

Chenxi Qing, Junxi Wu, Zheng Liu, Yixiang Qiu, Hongyao Yu, Bin Chen, Hao Wu, Shu-Tao Xia

机构 * Tsinghua University(清华大学) Nankai University(南开大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Peng Cheng Laboratory(鹏城实验室) Shannon InfoTech

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出C-ReD基准,用于检测AI生成的中文文本,通过解决模型多样性、领域覆盖和提示真实性等关键问题,提升检测性能和泛化能力。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18857 2026-05-20 cs.IR cs.AI cs.LG 73%

The 99% Success Paradox: When Near-Perfect Retrieval Equals Random Selection

99%成功悖论:当近完美检索等于随机选择

Vyzantinos Repantis, Harshvardhan Singh, Tony Joseph, Cien Zhang, Akash Vishwakarma, Svetlana Karslioglu, Michael Wyatt Thot, Ameya Gawde

机构 * Meta Platforms Inc.(Meta平台公司)

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 该研究引入了Bits-over-Random(BoR)指标,揭示了高成功率可能掩盖随机水平性能的现象,指出在大规模数据集上,即使检索结果覆盖率达到99%,其选择性仍可能接近零,从而表明需要重新考虑检索深度和传统指标的报告方式。

Comments 12 pages, 2 figures, 7 tables. Accepted at ICLR 2026 Blog Track, https://iclr-blogposts.github.io/2026/blog/2026/bits-over-random/

Journal ref ICLR Blog Track 2026, https://iclr.cc/virtual/2026/poster/10012083

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12217 2026-05-20 cs.CV 71%

HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models

HyperCap:面向视觉语言模型的超光谱土地覆盖描述数据集

Aryan Das, Tanishq Rachamalla, Pravendra Singh, Koushik Biswas, Vinay Kumar Verma, Salvador Garcia, Antonio Plaza, Swalpa Kumar Roy

机构 * Department of Computer Science and Engineering, Vellore Institute of Technology(计算机科学与工程系,维洛雷理工学院) Department of Information Technology, Siddhartha Academy of Higher Education(信息技术系,斯里达拉塔高等教育学院) Department of Computer Science and Engineering, Indian Institute of Technology, Roorkee(计算机科学与工程系,印度理工学院罗尔基分校) Department of Computer Science and Engineering, Indraprastha Institute of Information Technology Delhi(计算机科学与工程系,印度信息技术学院德里) Department of Computer Science and Engineering, Indian Institute of Technology, Kanpur(计算机科学与工程系,印度理工学院坎浦尔) Department of Computer Science and Artificial Intelligence, University of Granada(计算机科学与人工智能系,格拉纳达大学) Hyperspectral Computing Laboratory, Department of Computers and Communications, University of Extremadura(超光谱计算实验室,计算机与通信系,埃斯特拉达大学)

专题命中 评测与基准 :language model(title)

AI总结 本文提出HyperCap数据集,通过整合光谱数据与像素级文本标注,提升遥感应用中的模型性能,为未来研究提供基础资源。

Comments Accepted for publication in IEEE Geoscience and Remote Sensing Magazine (GRSM), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19798 2026-05-20 cs.CL 70%

Towards Trust Calibration in Socially Interactive Agents: Investigating Gendered Multimodal Behaviors Generation with LLMs

迈向社交互动代理的信任校准:探究LLMs生成性别化的多模态行为生成

Lucie Galland, Chloé Clavel, Magalie Ochs

机构 * LIS Laboratory, Amu(LIS实验室,Amu) Inria Paris(Inria巴黎)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文研究了LLMs生成多模态行为以反映能力与善意的不同层次,探讨了性别对行为生成的影响,并通过用户研究验证了方法的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13318 2026-05-20 cs.AI cs.ET 70%

VERA-MH: Validation of Ethical and Responsible AI in Mental Health

VERA-MH:心理健康领域伦理和负责任AI的验证

Luca Belli, Kate H. Bentley, Josh Gieringer, Emily Van Ark, Nilu Zhao, Pradip Thachile, Matt Hawrilenko, Millard Brown, Adam M. Chekroud

机构 * Spring Health Yale University(耶鲁大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.AI

AI总结 本研究提出VERA-MH,一种用于评估心理健康支持聊天机器人安全性的新型临床验证方法,重点评估聊天机器人在识别自杀倾向风险方面的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏