arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-05-05 至 2026-05-05 共收录 190 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 57 篇

2605.01264 2026-05-05 cs.SE cs.LG 81%

FeedbackLLM: Metadata driven Multi-Agentic Language Agnostic Test Case Generator with Evolving prompt and Coverage Feedback

FeedbackLLM: 基于元数据的多智能体语言无关测试用例生成器与演进提示和覆盖反馈

Kushal Jasti, Tejamani Prashanth Sahu, Rishitha Pentyala, Muvvala Mohit, Vivek Yelleti

机构 * Department of Computer Science and Engineering, SRM University AP, India(印度SRM大学计算机科学与工程系)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 FeedbackLLM通过双阶段方法生成测试用例,利用元数据反馈提升覆盖率,有效解决传统方法的计算开销和可扩展性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00957 2026-05-05 cs.IR cs.AI 81%

"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation

我不知道--迈向具有确定性意识的适当信任

Daan Di Scala, Maaike de Boer, Pınar Yolum

机构 * TNO Netherlands Organisation for Applied Scientific Research, Department Data Science(荷兰应用科学研究院,数据科学部门) Utrecht University, Department of Information and Computing Sciences(乌得勒支大学,信息与计算科学系)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出CERTA系统,通过结合问题、上下文和答案的相关性来反映不确定性,以建立适当的信任。研究创建了Certainty Benchmark,并通过实验验证了CERTA在减少过度同意和提供谨慎行为方面的有效性。

Comments To be published in VALE 2025 Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14607 2026-05-05 cs.AI 81%

GDPR Auto-Formalization with AI Agents and Human Verification

GDPR自动形式化与AI代理和人类验证

Ha Thanh Nguyen, Wachara Fungwacharakorn, Sabine Wehnert, May Myo Zin, Yuntao Kong, Jieying Xue, Michał Araszkiewicz, Randy Goebel, Ken Satoh

机构 * Center for Juris-Informatics, ROIS-DS(法律信息中心,ROIS-DS) Ruhr-University Bochum, RC-Trust(波恩鲁尔大学,RC-Trust) Alberta Machine Intelligence Institute, University of Alberta(阿尔伯塔机器智能研究所,阿尔伯塔大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文研究利用大语言模型在人类在环验证框架下实现GDPR条款自动形式化的整体过程,通过角色专业化工作流结合人类审查,构建高质量数据集并分析成功与问题案例,证明结构化验证和针对性人类监督对可靠法律形式化至关重要。

Comments Accepted at ICAIL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01311 2026-05-05 cs.LG econ.EM stat.AP stat.ML 79%

The Partial Testimony of Logs: Evaluation of Language Model Generation under Confounded Model Choice

日志的部分证词:在混杂模型选择下语言模型生成的评估

Jikai Jin, Vasilis Syrgkanis

专题命中 评测与基准 :language model(title,abstract);分类 cs.LG

AI总结 本文研究了在混杂模型选择下语言模型生成评估的偏差问题,提出结合大规模观察日志、小规模随机实验和离线模拟器的三源设计,通过识别定理证明随机实验和模拟器可恢复因果模型值,日志用于减少估计误差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00907 2026-05-05 cs.CV cs.AI cs.LG 79%

TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation

TRIP-Evaluate: 一个用于评估交通领域大模型的开放多模态基准

Han Gong, Zhen Zhou, Yunyang Shi, Yan Tan, Jinbiao Huo, Qi Hong, Zhiyuan Liu

机构 * School of Transportation(交通学院) Southeast University(东南大学) School of Artificial Intelligence and Computer Science(人工智能与计算机科学学院) Jiangnan University(江南大学) Department of Civil and Environmental Engineering(土木与环境工程系) Hong Kong Polytechnic University(香港理工大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.AI、cs.LG

AI总结 TRIP-Evaluate是首个开放多模态交通评估基准,通过837项任务覆盖车辆、交通管理、旅行者和规划设计功能,提供能力、模态和难度标签,支持跨模态诊断,揭示大模型在多步工程计算、规则约束推理和多模态场景理解方面的不足。

Comments 19 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24056 2026-05-05 cs.CR cs.CL cs.LG 79%

Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness

Logit-Gap Steering:一种对齐鲁棒性的前向传递诊断

Tung-Ling Li, Hongliang Liu

机构 * Palo Alto Networks(帕洛阿尔托网络)

专题命中 评测与基准 :RLHF(abstract,abstract_cn);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文提出logit-gap steering方法,通过前向传递发现短的分布内后缀以关闭对齐间隙,验证了当前对齐边距的薄且可测量,强调防御策略需考虑分布内后缀。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00970 2026-05-05 cs.IT math.IT 78%

Split and Aggregation Learning for Foundation Models Over Mobile Embodied AI Network (MEAN): A Comprehensive Survey

分割与聚合学习用于移动具身人工智能网络(MEAN)上的基础模型:综合综述

Qianzhou Chen, Siqi Sun, Minrui Xu, Sijie Ji, Jiawen Kang, Yijie Mao, Zhouxiang Zhao, Zhaohui Yang, Dusit Niyato

专题命中 评测与基准 :foundation model(title,abstract)

AI总结 本文综述了6G通信系统中分割学习与聚合学习的应用,分析了其架构、技术方法及与AI原生6G通信技术的结合,探讨了其在语义通信、RIS、SAGIN和量子通信中的应用,旨在提升分布式基础模型的效率、隐私保护与可扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00902 2026-05-05 cs.CV cs.IR 78%

Validation of Whole-Slide Foundation Models for Image Retrieval in TCGA Data

对TCGA数据中整张滑动图像检索的整张滑动图像基础模型进行验证

Tianhao Lei, Parsa Esmaeilkhani, Saghir Alfasly, Wataru Uegami, Judy C. Boughey, Matthew P. Goetz, Krishna R. Kalari, H. R. Tizhoosh

机构 * KIMIA Lab, Artificial Intelligence and Informatics, Mayo Clinic, Rochester, MN, USA(KIMIA实验室,人工智能与信息学,梅奥诊所,罗切斯特,明尼苏达州,美国) Department of Neurology, Northwestern University Feinberg School of Medicine, Chicago, IL, USA(神经病学系,北western大学费因伯格医学院,芝加哥,伊利诺伊州,美国) Department of Computer Science, Temple University, Philadelphia, PA, USA(计算机科学系,泰勒大学,费城,宾夕法尼亚州,美国) Department of Breast and Melanoma Surgical Oncology, Comprehensive Cancer Center, Mayo Clinic, Rochester, MN, USA(乳腺和黑色素瘤外科肿瘤学系,综合癌症中心,梅奥诊所,罗切斯特,明尼苏达州,美国) Department of Oncology, Comprehensive Cancer Center, Mayo Clinic, Rochester, MN, USA(肿瘤学系,综合癌症中心,梅奥诊所,罗切斯特,明尼苏达州,美国) Department of Quantitative Health Sciences, Mayo Clinic, Rochester, MN, USA(定量健康科学系,梅奥诊所,罗切斯特,明尼苏达州,美国)

专题命中 评测与基准 :foundation model(title,abstract)

AI总结 本研究验证了TCGA数据中整张滑动图像基础模型的检索性能,发现基于片段和监督聚合的方法在Top-1和Top-3准确率上表现相当,但整体性能受器官和诊断差异影响较大,形态学检索存在固有限制,需进一步改进特征表示和多模态框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01664 2026-05-05 cs.IR 75%

A Hybrid Retrieval and Reranking Framework for Evidence-Grounded Retrieval-Augmented Generation

一种混合检索与重排序框架用于证据导向的检索增强生成

Fariba Afrin Irany, Sampson Akwafuo

专题命中 评测与基准 :large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出一种混合检索与重排序框架,用于生物医学和医疗相关文档问答中的引用感知RAG,通过亚马逊Bedrock知识库进行文档处理,并通过Cohere重排序提升相关性,最终实现高准确率的证据基础生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00842 2026-05-05 cs.AI cs.LG 73%

Understanding Emergent Misalignment via Feature Superposition Geometry

通过特征叠加几何理解涌现对齐问题

Gouki Minegishi, Hiroki Furuta, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo

机构 * The University of Tokyo(东京大学) Google DeepMind(谷歌DeepMind)

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 研究通过特征叠加几何解释了微调窄任务导致有害行为的机制,发现有害特征在几何上更接近,并通过过滤接近毒特征的数据减少对齐问题。

Comments Accepted to ACL2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02293 2026-05-05 cs.CL cs.AI 73%

Breaking the Silence: A Dataset and Benchmark for Bangla Text-to-Gloss Translation

打破沉默:Bangla文本到词组翻译的数据集和基准

Sharif Mohammad Abdullah, Abhijit Paul, Shubhashis Roy Dipta, Zarif Masud, Shebuti Rayana, Ahmedul Kabir

机构 * IIT, University of Dhaka, Bangladesh(达卡大学理工学院,孟加拉国) University of Maryland, Baltimore County, USA(马里兰大学巴尔的摩县分校,美国) SUNY, Old Westbury, USA(SUNY,美国奥尔德韦斯特伯里)

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本文首次构建了Bangla文本到词组翻译的数据集,通过人工标注和合成数据评估低资源手语翻译的挑战,展示了系统生成的合成数据的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01604 2026-05-05 cs.AI 72%

Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework

在真实环境中评估代理AI:故障模式、漂移模式及生产评估框架

Mukund Pandey

机构 * Independent Researcher(独立研究者)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI;LLM(comments)

AI总结 本文针对代理AI在生产环境中持续运行时的评估挑战,提出七种故障模式分类及PAEF框架,揭示传统指标在检测故障模式上的不足。

Comments 11 pages, 6 tables, 1 figure. Reference implementation: https://github.com/mukund1985/llm-eval-toolkit

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01647 2026-05-05 cs.CL 70%

Beyond Perplexity: Character Distribution Signatures and the MDTA Benchmark for AI Text Detection

超越困惑度:字符分布签名与AI文本检测的MDTA基准

Priyadarshan Narayanasamy, Swastik Agrawal, Klint Faber, Fardina Fathmiul Alam

机构 * University of Maryland, College Park(马里兰大学学院市分校)

专题命中 评测与基准 :RLHF(abstract,abstract_cn);分类 cs.CL

AI总结 本文提出基于字符分布签名的AI文本检测方法,通过MDTA基准验证其有效性,显示与困惑度方法低相关性,并在特定领域取得显著提升。

Comments 11 figures, 10 tables, 24 pages, Under Review at COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01404 2026-05-05 cs.AR cs.AI 70%

AMSnet-q: Unsupervised Circuit Identification and Performance Labeling for AMS Circuits

AMSnet-q:无监督的电路识别与性能标注用于AMS电路

Ze Zhang, Junzhuo Zhou, Yichen Shi, Zhuofu Tao, Rui Ji, Zhiping Yu, Quan Chen, Ting-Jung Lin, Lei He

机构 * Southern University of Science and Technology(南方科技大学) University of California Los Angeles(加州大学洛杉矶分校) Tsinghua University(清华大学) Eastern Institute of Technology Ningbo(宁波东部技术研究院)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 AMSnet-q通过无监督方法自动构建AMS电路数据库,实现电路功能验证和性能标注,无需人工干预。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01359 2026-05-05 cs.AI 70%

Structural Ranking of the Cognitive Plausibility of Computational Models of Analogy and Metaphors with the Minimal Cognitive Grid

认知可信度的结构排序:基于最小认知网格的类比与隐喻计算模型评估

Alessio Donvito, Antonio Lieto

机构 * University of Bari(巴里大学) University of Salerno(萨勒诺大学) CIIT Lab / ICAR-CNR(CIIT实验室 / ICAR-CNR)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文利用最小认知网格框架系统评估类比与隐喻计算模型的认知可信度,通过分析三个核心维度比较不同模型与认知理论的一致性。

Comments 35 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01338 2026-05-05 cs.AI 70%

DiagramNet: An End-to-End Recognition Framework and Dataset for Non-Standard System-Level Diagrams

DiagramNet: 一种面向非标准系统级图表的端到端识别框架和数据集

Jincheng Lou, Ruohan Xu, Jiapeng Li, Junyin Pi, Runzhe Tao, Weijian Fan, Xiao Tan, Guojie Luo, Yibo Lin

机构 * School of IC, Peking University, Beijing, China(北京大学信息科学技术学院) College of Artificial Intelligence, Xi'an Jiaotong University, Xi'an, China(西安交通大学人工智能学院) Department of Precision Instruments, Tsinghua University, Beijing, China(清华大学精密仪器系) School of Software and Microelectronics, Peking University, Beijing, China(北京大学软件与微电子学院) School of Computer Science, Peking University, Beijing, China(北京大学计算机科学系) Institute of EDA, Peking University, Beijing, China(北京大学EDA研究院) Beijing Advanced Innovation Center for IC, Beijing, China(北京集成电路先进创新中心)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出DiagramNet数据集和训练框架,通过端到端方法在系统级图表识别中超越现有模型,提升多模态大语言模型的图表理解能力。

Comments 13 pages, 7 figures. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01336 2026-05-05 cs.CL 70%

A Multi-View Media Profiling Suite: Resources, Evaluation, and Analysis

多视角媒体画像套件:资源、评估与分析

Muhammad Arslan Manzoor, Dilshod Azizov, Daniil Orel, Umer Siddique, Zain Muhammad Mujahid, Yufang Hou, Preslav Nakov

机构 * MBZUAI(穆巴扎人工智能研究院) Interdisciplinary Transformation University(跨学科转型大学) University of Texas at San Antonio, USA(德克萨斯大学圣安东尼奥分校) University of Copenhagen, Denmark(哥本哈根大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.CL

AI总结 本文提出MBFC-2025大规模标签集,构建多视角表示,系统评估嵌入视图与融合策略,并在ACL-2020和MBFC-2025上取得最优结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01147 2026-05-05 cs.AI 70%

Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment

位置:代理AI的安全性与公平性取决于交互拓扑,而非模型规模或对齐

Tanav Singh Bajaj, Nikhil Singh, Karan Anand, Eishkaran Singh

机构 * Department of Computer Science, University of British Columbia, Vancouver, Canada(英属哥伦比亚大学计算机科学系) Department of Artificial Intelligence, IIT Hyderabad, Hyderabad, India(印度海得拉巴理工学院人工智能系) Amazon, Delhi, India(印度德里亚马逊公司)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文指出代理AI的安全性由交互拓扑决定,而非模型规模或对齐程度。研究揭示了顺序不稳定、信息级联和功能崩溃等拓扑驱动的问题,并强调需通过动态系统视角评估安全性和公平性。

Comments 18 pages, 8 figures. Position paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01024 2026-05-05 cs.CV cs.AI 70%

EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness

EmoMM:针对冲突和缺失性下的多模态情绪识别的MLLM基准测试与引导

Yueru Sun, Yimeng Zhang, Haoyu Gu, Nuo Chen, Dong She, Xianrong Yao, Yang Gao, Zhanpeng Jin

机构 * School of Future Technology, South China University of Technology(华南理工大学未来技术学院)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出EmoMM基准,用于研究多模态情绪识别中模态冲突和缺失性下的MLLM决策机制,发现视频贡献崩溃现象,并提出CHASE机制减轻决策偏差,提升模型在复杂情感场景中的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00964 2026-05-05 cs.IR cs.AI cs.HC 70%

Seeking Information with RAG-Assistants: Does Model Size Matter in Human-AI Collaborations?

通过RAG助手获取信息:在人机协作中模型大小是否重要?

Lennard C. Froma, Tom Kouwenhoven, Maaike H. T. de Boer, Catholijn M. Jonker, Max J. van Duijn

机构 * Leiden University(莱顿大学) TNO(荷兰代尔夫特理工大学) TU Delft(代尔夫特理工大学)

专题命中 评测与基准 :LLM(abstract,abstract_cn);分类 cs.AI

AI总结 研究评估基于RAG的助手在真实多轮信息检索场景中的表现,探讨模型大小对人机协作动态及用户感知的影响,发现混合系统在信息检索中具有优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00876 2026-05-05 cs.LG cs.CV 70%

GAZE: Grounded Agentic Zero-shot Evaluation with Viewer-Level Tools and Literature Retrieval on Rare Brain MRI

GAZE:基于视图级工具和文献检索的 grounded 零样本评估

Duaa Alim, Mogtaba Alim, Liam Chalcroft

机构 * Imperial College London, UK(伦敦帝国学院) University of Toronto, Canada(多伦多大学) University College London, UK(伦敦大学学院)

专题命中 评测与基准 :language model(abstract);prompting(abstract);分类 cs.LG

AI总结 GAZE框架通过调用视图级工具和文献检索工具,使医学VLM在迭代方式下工作,实现对罕见脑MRI的零样本评估,提升病变定位和诊断准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00841 2026-05-05 cs.AI econ.GN q-fin.EC 70%

AI Agents for Sustainable SMEs: A Green ESG Assessment Framework

为可持续中小企业的AI代理:一个绿色ESG评估框架

Viet Trinh, Tan Nguyen, Minh-Huyen Phan, Quan Luu

机构 * European Union(欧洲联盟)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出一个基于AI的ESG评估框架,利用AI代理系统对欧洲中小企业的ESG表现进行自动化分类和推荐,验证了其与人工结果的一致性,支持绿色协议下的有效监控。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01496 2026-05-05 cs.CV 67%

SF20K Competition 2025: Summary and findings

SF20K竞赛2025:总结与发现

Ridouane Ghermi, Xi Wang, Vicky Kalogeiton, Ivan Laptev

机构 * SLoMO Workshop(SLoMO工作坊) Hugging Face

专题命中 评测与基准 :LLM(abstract,abstract_cn)

AI总结 本文总结了首个SF20K竞赛的结果,探讨了视频问答任务中多模态理解的重要性,并指出信息选择和推理结构是长视频问答的主要瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01371 2026-05-05 cs.RO 67%

ESARBench: A Benchmark for Agentic UAV Embodied Search and Rescue

ESARBench: 一种用于代理无人机体感知搜索与救援的基准

Daoxuan Zhang, Ping Chen, Jianyi Zhou, Shuo Yang

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract)

AI总结 本文提出ESARBench,首个用于评估基于MLLM的无人机在高真实感搜索与救援场景中的基准,通过构建高保真环境和600个任务数据集,揭示了ESAR中的关键瓶颈。

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01238 2026-05-05 cs.HC cs.CV 67%

EduGage: Methods and Dataset for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning

EduGage:基于传感器的自我引导视频学习中即时参与度评估的方法与数据集

Zikang Leng, Edan Eyal, Yingtian Shi, Jiaman He, Yaqi Liu, Thomas Plötz

机构 * School of Interactive Computing, Georgia Institute of Technology(佐治亚理工学院交互计算学院) College of Computing, Georgia Institute of Technology(佐治亚理工学院计算机学院) School of Computing Technologies, RMIT University(皇家墨尔本理工大学计算技术学院)

专题命中 评测与基准 :LLM(abstract,abstract_cn)

AI总结 本文提出EduGage方法与数据集,通过可穿戴和摄像头设备收集生理和运动信号,评估学习者参与度,实验显示多模态建模在参与度评估中有效,且轻量级的生理与行为信号组合优于全多模态设备。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01310 2026-05-05 cs.LG cs.AI 62%

GraphSculptor: Sculpting Pre-training Coreset for Graph Self-supervised Learning

GraphSculptor: 为图自监督学习构建预训练核心集

Chuang Liu, Zelin Yao, Xueqi Ma, Luzhi Wang, Mukun Chen, Pinghua Xu, Wenbin Hu

机构 * Sangfor Technologies Inc.(Sangfor 技术公司) Wuhan University(武汉大学) The University of Melbourne(墨尔本大学) Dalian Maritime University(大连海事大学) Hunan University of Technology and Business(湖南工业大学;湘江实验室) Xiangjiang Laboratory

专题命中 评测与基准 :language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出GraphSculptor,通过内在结构和语义视角构建图预训练核心集,利用内在统计和预训练语言模型实现去重,理论分析证明其有效性,实验显示10%核心集性能达99.6%且训练时间减少90%。

Comments 9 pages, 5 figures, Accepted by IJCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12539 2026-05-05 cs.AI cs.CL 62%

MemeLens: Multilingual Multitask VLMs for Memes

MemeLens:多语言多任务VLMs用于表情包

Ali Ezzat Shahroor, Mohamed Bayan Kmainasi, Abul Hasnat, Dimitar Dimitrov, Giovanni Da San Martino, Preslav Nakov, Firoj Alam

机构 * Qatar Computing Research Institute(卡塔尔计算研究所) Sofia University "St. Kliment Ohridski"(索菲亚大学"圣克莱孟·奥赫里迪斯") University of Padova(帕多瓦大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 评测与基准 :language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MemeLens,一种统一的多语言多任务解释增强视觉语言模型,用于表情包理解。通过整合38个公开数据集,建立20个任务的共享分类体系,并分析不同模型范式和数据集的表现,发现多模态训练和统一训练对表情包理解至关重要。

Comments disinformation, misinformation, factuality, harmfulness, fake news, propaganda, hateful meme, multimodality, text, images

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01650 2026-05-05 cs.LG 57%

Geospatial foundation-model embeddings improve population estimation unevenly across space and scale

地理基础模型嵌入物在空间和尺度上不均等地提高人口估计

Wenbin Zhang, Eimear Cleary, Francisco Rowe, Somnath Chaudhuri, Maksym Bondarenko, Shengjie Lai, Andrew J. Tatem

机构 * WorldPop, School of Geography and Environmental Sciences, University of Southampton, United Kingdom(世界人口研究机构,地理与环境科学学院,南安普顿大学,英国) Geographic Data Science Lab, Department of Geography and Planning, School of Environmental Sciences, University of Liverpool, United Kingdom(地理数据科学实验室,地理与规划系,环境科学学院,利物浦大学,英国)

专题命中 评测与基准 :foundation model(abstract);分类 cs.LG

AI总结 本文评估了地理基础模型嵌入物在巴西、尼日利亚和美国的子国家人口估计中的表现,发现其在数据贫乏区域提升预测效果,但空间尺度不匹配时性能下降,揭示了当前地理AI的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00893 2026-05-05 cs.CV cs.AI cs.IR 57%

Retrieval-Guided Generation for Safer Histopathology Image Captioning

基于检索的生成用于更安全的病理科图像描述生成

Md. Enamul Hoq, Wataru Uegami, Saghir Alfasly, Ghazal Alabtah, Sahar Rahimi Malakshan, Armita Kazemi, Alex T. Schmitgen, Fred Prior, H. R. Tizhoosh

机构 * Kimia Lab, Department of Artificial Intelligence \& Informatics, Mayo Clinic, Rochester, MN, USA Department of Biomedical Informatics, University of Arkansas for Medical Sciences, Little Rock, AR, USA Lane Department of Computer Science Electrical Engineering, West Virginia University, Morgantown, WV, USA Department of Computer Science Engineering, Princeton University, Princeton, NJ, USA Department of Computer Sciences, University of Wisconsin--Madison, Madison, WI, USA

专题命中 评测与基准 :language model(abstract);分类 cs.AI

AI总结 本文提出检索引导生成方法,通过总结相似病例的专家文本生成描述,提升病理图像描述的准确性与可靠性,实验表明其在语义对齐和诊断一致性方面优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27033 2026-05-05 cs.LG eess.SP 57%

Cross-Subject Generalization for EEG Decoding: A Survey of Deep Learning Methods

跨受体解码的通用性:深度学习方法的综述

Taida Li, Yujun Yan, Fei Dou, Wenzhan Song, Xiang Zhang

机构 * Department of Computer Science, University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校计算机科学系) Department of Computer Science, Dartmouth College(达特茅斯学院计算机科学系) School of Computing, University of Georgia(佐治亚大学计算科学学院) School of Electrical and Computer Engineering, University of Georgia(佐治亚大学电气与计算机工程学院)

专题命中 评测与基准 :foundation model(abstract);分类 cs.LG

AI总结 本文综述了深度学习方法在跨受体解码中的应用,分析了多源领域问题的评估协议,并系统分类了特征对齐、对抗学习等方法,探讨了理论限制和EEG基础模型的发展。

Comments Accepted manuscript in Progress in Biomedical Engineering. Minor update: corrected author affiliation in comment

详情

展开后加载摘要…

URL PDF HTML 收藏