arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9133 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9133 篇

2606.20143 2026-06-19 cs.CV 新提交 79%

HEad and neCK TumOR (HECKTOR) 2025: Benchmark of Segmentation, Diagnosis, and Prognosis in Multimodal PET/CT

头颈肿瘤 (HECKTOR) 2025 挑战赛:多模态 PET/CT 中的分割、诊断与预后基准

Numan Saeed, Salma Hassan, Shahad Hardan, Lishan Cai, Xinglong Liang, Moona Mazher, Abdul Qayyum, Yansong Bu, Mengye Lyu, Yue Lin, Mingyuan Meng, Chuanyi Huang, Lisheng Wang, Dalal Chamseddine, Shamimeh Ahrari, Beining Wu, Yifei Chen, Fuyou Mao, Hao Zhang, Baixiang Zhao, Surajit Ray, Muzi Guo, Lei Xiang, Jakob Dexl, Michael Ingrisch, Adrien Depeursinge, Arman Rahmim, Mathieu Hatt, Vincent Andrearczyk, Mohammad Yaqub

机构 * Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学) Amsterdam UMC(阿姆斯特丹大学医学中心) The Netherlands Cancer Institute(荷兰癌症研究所) Radboud University Medical Centre(拉德堡德大学医学中心) University College London(伦敦大学学院) Imperial College London(帝国理工学院) Shenzhen Technology University(深圳技术大学) Shenzhen University(深圳大学) Newland Digital Technology(新大陆数字技术) The University of Sydney(悉尼大学) Shanghai Jiao Tong University(上海交通大学) University Hospital, Nantes(南特大学医院) Nantes Université, Centrale Nantes, CNRS, LS2N(南特大学、南特中央理工学院、法国国家科学研究中心、LS2N实验室) Hangzhou Dianzi University(杭州电子科技大学) Tsinghua University(清华大学) Central South University(中南大学) University of Glasgow(格拉斯哥大学) China Mobile System Integration Co., Ltd.(中移系统集成有限公司) Subtle Medical Inc.(Subtle Medical公司) University Hospital, LMU Munich(慕尼黑大学医院) Munich Center for Machine Learning(慕尼黑机器学习中心) BC Cancer Research Institute(不列颠哥伦比亚癌症研究所) HES-SO Valais-Wallis University of Applied Sciences and Arts(HES-SO瓦莱州应用科学与艺术大学) Lausanne University Hospital (CHUV)(洛桑大学医院) LaTIM, INSERM, UMR 1101, Univ Brest(LaTIM实验室、法国国家健康与医学研究院、UMR 1101、布雷斯特大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 HECKTOR 2025 挑战赛利用多模态 PET/CT 和电子健康记录,建立了头颈癌自动分析的基准,涵盖肿瘤分割、复发预测和 HPV 分类三个任务,最佳算法分别达到 Dice 0.75、C-index 0.66 和平衡准确率 0.56。

Comments 17 pages, 4 figures, 4 tables. Overview paper for the HECKTOR 2025 challenge, held as a satellite event at MICCAI 2025. Challenge website: https://hecktor.grand-challenge.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28387 2026-06-19 cs.AI cs.LG 版本更新 79%

The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation

脚手架效应:提示框架如何驱动临床VLM评估中的表面多模态增益

Doan Nam Long Vu, Simone Balloccu

机构 * Technical University of Darmstadt(达姆施塔特技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 研究发现,在临床VLM评估中,提示中提及MRI可用性即可解释70-80%的性能提升,与图像数据是否存在无关,这种“脚手架效应”揭示了表面评估无法反映真实多模态推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05368 2026-06-18 cs.CV 版本更新 79%

Biomazon: A Multimodal Dataset for 3D Forest Structure and Biomass Modeling in the Amazon Basin

Biomazon:亚马逊盆地三维森林结构与生物量建模的多模态数据集

Sayan Mandal, Rocco Sedona, Simon Besnard, Mikhail Urbazaev, Morris Riedel, Ehsan Zandi, Gabriele Cavallaro

机构 * Jülich Supercomputing Centre (JSC), Forschungszentrum Jülich(julich超级计算中心(JSC),julich研究所) School of Engineering and Natural Sciences (SENS), University of Iceland(工程与自然科学学院(SENS),冰岛大学) Global Land Monitoring Group, GFZ Helmholtz Centre for Geosciences(全球土地监测组,geofz赫尔姆霍兹研究中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 针对现有方法未将森林垂直结构作为有序轮廓学习的问题,提出Biomazon多模态基准数据集,结合GEDI RH和AGBD目标与多传感器预测因子,通过共享编码器-解码器框架进行消融研究,为热带森林结构一致RH轮廓预测和结构-生物量建模建立参考基准。

Comments 32 pages, 21 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18109 2026-06-18 cs.CL cs.SD 版本更新 79%

FLiP: Towards understanding and interpreting multimodal multilingual sentence embeddings

FLiP:理解和解释多模态多语句子嵌入

Santosh Kesiraju, Bolaji Yusuf, Šimon Sedláček, Oldřich Plchot, Petr Schwarz

机构 * Brno University of Technology(布拉格技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 提出因子化线性投影(FLiP)模型,从多语言、多模态句子嵌入中恢复词汇内容,揭示编码器的模态和语言偏差。

Comments Accepted to Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17888 2026-06-17 cs.AI 新提交 79%

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning

MathVis-Fine:通过渐进式依赖引导训练将视觉监督与必要性对齐的多模态数学推理

Wanshi Xu, Haokun Zhao, Haidong Yuan, Songjun Cao, Long Ma

机构 * School of ECE, Peking University(北京大学电子与计算机工程学院) College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与技术学院) School of Software and Microelectronics, Peking University(北京大学软件与微电子学院) Tencent Youtu Lab(腾讯优图实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 提出MathVis-Fine框架,通过构建细粒度视觉标注数据集和两阶段渐进式训练,根据样本的视觉依赖程度平衡答案正确性和视觉基础奖励,提升多模态数学推理的监督精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17696 2026-06-17 cs.AI cs.GR 新提交 79%

FllumaOne: A Code-Native Multimodal CAD Dataset with Executable Programs and Kernel-Validated Feature Histories

FllumaOne:一个代码原生多模态CAD数据集,包含可执行程序与内核验证的特征历史

Jizong Zhan

机构 * Qt/C++ OpenCASCADE-based CAD system(基于Qt/C++ OpenCASCADE的CAD系统)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 提出FllumaOne数据集,通过可执行Python程序生成CAD模型,对齐程序、特征树、几何等模态,支持可编辑逆向工程等任务。

Comments 24 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17376 2026-06-17 cs.RO cs.CV 新提交 79%

Contactless Respiratory Monitoring on Heterogeneous Mobile Robots: A Multimodal Edge-Computing Framework

异构移动机器人上的非接触式呼吸监测:一种多模态边缘计算框架

Milind Rampure, Shadman Sakib, Haley Patel, Zahid Hasan, Nirmalya Roy

机构 * University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 提出一种适用于异构移动机器人的多模态非接触式呼吸率监测框架,通过自适应传感器选择、关键点引导的ROI提取和信号质量过滤,在多种平台和光照条件下实现鲁棒监测,无需平台特定调参。

Comments 8 pages, 6 figures. To appear in Proceedings of the 8th International Workshop on IoT Applications and Industry 5.0 (IoTI5 2026), co-located with IEEE DCOSS-IoT 2026, Reykjavik, Iceland, June 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17115 2026-06-17 cs.LG cs.AI q-bio.QM 新提交 79%

Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis

探测、融合与可信度:基础模型表示在多模态癌症分析中的系统评估

Jingyu Hu, Giuseppe Tripodi, Reed Naidoo, Sarah F. McGough, Tapabrata Chakraborti

机构 * The Alan Turing Institute(艾伦·图灵研究所) University of Bristol(布里斯托大学) University of Manchester(曼彻斯特大学) The Institute of Cancer Research(癌症研究所) Genentech(基因泰克)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 系统评估基础模型表示在计算病理学任务中的性能,发现图像和组学表示互补,多模态融合在单模态不占优时有效,并利用共形预测验证了不确定性感知推理的临床价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10384 2026-06-17 cs.CL 版本更新 79%

When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documents

当表格失控:评估多模态模型在法语金融文档上的表现

Virginie Mouilleron, Théo Lasnier, Anna Mosolova, Djamé Seddah

机构 * Inria Paris(巴黎国家信息与自动化研究所) Sorbonne Université(索邦大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 提出Scribe Finance基准,评估多模态模型在法语金融文档上的文本、表格、图表及多轮对话理解能力,发现模型在图表和多轮对话中表现脆弱。

Comments 16 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15890 2026-06-16 cs.AI 新提交 79%

UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics

UrbanWell: 面向时空城市福祉分析的多模态大语言模型基准测试

Yanxin Xi, Xiang Su, Jie Feng, Yu Liu, Sasu Tarkoma, Pan Hui

机构 * University of Helsinki(赫尔辛基大学) Zhongguancun Academy(中关村学院) University of Oxford(牛津大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 提出UrbanWell基准,通过卫星和街景图像联合建模,系统评估多模态大语言模型在环境、空间可达性、城市形态、活力和主观感知等5类城市福祉指标上的时空推理能力,并定义时序预测和趋势分类任务。

Comments accepted by KDD Datasets and Benchmarks Track 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15152 2026-06-16 cs.CL 新提交 79%

Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation

智能体能否读懂房间气氛?多模态模拟中的视觉社交智能基准测试

Shijun Wan, Xuehai Wu, Jiwen Zhang, Siyuan Wang, Zhongyu Wei

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 提出AgentViSS基准,通过240个场景和四项角色级任务评估多模态大模型的视觉社交智能,发现局部角色扮演接近饱和而交互调控仍困难。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18827 2026-06-16 q-bio.NC cs.AI 版本更新 79%

OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens

OmniMouse: 基于1500亿神经令牌的多模态多任务脑模型的可扩展性

Konstantin F. Willeke, Polina Turishcheva, Alex Gilbert, Goirik Chakrabarty, Hasan A. Bedel, Paul G. Fahey, Yongrong Qiu, Marissa A. Weis, Michaela Vystrčilová, Taliah Muhammad, Lydia Ntanavara, Rachel E. Froebe, Kayla Ponder, Zheng Huan Tan, Emin Orhan, Erick Cobos, Sophia Sanborn, Katrin Franke, Fabian H. Sinz, Alexander S. Ecker, Andreas S. Tolias

机构 * Department of Ophthalmology, Byers Eye Institute, Stanford University(斯坦福大学眼科学系、比尔斯眼科研究所) Stanford Bio-X, Stanford University(斯坦福大学生物交叉学科) Wu Tsai Neurosciences Institute, Stanford University(斯坦福大学吴泰教授神经科学研究所) Institute of Computer Science and Campus Institute Data Science, University Göttingen(哥廷根大学计算机科学研究所和校园数据科学研究所)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

AI总结 利用小鼠视觉皮层31亿神经元数据,训练多模态多任务模型OmniMouse,在神经预测、行为解码等任务上达到最优,发现性能随数据量可靠提升但模型规模收益饱和,与AI领域标准扩展规律相反。

Comments Published at ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22391 2026-06-16 cs.CL 版本更新 79%

Detecting Hate and Inflammatory Content in Bengali Memes: A New Multimodal Dataset and Co-Attention Framework

检测孟加拉语模因中的仇恨和煽动性内容:一个新的多模态数据集和共注意力框架

Rakib Ullah, Mominul islam, Md Sanjid Hossain, Md Ismail Hossain

机构 * University of California, Berkeley(加州大学伯克利分校) University of Washington(华盛顿大学)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CL

AI总结 针对孟加拉语模因中仇恨和煽动性内容检测的研究空白,构建了首个区分煽动性内容与直接仇恨言论的数据集Bn-HIB,并提出多模态共注意力融合模型MCFM,通过联合分析视觉和文本特征实现更准确分类。

Comments Added public link to dataset and fixed typo in abstract

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05888 2026-06-16 cs.CV 版本更新 79%

BioAutoML-NAS: An End-to-End AutoML Framework for Multimodal Insect Classification via Neural Architecture Search on Large-Scale Biodiversity Data

BioAutoML-NAS:基于大规模生物多样性数据通过神经架构搜索进行多模态昆虫分类的端到端AutoML框架

Arefin Ittesafun Abian, Debopom Sutradhar, Md Rafi Ur Rashid, Reem E. Mohamed, Md Rafiqul Islam, Asif Karim, Kheng Cher Yeo, Sami Azam

机构 * Department of Computer Science and Engineering, United International University, Dhaka, Bangladesh(乌姆国际大学计算机科学与工程系,达卡,孟加拉国) Applied Artificial Intelligence and Intelligent Systems (AAIINS) Laboratory, 1217, Dhaka, Bangladesh(应用人工智能与智能系统实验室(AAIINS),1217号,达卡,孟加拉国) Department of Computer Science and Engineering, Penn State University, University Park, PA, USA(宾夕法尼亚州立大学计算机科学与工程系,University Park,PA,美国) Faculty of Science and Information Technology, Charles Darwin University, Sydney, NSW, Australia(查尔斯达尔文大学科学与信息技术学院,悉尼,新南威尔士州,澳大利亚) Faculty of Science and Technology, Charles Darwin University, Casuarina, 0909, NT, Australia(查尔斯达尔文大学科学与技术学院,Casuarina,0909,北领地,澳大利亚)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 提出首个多模态BioAutoML模型BioAutoML-NAS,利用神经架构搜索自动学习图像操作,结合元数据融合与交替双层优化,在BIOSCAN-5M数据集上以96.81%准确率超越现有方法。

Comments Accepted in IEEE Transactions on Big Data

Journal ref IEEE Transactions on Big Data (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.21036 2026-06-16 cs.CL 版本更新 79%

GePBench: Evaluating Fundamental Geometric Perception for Multimodal Large Language Models

GePBench:评估多模态大语言模型的基础几何感知能力

Shangyu Xing, Changhao Xiang, Yuteng Han, Yifan Yue, Zhen Wu, Xinyu Liu, Zhangtai Wu, Fei Zhao, Xinyu Dai

机构 * University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 提出GePBench基准,系统评估多模态大语言模型的几何形状识别与空间关系感知能力,发现现有模型存在显著缺陷,而基于该基准训练可提升下游任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14562 2026-06-15 cs.CV cs.LG 新提交 79%

NEST3D: A High-Resolution Multimodal Dataset of Sociable Weaver Tree Nests

NEST3D:织布鸟树巢的高分辨率多模态数据集

Constanza A. Molina Catricheo, Simon Boeder, Ting-Jia Guo, Giacomo May, Clément Berthelot, Devis Tuia, Friedrich Fedor Reinhard, Fabio Remondino, Benjamin Risse

机构 * Institute for Geoinformatics (ifgi), University of Münster(明斯特大学地理信息学研究所) École Polytechnique Fédérale de Lausanne (EPFL)(洛桑联邦理工学院) Max Planck Institute of Animal Behavior(马克斯·普朗克动物行为研究所) University of Konstanz(康斯坦茨大学) Kuzikus Research Station(库兹库斯研究站) Fondazione Bruno Kessler (FBK)(布鲁诺·凯斯勒基金会)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 针对织布鸟巢缺乏精细3D结构数据的问题,提出包含104棵巢树、1.4TB多模态无人机数据集,并基准测试语义分割方法,PT-v3达86.35% mIoU。

Comments 14 pages, 4 figures. Dataset available at https://huggingface.co/NEST3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00336 2026-06-15 cs.CV 版本更新 79%

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection

MVAD:多模态AI生成视频-音频检测基准数据集

Mengxue Hu, Yunfeng Diao, Changtao Miao, Tairui Ge, Taize Ge, Zhiqing Guo, Jianshu Li, Zhe Li, Zhongjie Ba, Joey Tianyi Zhou

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 针对现有数据集缺乏多模态真实性的问题,提出MVAD数据集,包含三种伪造模式、高质量样本及多样化的视觉风格和内容类别,用于检测AI生成的视频-音频内容。

Comments 10 pages,2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12419 2026-06-12 cs.CY cs.AI 新提交 79%

GeoDial: A Multimodal Conversational Tutoring Dataset for Geometry Problem-Solving with Visual Tutor Turns

GeoDial:面向几何问题求解的多模态对话式辅导数据集,包含可视化辅导轮次

Sankalan Pal Chowdhury, Junling Wang, Donya Rooein, April Yi Wang, Mrinmaya Sachan

机构 * ETH Zurich(苏黎世联邦理工学院) ETH AI Center(苏黎世联邦理工学院人工智能中心) Bocconi University(博科尼大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 提出GeoDial数据集,包含1300+几何师生对话,通过可扩展标注协议整合对话行为、视觉高亮和反馈,微调视觉语言模型发现其难以生成准确图解高亮。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19652 2026-06-12 cs.CV 版本更新 79%

Navigating Gigapixel Pathology Images with Large Multimodal Models

利用大型多模态模型导航千兆像素病理图像

Thomas A. Buckley, Kian R. Weihrauch, Katherine Latham, Andrew Z. Zhou, Padmini A. Manrai, Arjun K. Manrai

机构 * Department of Biomedical Informatics, Harvard Medical School(哈佛医学院生物医学信息学系) Department of Pathology, Massachusetts General Hospital(麻省总医院病理学系) Department of Pathology and Laboratory Medicine, Brown University(布朗大学病理学与实验室医学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 提出GIANT方法,无需训练即可让通用多模态模型自主导航WSI,通过迭代选择多放大倍数裁剪并聚合证据,在MultiPathQA基准上实现SOTA。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11563 2026-06-11 cs.CV cs.RO 新提交 79%

Cross-Modal Benchmarking for Robotic Perception in Natural Environments

自然环境中机器人感知的跨模态基准测试

David Hall, Joshua Knights, Mark Cox, Peyman Moghadam

机构 * CSIRO Robotics, CSIRO, Australia(CSIRO机器人研究所,CSIRO,澳大利亚) University of Sydney (USyd), Australia(悉尼大学(USyd),澳大利亚) Queensland University of Technology (QUT), Australia(昆士兰理工大学(QUT),澳大利亚)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

AI总结 针对自然环境中机器人感知的挑战,提出WildCross跨模态基准,用于大规模自然场景下的地点识别和度量深度估计,并扩展了度量深度估计实验。

Comments Accepted to the IEEE ICRA Workshop on Open Challenges for Rigorous Robot Perception 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19502 2026-06-11 cs.AI cs.LG 版本更新 79%

Human-Guided Agentic AI for Multimodal Clinical Prediction: Lessons from the AgentDS Healthcare Benchmark

人类引导的智能体AI用于多模态临床预测:来自AgentDS医疗基准的教训

Lalitha Pranathi Pulavarthy, Raajitha Muthyala, Aravind V Kuruvikkattil, Zhenan Yin, Rashmita Kudamala, Saptarshi Purkayastha

机构 * University of California, Berkeley(加州大学伯克利分校) University of Washington(华盛顿大学) Stanford University(斯坦福大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 通过人类引导智能体AI在多模态临床预测任务中取得领先性能,提炼出领域知识引导特征工程、任务特定多模态融合和临床动机模型集成三大通用经验。

Comments Presented at the Data Challenge track at the 14th IEEE International Conference on Healthcare Informatics (ICHI) 2026 on June 3, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23508 2026-06-11 cs.CL 版本更新 79%

M4FC: a Multimodal, Multilingual, Multicultural, Multitask Real-World Fact-Checking Dataset

M4FC:一个多模态、多语言、多文化、多任务的真实世界事实验证数据集

Jiahui Geng, Jonathan Tonglet, Iryna Gurevych

机构 * Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学) Ubiquitous Knowledge Processing Lab(ubiquitous知识处理实验室) Department of Computer Science, TU Darmstadt(TU Darmstadt计算机科学系) National Research Center for Applied Cybersecurity ATHENE(应用网络安全国家研究中心ATHENE) Department of Electrical Engineering, KU Leuven(KU Leuven电气工程系) Department of Computer Science, KU Leuven(KU Leuven计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 为解决现有事实验证数据集规模小、语言单一、任务局限等问题,提出包含4982张图片和6980条声明的多模态数据集M4FC,覆盖6个验证任务,并提供基线结果。

Comments Preprint under review. Code and data available at: https://github.com/UKPLab/M4FC

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17012 2026-06-11 cs.AI cs.CE 版本更新 79%

Sustainability assessment using multimodal AI agents

使用多模态AI代理进行可持续性评估

Zhihan Zhang, Alexander Metzger, Yuxuan Mei, Felix Hähnlein, Zachary Englhardt, Tingyu Cheng, Gregory D. Abowd, Shwetak Patel, Adriana Schulz, Vikram Iyer

机构 * Paul G. Allen School of Computer Science & Engineering, University of Washington(保罗·G·艾伦计算机科学与工程学院,华盛顿大学) Computer Science and Engineering, University of Notre Dame(计算机科学与工程,诺丁汉大学) Electrical and Computer Engineering, Northeastern University(电气与计算机工程,东北大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 提出多模态多代理AI系统,模拟生命周期评估专家与利益相关者协作,自动估算电子设备碳足迹,将数据收集时间从数周缩短至一分钟,误差在19%以内。

Comments This article is published in Nature Electronics, and is available online at: https://www.nature.com/articles/s41928-026-01653-w

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10790 2026-06-10 cs.CV 新提交 79%

A Multimodal RGB and Events Dataset for Hand Detection in First-Person View

第一人称视角下用于手部检测的多模态RGB和事件数据集

Bharghav Kota, Yulia Sandamirskaya

机构 * Zurich University of Applied Sciences(苏黎世应用科技大学)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CV

AI总结 针对移动机器人系统中传统相机在暗光下运动模糊的问题,提出利用事件相机与RGB相机结合的多模态手部检测方法,并通过合成事件数据集实现与现有方法相当的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09882 2026-06-10 cs.CV cs.LG 新提交 79%

WHU-Infra3D: A Full-stack Multi-modal Dataset and Benchmark for 3D Roadside Infrastructure Inventory

WHU-Infra3D:面向3D路边基础设施清单的全栈多模态数据集与基准

Chong Liu, Luxuan Fu, Xuyu Feng, Zhen Dong, Bisheng Yang

机构 * State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing (LIESMARS)(信息工程测绘遥感国家重点实验室) Wuhan University(武汉大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

AI总结 提出WHU-Infra3D多模态基准数据集,覆盖三城市53.8公里,融合全景图像与LiDAR点云,提供2D-3D实例关联和跨帧跟踪,支持基础设施状态诊断与属性识别,填补自动化维护数据集空白。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09461 2026-06-09 cs.CL 新提交 79%

H2HMem: A Multimodal Memory Benchmark for Agents in Human-Human Interactions

H2HMem: 面向人际交互中智能体的多模态记忆基准

Shiping Zhu, Yibo Yang, Zhengyang Wang, Tiancheng Shen, Dandan Guo, Ming-Hsuan Yang

机构 * Jilin University(吉林大学) Shanghai Jiao Tong University(上海交通大学) University of California at Merced(加州大学默塞德分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 提出H2HMem基准,通过双人和多人多模态对话评估智能体在记忆召回、推理和应用方面的能力,揭示现有模型在多模态、多参与者场景下的显著局限。

Comments 22 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09125 2026-06-09 cs.CR cs.AI 新提交 79%

Unveiling Privacy Risks in Multi-modal Large Language Models: Task-specific Vulnerabilities and Mitigation Challenges

多模态大语言模型中的隐私风险揭示:任务特定漏洞与缓解挑战

Tiejin Chen, Pingzhi Li, Kaixiong Zhou, Tianlong Chen, Hua Wei

机构 * Arizona State University(亚利桑那州立大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) North Carolina State University(北卡罗来纳州立大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

AI总结 本研究揭示了多模态大语言模型在处理图像和文本时存在的隐私泄露风险,通过构建MM-Privacy数据集评估了不同任务下的披露风险与保留风险,并强调了任务不一致性对隐私风险的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08394 2026-06-09 cs.CL 新提交 79%

When Correct Decisions Hide Internal Stress: Decision-State Probing in Multimodal Language Models

当正确决策隐藏内部压力:多模态语言模型中的决策状态探测

Haoran Zhao, Soyeon Caren Han, Eduard Hovy

机构 * The University of Melbourne(墨尔本大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 提出S³E框架,通过正锚定A/B强制选择任务和隐藏状态分析,发现多模态语言模型在正确行为下仍存在语义压力导致的决策状态位移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07661 2026-06-09 cs.CV cs.DL 新提交 79%

PereStruct: Multimodal Semantic Assembly for Robust Historical Document Parsing

PereStruct: 面向鲁棒历史文档解析的多模态语义组装

Maksim Shandybo, Ivan Bespalov, Daniil Yefimov, Marina Kosheleva, Alexander Loukianov

机构 * IGIC RAS(俄罗斯科学院信息传输问题研究所) Yandex Cloud National University of Science and Technology MISIS(莫斯科国立钢铁合金学院) Nekrasov Central Universal Scientific Library(涅克拉索夫中央综合科学图书馆)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CV

AI总结 针对历史报纸复杂多栏布局的解析难题,提出结合微调YOLO与语义组装模块的多模态方法,在块到文章映射上F1达0.904,BLEU约0.96,显著优于通用视觉语言模型。

Comments Code and data available at https://github.com/makSShandybo/PereStruct

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06966 2026-06-08 cs.CV 新提交 79%

From Vision to Text: A Compact Multimodal Approach for Robust, Cross-Domain Presentation Attack Detection on ID Cards

从视觉到文本:一种用于身份证件跨域鲁棒演示攻击检测的紧凑多模态方法

Qingwen Zeng, Juan E. Tapia, Sneha Das, Christoph Busch

机构 * da/sec-Biometrics and Security Research Group, Hochschule Darmstadt(da/sec生物安全研究组,达姆施塔特应用技术大学) Technical University of Denmark (DTU)(丹麦技术大学(DTU))

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 针对身份证件演示攻击检测中的跨域迁移问题,提出一种结合视觉与文本数据的紧凑多模态模型,发现监督微调后泛化强但零样本设置下失效,强调模型容量和真实数据的重要性。

Comments Publication under the revision process on IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏