arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-06-11 至 2026-06-11 共收录 14 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 14 篇

2606.11563 2026-06-11 cs.CV cs.RO 新提交 79%

Cross-Modal Benchmarking for Robotic Perception in Natural Environments

自然环境中机器人感知的跨模态基准测试

David Hall, Joshua Knights, Mark Cox, Peyman Moghadam

机构 * CSIRO Robotics, CSIRO, Australia(CSIRO机器人研究所,CSIRO,澳大利亚) University of Sydney (USyd), Australia(悉尼大学(USyd),澳大利亚) Queensland University of Technology (QUT), Australia(昆士兰理工大学(QUT),澳大利亚)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

AI总结 针对自然环境中机器人感知的挑战,提出WildCross跨模态基准,用于大规模自然场景下的地点识别和度量深度估计,并扩展了度量深度估计实验。

Comments Accepted to the IEEE ICRA Workshop on Open Challenges for Rigorous Robot Perception 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19502 2026-06-11 cs.AI cs.LG 版本更新 79%

Human-Guided Agentic AI for Multimodal Clinical Prediction: Lessons from the AgentDS Healthcare Benchmark

人类引导的智能体AI用于多模态临床预测:来自AgentDS医疗基准的教训

Lalitha Pranathi Pulavarthy, Raajitha Muthyala, Aravind V Kuruvikkattil, Zhenan Yin, Rashmita Kudamala, Saptarshi Purkayastha

机构 * University of California, Berkeley(加州大学伯克利分校) University of Washington(华盛顿大学) Stanford University(斯坦福大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 通过人类引导智能体AI在多模态临床预测任务中取得领先性能,提炼出领域知识引导特征工程、任务特定多模态融合和临床动机模型集成三大通用经验。

Comments Presented at the Data Challenge track at the 14th IEEE International Conference on Healthcare Informatics (ICHI) 2026 on June 3, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23508 2026-06-11 cs.CL 版本更新 79%

M4FC: a Multimodal, Multilingual, Multicultural, Multitask Real-World Fact-Checking Dataset

M4FC:一个多模态、多语言、多文化、多任务的真实世界事实验证数据集

Jiahui Geng, Jonathan Tonglet, Iryna Gurevych

机构 * Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学) Ubiquitous Knowledge Processing Lab(ubiquitous知识处理实验室) Department of Computer Science, TU Darmstadt(TU Darmstadt计算机科学系) National Research Center for Applied Cybersecurity ATHENE(应用网络安全国家研究中心ATHENE) Department of Electrical Engineering, KU Leuven(KU Leuven电气工程系) Department of Computer Science, KU Leuven(KU Leuven计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 为解决现有事实验证数据集规模小、语言单一、任务局限等问题,提出包含4982张图片和6980条声明的多模态数据集M4FC,覆盖6个验证任务,并提供基线结果。

Comments Preprint under review. Code and data available at: https://github.com/UKPLab/M4FC

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17012 2026-06-11 cs.AI cs.CE 版本更新 79%

Sustainability assessment using multimodal AI agents

使用多模态AI代理进行可持续性评估

Zhihan Zhang, Alexander Metzger, Yuxuan Mei, Felix Hähnlein, Zachary Englhardt, Tingyu Cheng, Gregory D. Abowd, Shwetak Patel, Adriana Schulz, Vikram Iyer

机构 * Paul G. Allen School of Computer Science & Engineering, University of Washington(保罗·G·艾伦计算机科学与工程学院,华盛顿大学) Computer Science and Engineering, University of Notre Dame(计算机科学与工程,诺丁汉大学) Electrical and Computer Engineering, Northeastern University(电气与计算机工程,东北大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 提出多模态多代理AI系统,模拟生命周期评估专家与利益相关者协作,自动估算电子设备碳足迹,将数据收集时间从数周缩短至一分钟,误差在19%以内。

Comments This article is published in Nature Electronics, and is available online at: https://www.nature.com/articles/s41928-026-01653-w

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11683 2026-06-11 stat.AP 78%

MealMeter: Using Multimodal Sensing and Machine Learning for Automatically Estimating Nutrition Intake

MealMeter:利用多模态感知与机器学习自动估计营养摄入

Asiful Arefeen, Samantha Fessler, Sayyed Mostafa Mostafavi, Carol S Johnston, Hassan Ghasemzadeh

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 MealMeter通过整合多模态传感器数据与轻量级机器学习模型,实现高精度的营养摄入估计,其在碳水化合物的MAE和RMSRE分别达到13.2克和0.37,优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12106 2026-06-11 cs.CV cs.AI 新提交 76%

MSUE: Multi-Modal Soccer Understanding Expert

MSUE:多模态足球理解专家

Litao Li, Yibo Yu, Yufeng Hu, Zhuo Yang, Jiali Wen, Yixin Chen, Yixi Zhou

机构 * South China University of Technology(华南理工大学) Johns Hopkins University(约翰霍普金斯大学) Peking University(北京大学) University of Electronic Science and Technology of China(电子科技大学)

专题命中 多模态评测 :multi-modal(title);分类 cs.CV、cs.AI

AI总结 提出MSUE多专家问答架构,结合VLM数据合成管道与LLM动态调度文本、图像、视频专家,在SoccerNet VQA挑战中达到0.95准确率,获第三名。

Comments 6 pages, 1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10242 2026-06-11 cs.CV 版本更新 70%

MedVeriSeg: Teaching LISA-Like Medical Segmentation Models to Verify Query Validity Without Extra Training

MedVeriSeg: 教授LISA-like医学分割模型验证查询的有效性而无需额外训练

Qinyue Tong, Xiaozhen Wang, Ziqian Lu, Jun Liu, Yunlong Yu, Zheming Lu

机构 * School of Aeronautics and Astronautics, Zhejiang University(浙江大学航空宇航学院) Southern Medical University(南方医科大学) School of Computer Science and Technology (School of Artificial Intelligence), Zhejiang Sci-Tech University(浙江科技学院计算机科学与技术学院(人工智能学院)) College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院)

专题命中 多模态评测 :MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出MedVeriSeg,一种无需训练的查询验证框架,使LISA-like医学分割模型能够拒绝虚假分割查询。通过相似性响应质量评分模块和轻量级路由多代理验证模块,提升验证鲁棒性,并构建MedVeriSeg-Bench基准,有效减少幻觉分割。

Comments 13 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12169 2026-06-11 cs.CV cs.AI cs.CL cs.LG 新提交 67%

OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

OpenMedReason: 医学视觉语言模型的科学推理监督

Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi

机构 * York University(约克大学) Vector Institute(向量研究所) University of British Columbia(不列颠哥伦比亚大学) University of Toronto(多伦多大学) Unity Health Toronto / St. Michael’s Hospital(多伦多联合健康/圣迈克尔医院) University Health Network(大学健康网络) Arc Institute(弧研究所) Queen's University(女王大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出OpenMedReason,一个包含约45万图像-问题-答案实例的大规模开放医学推理语料库,其推理轨迹主要来自生物医学科学文章,并配套基准OpenMedReason-Bench进行细粒度评估,在监督微调和强化对齐中有效提升模型性能。

Comments 42 pages, 9 figures, 24 tables. Dataset and code: https://huggingface.co/datasets/neginb/OpenMedReason

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11702 2026-06-11 cs.CV cs.AI cs.CL 新提交 67%

MedCTA: A Benchmark for Clinical Tool Agents

MedCTA: 临床工具智能体基准

Tajamul Ashraf, Hyewon Jeong, Fida Mohammad Thoker, Bernard Ghanem

机构 * King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学) Massachusetts Institute of Technology (MIT)(麻省理工学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出MedCTA基准,基于放射影像、病理切片和报告等真实临床多模态输入,评估医疗AI智能体在工具检索、证据获取和集成方面的规划与执行能力。

Comments Project Page: https://ivul-kaust.github.io/MedCTA/ Code: https://github.com/IVUL-KAUST/MedCTA Data: https://huggingface.co/datasets/IVUL-KAUST/MedCTA

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22725 2026-06-11 cs.CV cs.AI 版本更新 62%

OpenVTON-Bench: A Large-Scale High-Resolution Benchmark for Controllable Virtual Try-On Evaluation

OpenVTON-Bench:用于可控虚拟试穿评估的大规模高分辨率基准

Jin Li, Tao Chen, Kai Wen, Siqi Yin, Shuai Jiang, Weijie Wang, Jingwen Luo, Chenhui Wu

机构 * Renxing Intelligence, Hangzhou, China Hangzhou Dianzi University, Hangzhou, China(杭州电子科技大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 提出OpenVTON-Bench,包含约10万对高分辨率图像,通过DINOv3聚类和Gemini描述构建,并设计多模态评估协议,沿五个维度衡量试穿质量,与人类判断高度一致。

Comments Under review for the NeurIPS 2026 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12207 2026-06-11 cs.RO cs.AI 新提交 57%

Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simulators, and Trends

具身基准构建的智能自动化:流程、具身、模拟器与趋势

Jinshan Lai, Jianwei Hu, Baoyang Jiang, Fengchun Zhang, Leyuan Wang, Haotian Li, Yida Wang, Tingxuan Huang, Xi Ren, Qiang Ma

机构 * University of Electronic Science and Technology of China(电子科技大学) Qiyuan Lab(启元实验室) Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学) Beihang University(北京航空航天大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 本文综述具身智能基准构建的五阶段流程,分析从人工到自动化再到智能体闭环的转变,指出自动化将成本转向验证与治理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11573 2026-06-11 cs.CV 新提交 57%

Understanding Cross-Sensor Feature Variations for Generalizable 3D Perception

理解跨传感器特征变化以实现可泛化的3D感知

Xin Qiu, Wenjie Liu, Fuyuan Ai, YuChen Tan, Zhiwei Xu, Chunyi Song

机构 * Zhejiang University(浙江大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

AI总结 针对雷达-相机BEV感知跨数据集性能下降问题,提出频域场景变化建模框架,通过合成多样源域视图并正则化融合表示,提升3D检测器鲁棒性,无需目标域样本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11260 2026-06-11 cs.SD cs.AI 新提交 57%

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark

RAIL: 基于CHC框架重新思考大型音频语言模型中的听觉智能

Hongyu Jin, Siyi Wang, Yang Xiao, Jiaheng Dong, Shihong Tan, Kaiyuan peng, Georgiana Juravle, Shanquan Chen, Gongping Huang, Hong Jia, Eun-Jung Holden, James Bailey, Ting Dang

机构 * School of Computing and Information Systems, The University of Melbourne(墨尔本大学计算与信息系统学院) Faculty of Psychology and Educational Sciences, Alexandru Ioan Cuza University of Iași(亚历山德鲁伊万库扎大学心理学与教育科学学院) School of Electronic Information, Wuhan University(武汉大学电子信息学院) School of Public Health, The University of Hong Kong(香港大学公共卫生学院) School of Computer Science, The University of Auckland(奥克兰大学计算机科学学院) Department of Data Science and Artificial Intelligence, Monash University(莫纳什大学数据科学与人工智能系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 提出RAIL基准,基于CHC认知框架将听觉智能分解为五种核心能力,构建结构化评估任务,系统评测大型音频语言模型的认知行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11204 2026-06-11 cs.CL cs.IR 新提交 57%

Benchmarking Large Language Models for Safety Data Extraction

大型语言模型在安全数据提取中的基准测试

Jonas Grill, Thomas Bayer, Sören Berlinger

机构 * SAP SE(SAP公司) Institute for Digital Transformation, Ravensburg-Weingarten University(拉文斯堡-魏恩加滕大学数字化转型研究所)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

AI总结 针对安全数据表(SDS)的异构格式,本研究基准测试了四种大型语言模型(LLM)在文本与多模态处理下的提取性能,发现文本结合思维链提示的Gemini 1.5 Pro准确率最高(84%),但均未达到90%的可靠部署阈值。

Comments 18 pages, 8 figures, submitted to Applied Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏