arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9087 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9087 篇

2512.20548 2025-12-29 cs.AI 83%

Advancing Multimodal Teacher Sentiment Analysis:The Large-Scale T-MED Dataset & The Effective AAM-TSA Model

推进多模态教师情感分析:大规模T-MED数据集与有效的AAM-TSA模型

Zhiyi Duan, Xiangren Wang, Hongyu Yuan, Qianli Xing

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出T-MED数据集和AAM-TSA模型,通过多模态融合提升教师情感分析的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21616 2025-12-29 cs.CV 83%

TAMEing Long Contexts in Personalization: Towards Training-Free and State-Aware MLLM Personalized Assistant

在个性化中延长上下文:迈向无训练和状态感知的MLLM个性化助手

Rongpei Hong, Jian Lang, Ting Zhong, Yong Wang, Fan Zhou

机构 * University of Electronic Science and Technology of China(电子科技大学) Aiwen Technology Co., Ltd.(Aiwen科技有限公司) Intelligent Digital Media Technology Key Laboratory of Sichuan Province(四川省智能数字媒体技术重点实验室)

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 本文提出TAME框架,通过双记忆和RA2G范式实现无训练、状态感知的MLLM个性化,提升长上下文对话能力。

Comments Accepted by KDD 2026 research track. Code and data are available at https://github.com/ronpay/TAME

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21360 2025-12-29 cs.AI cs.SY eess.SY 83%

From Visual Perception to Deep Empathy: An Automated Assessment Framework for House-Tree-Person Drawings Using Multimodal LLMs and Multi-Agent Collaboration

从视觉感知到深度共情:一种利用多模态大语言模型和多智能体协作的房屋-树-人绘画自动评估框架

Shuide Wen, Yu Sun, Beier Ku, Zhi Gao, Lijun Ma, Yang Yang, Can Jiao

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 本文提出利用多模态大语言模型和多智能体协作,开发自动评估房屋-树-人绘画测试的框架,以提升投射评估的标准化和效率。

Comments 16 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09603 2025-12-18 cs.CV 83%

Do MLLMs Exhibit Human-like Perceptual Behaviors? HVSBench: A Benchmark for MLLM Alignment with Human Perceptual Behavior

多模态大语言模型是否表现出类人感知行为?HVSBench:一个用于多模态大语言模型对齐人类视觉行为的基准测试

Jiaying Lin, Shuquan Ye, Dan Xu, Wanli Ouyang, Rynson W. H. Lau

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 HVSBench通过测试多模态大语言模型与人类视觉行为的对齐性,揭示了其在感知任务上的显著差距,并强调了开发更符合人类认知的AI的重要性。

Comments Project page: https://jiaying.link/HVSBench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14594 2025-12-17 cs.CV 83%

LLM-driven Knowledge Enhancement for Multimodal Cancer Survival Prediction

基于大语言模型的知识增强的多模态癌症生存预测

Chenyu Zhao, Yingxue Xu, Fengtao Zhou, Yihui Wang, Hao Chen

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系) Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology(香港科技大学化学与生物工程系) Division of Life Science, The Hong Kong University of Science and Technology(香港科技大学生命科学系) HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, The Hong Kong University of Science and Technology(香港科技大学深圳-香港协同创新研究院) State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology(香港科技大学神经系统疾病国家重点实验室)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出KEMM模型,通过整合专家报告和预后背景知识,利用知识增强的跨模态注意力模块提升癌症生存预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10576 2025-12-16 cs.CV 83%

HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs

HumanSense: 从多模态感知到通过推理的 empathetic 上下文感知响应

Zheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li, Yi Yuan, Jingdong Chen, Le Wang

专题命中 多模态评测 :multimodal(title,abstract);omni-modal(abstract);分类 cs.CV

AI总结 HumanSense通过多阶段模态渐进强化学习提升MLLMs在多模态感知和交互任务中的性能,强调推理能力对共情反馈的重要性。

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09872 2025-12-11 cs.CR cs.AI 83%

FlipLLM: Efficient Bit-Flip Attacks on Multimodal LLMs using Reinforcement Learning

FlipLLM: 通过强化学习高效攻击多模态大语言模型的位翻转攻击

Khurram Khalil, Khaza Anuarul Hoque

机构 * Department of Electrical Engineering and Computer Science(电气工程与计算机科学系) University of Missouri-Columbia(密苏里大学哥伦比亚分校)

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 FlipLLM通过强化学习高效识别多模态大语言模型的位翻转攻击漏洞,显著提升攻击检测速度和防御指导价值。

Comments Accepted in IEEE HOST 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02022 2025-12-11 cs.CV 83%

Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs

你看见我:一个多维基准用于评估多模态大语言模型的视觉感知

Aditya Kanade, Tanuja Ganu

机构 * Microsoft Research India(微软印度研究院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 该研究提出‘你看见我’基准,评估多模态大语言模型的视觉感知能力,发现其在复杂任务中表现远低于人类。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17397 2025-12-09 cs.CV 83%

MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality Assessment

MCMoE:基于专家混合的缺失模态补全用于不完整多模态动作质量评估

Huangbiao Xu, Huanqi Wu, Xiao Ke, Junyi Wu, Rui Xu, Jinglin Xu

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 MCMoE通过专家混合方法解决多模态动作质量评估中缺失模态的问题,实现单阶段训练下的多模态学习和生成。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03783 2025-12-05 cs.AI cs.SD 83%

Omni-AutoThink: Adaptive Multimodal Reasoning via Reinforcement Learning

Omni-AutoThink: 通过强化学习实现自适应多模态推理

Dongchao Yang, Songxiang Liu, Disong Wang, Yuanyuan Wang, Guanglu Wan, Helen Meng

机构 * The Chinese University of Hong Kong(香港中文大学) Meituan(美团)

专题命中 多模态评测 :multimodal(title,abstract);audio-visual(abstract);分类 cs.AI

AI总结 Omni-AutoThink通过强化学习实现自适应多模态推理,提升模型在不同任务难度下的推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02668 2025-12-03 cs.CV 83%

UAUTrack: Towards Unified Multimodal Anti-UAV Visual Tracking

UAUTrack:迈向统一的多模态反无人机视觉跟踪

Qionglin Ren, Dawei Zhang, Chunxu Tian, Dan Zhang

机构 * College of Intelligent Robotics and Advanced Manufacturing(智能机器人与先进制造学院) Fudan University(复旦大学) School of Computer Science and Technology(计算机科学与技术学院) Zhejiang Normal University(浙江师范大学) Department of Mechanical Engineering(机械工程系) The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 UAUTrack通过统一多模态框架实现反无人机跟踪,结合文本先验提示策略提升跨模态信息整合效率,取得优异性能与效率平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22404 2025-12-01 cs.CV 83%

UAV-MM3D: A Large-Scale Synthetic Benchmark for 3D Perception of Unmanned Aerial Vehicles with Multi-Modal Data

UAV-MM3D: 一种大规模合成基准,用于多模态数据下的无人机三维感知

Longkun Zou, Jiale Wang, Rongqin Liang, Hai Wu, Ke Chen, Yaowei Wang

机构 * Pengcheng Laboratory(鹏城实验室) University of Southern California(南加州大学)

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 UAV-MM3D通过多模态合成数据提升无人机三维感知能力,提供高保真数据集和多任务基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07984 2025-12-01 cs.CV 83%

SAMChat: Introducing Chain of Thought Reasoning and GRPO to a Multimodal Small Language Model for Small Scale Remote Sensing

SAMChat:引入链式推理和GRPO以增强小规模遥感遥感小语言模型

Aybora Koksal, A. Aydin Alatan

机构 * Center for the Image Analysis (OGAM) and Department of Electrical and Electronics Engineering of Middle East Technical University (METU)(图像分析中心(OGAM)和中东部技术大学(METU)电子与电气工程系)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 SAMChat通过引入链式推理和GRPO,专为遥感影像分析优化,实现了在开放描述和分类任务上的高精度表现。

Comments Accepted to Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS) Special Issue on Foundation and Large Vision Models for Remote Sensing. Code and dataset are available at https://github.com/aybora/SAMChat

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20490 2025-11-26 cs.LG cs.AI 83%

MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology

MTBBench: 一个用于肿瘤学的多模态序列临床决策制定基准

Kiril Vasilev, Alexandre Misrahi, Eeshaan Jain, Phil F Cheng, Petros Liakopoulos, Olivier Michielin, Michael Moor, Charlotte Bunne

机构 * ETH Zürich(苏黎世联邦理工学院) EPFL(洛桑联邦理工学院) HUG

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

AI总结 MTBBench通过模拟分子肿瘤委员会的决策流程,提出一个多模态、纵向的肿瘤学基准测试,评估LLMs在复杂临床任务中的表现并提升其推理与可靠性。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15165 2025-11-25 cs.CR cs.AI 83%

Can MLLMs Detect Phishing? A Comprehensive Security Benchmark Suite Focusing on Dynamic Threats and Multimodal Evaluation in Academic Environments

MLLMs能否检测钓鱼?一个全面的安全基准套件,聚焦于动态威胁和学术环境中的多模态评估

Jingzhuo Zhou

机构 * School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 本文提出AdapT-Bench,一个针对学术环境动态钓鱼攻击的多模态安全基准套件,旨在评估MLLM的防御能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14439 2025-11-20 cs.CL 83%

MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents

Jinru Ding, Lu Lu, Chao Ding, Mouxiao Bian, Jiayuan Chen, Wenrao Pang, Ruiyao Chen, Xinwei Peng, Renjie Lu, Sijie Ren, Guanxu Zhu, Xiaoqin Wu, Zhiqiang Liu, Rongzhao Zhang, Luyi Jiang, Bing Han, Yunqiu Wang, Jie Xu

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13242 2025-11-18 cs.CV 83%

MMD-Thinker: Adaptive Multi-Dimensional Thinking for Multimodal Misinformation Detection

Junjie Wu, Guohong Fu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院)

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05982 2025-11-18 cs.CV 83%

MCA-Bench: A Multimodal Benchmark for Evaluating CAPTCHA Robustness Against VLM-based Attacks

Zonglin Wu, Yule Xue, Yaoyao Feng, Xiaolong Wang, Yiren Song

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments we update the paper supplement

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09415 2025-11-18 cs.CV 83%

FaceShield: Explainable Face Anti-Spoofing with Multimodal Large Language Models

Hongyang Wang, Yichen Shi, Zhuofu Tao, Yuhao Gao, Liepiao Zhang, Xun Lin, Jun Feng, Xiaochen Yuan, Zitong Yu, Xiaochun Cao

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by AAAI 2025. Hongyang Wang and Yichen Shi contribute equally. Corresponding author: Zitong Yu

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10075 2025-11-14 cs.CL 83%

Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts

Xanh Ho, Yun-Ang Wu, Sunisth Kumar, Florian Boudin, Atsuhiro Takasu, Akiko Aizawa

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21572 2025-11-14 cs.CL 83%

Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling

Shengwu. Xiong, Tianyu. Zou, Cong. Wang, Xuelong Li

机构 * Interdisciplinary Artificial Intelligence Research Institute, Wuhan College(交叉学科人工智能研究 institute,武汉学院) School of Computer and Artificial Intelligence, Wuhan University of Technology(计算机与人工智能学院,武汉理工大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Sanya Science and Education Innovation Park, Wuhan University of Technology(三亚科学教育创新园,武汉理工大学) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) School of Mathematics and Statistics, Northwestern Polytechnical University(数学与统计学院,西北工业大学) Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究所(TeleAI))

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.CL

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10975 2025-11-13 cs.IR cs.CL 83%

ReFineG: Synergizing Small Supervised Models and LLMs for Low-Resource Grounded Multimodal NER

Jielong Tang, Shuang Wang, Zhenxing Wang, Jianxing Yu, Jian Yin

机构 * School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院) Key Laboratory of Sustainable Tourism Smart Assessment Technology, Ministry of Culture and Tourism, Sun Yat-sen University(文化旅游可持续评估技术重点实验室,中华人民共和国文化和旅游部,中山大学) Beijing Normal University(北京师范大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments CCKS 2025 Shared Task Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07812 2025-11-12 cs.CV 83%

Revisiting MLLM Based Image Quality Assessment: Errors and Remedy

Zhenchen Tang, Songlin Yang, Bo Peng, Zichuan Wang, Jing Dong

专题命中 多模态评测 :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19875 2025-11-11 cs.CV 83%

InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows

Kirolos Ataallah, Eslam Abdelrahman, Mahmoud Ahmed, Chenhui Gou, Khushbu Pahwa, Jian Ding, Mohamed Elhoseiny

机构 * KAUST(卡塔尔科技大学) Monash University(墨尔本大学) RICE University(里士满大学)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted for oral presentation at the EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03845 2025-11-07 cs.AI cs.LG 83%

To See or To Read: User Behavior Reasoning in Multimodal LLMs

Tianning Dong, Luyi Ma, Varun Vasudevan, Jason Cho, Sushant Kumar, Kannan Achan

机构 * Personalization Team, Walmart Global Tech(Walmart全球科技个性化团队)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Accepted by the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Efficient Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04458 2025-10-30 cs.CL 83%

Think Twice Before You Judge: Mixture of Dual Reasoning Experts for Multimodal Sarcasm Detection

Soumyadeep Jana, Abhrajyoti Kundu, Sanasam Ranbir Singh

机构 * Indian Institute of Technology Guwahati(印度理工学院古瓦哈提)

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12841 2025-10-29 cs.CV 83%

AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning

Yiming Ren, Zhiqiang Lin, Yu Li, Gao Meng, Weiyun Wang, Junjie Wang, Zicheng Lin, Jifeng Dai, Yujiu Yang, Wenhai Wang, Ruihang Chu

机构 * Tsinghua University(清华大学) Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态评测 :omni-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22443 2025-10-28 cs.CV cs.LG 83%

Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable Agents

Vijay Veerabadran, Fanyi Xiao, Nitin Kamra, Pedro Matias, Joy Chen, Caley Drooff, Brett D Roads, Riley Williams, Ethan Henderson, Xuanyi Zhao, Kevin Carlberg, Joseph Tighe, Karl Ridgeway

机构 * Meta Reality Labs(Meta 现实实验室) Meta FAIR

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted as a spotlight paper at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12520 2025-10-28 cs.CV 83%

SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine Unlearning

Junkai Chen, Zhijie Deng, Kening Zheng, Yibo Yan, Shuliang Liu, PeiJun Wu, Peijie Jiang, Jia Liu, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Southeast University(东南大学) Ant Group, Alibaba(蚂蚁集团)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21603 2025-10-27 cs.IR cs.CL 83%

Doc-Researcher: A Unified System for Multimodal Document Parsing and Deep Research

Kuicai Dong, Shurui Huang, Fangda Ye, Wei Han, Zhi Zhang, Dexun Li, Wenjun Li, Qu Yang, Gang Wang, Yichao Wang, Chen Zhang, Yong Liu

机构 * Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏