arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4868 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4868 篇

2111.02922 2022-07-07 cs.LG q-bio.QM stat.ML 82%

Reconstructing Nonlinear Dynamical Systems from Multi-Modal Time Series

Daniel Kramer, Philine Lou Bommer, Carlo Tombolini, Georgia Koppe, Daniel Durstewitz

专题命中 其他多模态 :multi-modal(title,abstract);cross-modal(abstract)

Comments 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.06299 2022-04-14 cs.CL cs.AI cs.CV 82%

TIB-VA at SemEval-2022 Task 5: A Multimodal Architecture for the Detection and Classification of Misogynous Memes

Sherzod Hakimov, Gullal S. Cheema, Ralph Ewerth

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted for publication at SemEval-2022 Workshop, Task 5: MAMI - Multimedia Automatic Misogyny Identification co-located with NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.12033 2022-02-25 cs.RO 82%

Multi-Modal Legged Locomotion Framework with Automated Residual Reinforcement Learning

Chen Yu, Andre Rosendo

专题命中 其他多模态 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.04692 2021-12-23 cs.CL cs.CV cs.MM 82%

Semiotically-grounded distant viewing of diagrams: insights from two multimodal corpora

Tuomo Hiippala, John A. Bateman

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments 22 pages, 11 figures. Under review at Digital Scholarship in the Humanities

Journal ref Digital Scholarship in the Humanities, 2021 (ahead of press)

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.08717 2021-05-27 cs.CV cs.AI cs.CL 82%

Giving Commands to a Self-driving Car: A Multimodal Reasoner for Visual Grounding

Thierry Deruyttere, Guillem Collell, Marie-Francine Moens

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Updated acknowledgements

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.11560 2021-04-26 cs.CL cs.CV cs.MM 82%

Weakly-supervised Multi-task Learning for Multimodal Affect Recognition

Wenliang Dai, Samuel Cahyawijaya, Yejin Bang, Pascale Fung

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments 13 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.04965 2021-04-06 cs.CV cs.AI cs.CL cs.LG 82%

Visual Relationship Detection with Visual-Linguistic Knowledge from Multimodal Representations

Meng-Jiun Chiou, Roger Zimmermann, Jiashi Feng

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published in IEEE Access

Journal ref IEEE Access, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.05208 2021-01-14 cs.CV cs.AI cs.CL 82%

Efficient Object-Level Visual Context Modeling for Multimodal Machine Translation: Masking Irrelevant Objects Helps Grounding

Dexin Wang, Deyi Xiong

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.05669 2020-06-11 eess.SY cs.LG cs.SY 82%

Interpretable Multimodal Learning for Intelligent Regulation in Online Payment Systems

Shuoyao Wang, Diwei Zhu

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract)

Comments Accepted by IJCAI 2020. SOLE copyright holder is IJCAI (international Joint Conferences on Artificial Intelligence)

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.03848 2019-08-13 cs.LG stat.ML 82%

Deep Structured Cross-Modal Anomaly Detection

Yuening Li, Ninghao Liu, Jundong Li, Mengnan Du, Xia Hu

专题命中 其他多模态 :cross-modal(title,abstract);multi-modal(abstract)

Comments 8 pages, in Proceedings of the 2019 International Joint Conference on Neural Networks (IJCNN)

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.11395 2019-05-29 cs.LG stat.ML 82%

Multi-Modal Graph Interaction for Multi-Graph Convolution Network in Urban Spatiotemporal Forecasting

Xu Geng, Xiyu Wu, Lingyu Zhang, Qiang Yang, Yan Liu, Jieping Ye

专题命中 其他多模态 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.01070 2019-04-03 cs.LG cs.NE q-bio.NC stat.ML 82%

Multimodal Sparse Classifier for Adolescent Brain Age Prediction

Peyman Hosseinzadeh Kassani, Alexej Gossmann, Yu-Ping Wang

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.01385 2018-06-12 cs.NE 82%

Research on the Brain-inspired Cross-modal Neural Cognitive Computing Framework

Yang Liu

专题命中 其他多模态 :cross-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1609.03330 2016-09-14 physics.med-ph 82%

Hearables: Multimodal physiological in-ear sensing

Valentin Goverdovsky, Wilhelm von Rosenberg, Takashi Nakamura, David Looney, David J Sharp, Christos Papavassiliou, Mary J Morrell, Danilo P Mandic

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15736 2024-04-26 cs.CV cs.AI 82%

What Makes Multimodal In-Context Learning Work?

Folco Bertini Baldassini, Mustafa Shukor, Matthieu Cord, Laure Soulier, Benjamin Piwowarski

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 20 pages, 16 figures. Accepted to CVPR 2024 Workshop on Prompting in Vision. Project page: https://folbaeni.gitlab.io/multimodal-icl

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.03280 2022-11-08 cs.CV cs.AI 82%

Multimodal Learning for Non-small Cell Lung Cancer Prognosis

Yujiao Wu, Yaxiong Wang, Xiaoshui Huang, Fan Yang, Sai Ho Ling, Steven Weidong Su

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 11 pages, 6 figures, Multimodal learning, NSCLC, Survival analysis, Transformer

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17343 2026-07-14 cs.CV 版本更新 81%

EvoGuard: An Extensible Agentic RL-based Framework for Practical and Evolving AI-Generated Image Detection

EvoGuard: 一种基于代理强化学习的可扩展框架,用于实际和演化的AI生成图像检测

Chenyang Zhu, Maorong Wang, Jun Liu, Ching-Chun Chang, Isao Echizen

机构 * The University of Tokyo(东京大学) National Institute of Informatics(国家信息研究所)

专题命中 其他多模态 :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 本文提出EvoGuard框架,利用多模态大语言模型和非MLLM检测器,通过能力感知的动态编排机制实现高效AIGI检测,优化后无需细粒度标注,实验表明其在准确性和可扩展性上均达最优。

Comments Template changed

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09741 2026-04-21 cs.CV cs.LG 81%

Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping

构造性失真:通过注意力引导的图像扭曲改进MLLMs

Dwip Dalal, Gautam Vashishtha, Utkarsh Mishra, Jeonghwan Kim, Madhav Kanda, Hyeonjeong Ha, Svetlana Lazebnik, Heng Ji, Unnat Jain

机构 * University of Illinois Urbana–Champaign(伊利诺伊大学厄巴纳-香槟分校) Skan AI Texas A&M University(德克萨斯大学阿姆斯特朗分校) University of California, Irvine(加州大学 Irvine 分校)

专题命中 其他多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出AttWarp方法,通过注意力引导的图像扭曲增强MLLMs对细节和空间关系的感知能力,提升准确性和推理能力。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10140 2024-05-17 cs.CV 81%

Libra: Building Decoupled Vision System on Large Language Models

Yifan Xu, Xiaoshan Yang, Yaguang Song, Changsheng Xu

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);multimodal foundation model(abstract)

Comments ICML2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10154 2026-08-12 cs.CL cs.AI 新提交 81%

Multimodal Item Parameter Estimation using Simulated Response Probabilitie

基于模拟响应概率的多模态题目参数估计

Christopher Ormerod, YoungKoung Kim

机构 * College Board(大学理事会)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本研究基于Qwen3.5微调多模态大语言模型,利用含图像与文本的多项选择题训练语料,通过学习学生能力相关的选择概率,实现了题目参数的准确估计。

Comments Submitted and Accepted for AIME-Con 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29064 2026-08-11 cs.CL cs.CV cs.HC cs.MA 版本更新 81%

Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation

分析多模态大语言模型代理在城市感知中生成解释的角色效应

Neemias da Silva, Matt Ratto, Myriam Delgado, Rodrigo Minetto, Daniel Silver, Thiago H Silva

机构 * Universidade Tecnologica Federal do Parana(巴西南里奥格兰德联邦技术大学) University of Toronto(多伦多大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 通过对比不同角色提示和无角色设置下多模态大语言模型生成的文本,发现标题描述趋同,但理由描述随社会经济和政治属性系统变化,感知标签无显著差异。

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04348 2026-08-06 cs.CV cs.AI cs.LG stat.ML 新提交 81%

iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data

iStructTab:面向图像与表格数据多模态学习的结构化特征排序

Al Zadid Sultan Bin Habib, Md Younus Ahamed, Prashnna Gyawali, Gianfranco Doretto, Donald A. Adjeroh

机构 * West Virginia University(西弗吉尼亚大学) The University of Utah(犹他大学) Scientific Computing and Imaging Institute(科学计算与成像研究所)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 iStructTab基于图增强描述子排序算法(GEDS),将其融入感知顺序的高效Transformer框架,在多模态基准上有效减少特征分散,提升了图像与表格数据多模态学习的预测性能及鲁棒性。

Comments This paper has been accepted for presentation at the 28th International Conference on Pattern Recognition (ICPR 2026) in Lyon, France Code: https://github.com/zadid6pretam/iStructTab PyPI: pip install istructtab

Journal ref International Conference on Pattern Recognition (ICPR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11621 2026-07-14 cs.AI cs.CL cs.LG 新提交 81%

Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns

受损多模态语言模型再现失语症患者的图片命名模式

Yong Yang, Xiang Guan, Sophie Arheix-Parras, Saeed Ahmadi, Roger Newman-Norlund, Leonardo Bonilha, Christopher Rorden, Julius Fridriksson, Rutvik H. Desai, Srihari Nelakuditi

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 研究未针对临床模拟设计的通用语言模型能否再现失语症患者图片命名错误模式,通过对多模态语言模型进行扰动配置,建立定量框架再现个体失语症错误模式,表明语言模型或可成为中风后失语症患者数字替身。

Comments 15 pages, 8 figures; supplementary materials (18 pages, 6 sections) included

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07328 2026-07-10 cs.LG cs.AI cs.CV cs.CY 版本更新 81%

MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation

MultiFair:具有双级梯度调制的多模态平衡公平感知医学分类

Md Zubair, Hao Zheng, Grayson W. Armstrong, Lucy Q. Shen, Gabriela Wilson, Yu Tian, Xingquan Zhu

机构 * School of Computing and Informatics, University of Louisiana at Lafayette(路易斯安那州立大学拉法叶分校计算机与信息学学院) Louisiana Center for Health Innovation and College of Nursing & Health Sciences, University of Louisiana at Lafayette(路易斯安那州立大学拉法叶分校健康创新中心及护理与健康科学学院) Massachusetts Eye and Ear, Harvard Medical School(哈佛医学院马萨诸塞眼耳医院) Department of Computer Science, University of Central Florida(佛罗里达州立大学计算机科学系) Department of Electrical Engineering and Computer Science, Florida Atlantic University(佛罗里达Atlantic大学电子工程与计算机科学系)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 针对多模态医学分类中数据模态学习不均衡和模型对特定群体不公平的问题,提出MultiFair方法,通过双级梯度调制过程动态调整训练梯度,在多数据集上评估,有效解决了上述挑战。

Comments This work has been accepted for publication in IEEE Transactions on Medical Imaging

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05798 2026-07-08 cs.CV cs.AI 新提交 81%

Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning

回答前分割:用于多模态大语言模型视觉推理的像素定位

Yake Wei, Yuan Wang, Fengyun Rao, Jing Lyu, Di Hu

专题命中 其他多模态 :MLLM(title);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文针对多模态大语言模型视觉推理,提出SegAnswer方法,将放大单元从边界框转为像素级分割掩码,能精准定位感兴趣区域,过滤冗余信息,与视觉令牌构建方式更契合,经多基准测试验证了其在像素定位及分割任务上的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30355 2026-06-30 cs.CV cs.AI 81%

Residual-Guided Expert Specialization for Incomplete Multimodal Learning

残差引导的专家特化用于不完整多模态学习

Seunghun Baek, Jihwan Park, Jaeyoon Sim, Minjae Jeong, Hoseok Lee, Won Hwa Kim

机构 * Pohang University of Science and Technology (POSTECH)(釜山科学技术大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 提出MARS框架,利用缺失引起的表示偏差引导专家特化,通过残差信号路由样本到对应专家,并引入差异感知噪声正则化缓解训练-测试路由差距,在缺失场景下提升多模态分类与分割性能。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13018 2026-06-23 cs.CY cs.AI cs.CV cs.SI 81%

GeoLocator: a location-integrated large multimodal model for inferring geo-privacy

GeoLocator:一种集成位置的大型多模态模型,用于推断地理隐私

Yifan Yang, Siqin Wang, Daoyang Li, Yixian Zhang, Shuju Sun, Junzhou He

机构 * Spatial Sciences Institute, University of Southern California(南加州大学空间科学研究所) Evolutionary Assets(进化资产) Viterbi school of engineering, University of Southern California(南加州大学维特比工程学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出GeoLocator模型,用于推断输入影像和社会媒体中的地理位置信息,揭示了模型在地理隐私保护中的潜在风险。

Comments 16pages, 2 figures

Journal ref Appl. Sci. 14(16), 7091 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00215 2026-06-17 cs.CV cs.CL 版本更新 81%

Disentangling Perception and Reasoning in Multimodal LLMs via Reward Design

通过奖励设计解耦多模态大语言模型中的感知与推理

Omar Sharif, Eftekhar Hossain, Nikhil Singh, Patrick Ng

机构 * Department of Computer Science, Dartmouth College(达特茅斯大学计算机科学系) Department of Computer Science, University of Central Florida(中央佛罗里达大学计算机科学系) Independent Researcher(独立研究者)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 研究多模态大模型中感知与推理的瓶颈,发现感知是主要约束,并通过奖励设计提升视觉基础推理,平均提升5.56分。

Comments 24 pages, 15 Figures, 10 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01444 2026-06-16 cs.AI cs.CL cs.LG 版本更新 81%

Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning

双不确定性引导的多模态推理策略学习

Rui Liu, Dian Yu, Tong Zheng, Runpeng Dai, Zongxia Li, Wenhao Yu, Zhenwen Liang, Linfeng Song, Haitao Mi, Pratap Tokekar, Dong Yu

机构 * Tencent Hunyuan(腾讯文汇) University of Maryland(马里兰大学) University of North Carolina(北卡罗来纳大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 提出DUPL方法,通过量化感知不确定性和输出不确定性来引导策略更新,在多个多模态推理基准上显著提升模型准确率,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04415 2026-05-12 cs.CL cs.CV 81%

Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training

双调优:面向多模态大语言模型训练的推理效能驱动的数据筛选

Ruobing Zheng, Tianqi Li, Jianing Li, Qingpei Guo, Yi Yuan, Jingdong Chen

机构 * Ant Group(蚂蚁集团)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出双调优框架,通过评估训练数据对推理训练的收益及推理训练的效能,指导多模态大语言模型的数据筛选与训练策略匹配。

Comments Project Page: https://digital-avatar.github.io/ai/ThinkingBoundary/

详情

展开后加载摘要…

URL PDF HTML 收藏