arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4878 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4878 篇

1801.01560 2018-01-08 cs.CV 74%

On-the-fly Augmented Reality for Orthopaedic Surgery Using a Multi-Modal Fiducial

Sebastian Andress, Alex Johnson, Mathias Unberath, Alexander Winkler, Kevin Yu, Javad Fotouhi, Simon Weidert, Greg Osgood, Nassir Navab

专题命中 其他多模态 :multi-modal(title);分类 cs.CV

Comments S. Andress, A. Johnson, M. Unberath, and A. Winkler have contributed equally and are listed in alphabetical order

Journal ref J. Med. Imag. 5(2), 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.09458 2017-12-29 cs.CV 74%

Multi-modal Geolocation Estimation Using Deep Neural Networks

Jesse M. Johns, Jeremiah Rounds, Michael J. Henry

专题命中 其他多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.00602 2016-10-04 cs.CL 74%

Multimodal Semantic Simulations of Linguistically Underspecified Motion Events

Nikhil Krishnaswamy, James Pustejovsky

专题命中 其他多模态 :multimodal(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.03369 2015-11-12 cs.CV physics.data-an physics.med-ph 74%

Multimodal MRI Neuroimaging with Motion Compensation Based on Particle Filtering

Yu-Hui Chen, Roni Mittelman, Boklye Kim, Charles Meyer, Alfred Hero

专题命中 其他多模态 :multimodal(title);分类 cs.CV

Comments This paper has been submitted to Transaction on Medical Imaging

详情

展开后加载摘要…

URL PDF HTML 收藏
1504.04763 2015-04-21 cs.CV 74%

Understanding the Fisher Vector: a multimodal part model

David Novotný, Diane Larlus, Florent Perronnin, Andrea Vedaldi

专题命中 其他多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1301.3666 2013-03-21 cs.CV cs.LG 74%

Zero-Shot Learning Through Cross-Modal Transfer

Richard Socher, Milind Ganjoo, Hamsa Sridhar, Osbert Bastani, Christopher D. Manning, Andrew Y. Ng

专题命中 其他多模态 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28149 2026-07-17 cs.CV cs.AI 版本更新 73%

Toward Robust In-Context Segmentation via Concept Guidance

通过概念引导实现鲁棒的上下文分割

Zhigang Chen, Xiawu Zheng, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)

专题命中 其他多模态 :MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 提出概念引导的上下文分割(CG-ICS),通过提取参考图像的高层语义概念而非仅依赖低层视觉匹配,结合文本概念与视觉示例,显著提升分割准确性和鲁棒性。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03647 2026-04-09 cs.CV cs.AI 73%

Stabilizing Unsupervised Self-Evolution of MLLMs via Continuous Softened Retracing reSampling

通过连续软化回溯重采样稳定多模态大语言模型的无监督自进化

Yunyao Yu, Zhengxian Wu, Zhuohong Chen, Hangrui Xu, Zirui Liao, Xiangwen Deng, Zhifang Liu, Senyuan Shi, Haoqian Wang

机构 * Tsinghua University(清华大学) Hefei University of Technology(合肥工业大学) University of Arizona(亚利桑那大学) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统实验室)

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 本文提出CSRS方法,通过回溯推理机制和软化频率奖励提升多模态大语言模型的无监督自进化稳定性,实验显示在MathVision等基准上表现优异。

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16590 2026-03-18 cs.CL cs.AI 73%

BATQuant: Outlier-resilient MXFP4 Quantization via Learnable Block-wise Optimization

BATQuant: 通过可学习的分块优化实现抗异常的MXFP4量化

Ji-Fu Li, Manyi Zhang, Xiaobo Xia, Han Bao, Haoli Bai, Zhenhua Dong, Xianzhi Yu

机构 * Huawei Technologies University of Science(华为技术大学科学)

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出BATQuant方法,通过可学习的分块优化解决MXFP4量化中异常传播问题,实现高性能量化方案。

Comments 30 pages, 13 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02789 2026-03-04 cs.CL cs.AI 73%

OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets

OCR 或不是?在 MLLMs 时代重新思考文档信息提取:基于真实世界的大规模数据集

Jiyuan Shen, Peiyue Yuan, Atin Ghosh, Yifan Mai, Daniel Dahlmeier

机构 * SAP(SAP公司) Stanford University(斯坦福大学)

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

AI总结 本文探讨了在 MLLMs 时代是否仍需 OCR,通过大规模数据集评估发现,仅图像输入可达到与 OCR 增强方法相当的性能,并展示了通过精心设计的模式、示例和指示可进一步提升 MLLMs 的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15915 2026-02-19 cs.CV cs.AI 73%

MaS-VQA: A Mask-and-Select Framework for Knowledge-Based Visual Question Answering

MaS-VQA: 一种基于掩码和选择的基于知识的视觉问答框架

Xianwei Mao, Kai Ye, Sheng Zhou, Nan Zhang, Haikuan Huang, Bin Li, Jiajun Bu

机构 * Zhejiang University, Hangzhou, China(浙江大学, 杭州, 中国) Alibaba Group, Hangzhou, China(阿里巴巴集团, 杭州, 中国)

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 MaS-VQA通过结合显式知识过滤与隐式知识推理,提升基于知识的视觉问答任务的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08717 2026-02-10 cs.CV cs.AI 73%

Zero-shot System for Automatic Body Region Detection for Volumetric CT and MR Images

零样本系统用于体积CT和MRI图像的自动身体区域检测

Farnaz Khun Jush, Grit Werner, Mark Klemens, Matthias Lenga

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 本文提出零样本系统用于体积CT和MRI图像自动身体区域检测,通过预训练模型实现无监督分割,展示了基于规则和多模态语言模型的性能对比。

Comments 8 pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07104 2026-02-10 cs.CV cs.AI 73%

Extended to Reality: Prompt Injection in 3D Environments

扩展到现实:3D环境中的提示注入

Zhuoheng Li, Ying Chen

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 研究提出PI3D攻击,通过在3D环境中放置带文本的物理物体来注入提示,挑战多模态大语言模型的安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10111 2026-01-22 cs.CV cs.AI cs.CR 73%

Training-Free In-Context Forensic Chain for Image Manipulation Detection and Localization

无需训练的上下文取证链用于图像篡改检测与定位

Rui Chen, Bin Liu, Changtao Miao, Xinghao Wang, Yi Li, Tao Gong, Qi Chu, Nenghai Yu

专题命中 其他多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 ICFC提出一种无需训练的多模态大语言模型框架,用于图像篡改检测与定位,通过可解释的推理流程实现高效且准确的图像分析。

Comments This version was uploaded in error and contains misleading information found in an early draft. The manuscript requires extensive and long-term revisions

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20531 2025-10-24 cs.CV cs.AI 73%

Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis

Lixiong Qin, Yang Zhang, Mei Wang, Jiani Hu, Weihong Deng, Weiran Xu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Normal University(北京师范大学)

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 25 pages, 9 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08830 2025-08-13 cs.AI cs.CV cs.CY 73%

Silicon Minds versus Human Hearts: The Wisdom of Crowds Beats the Wisdom of AI in Emotion Recognition

Mustafa Akben, Vinayaka Gude, Haya Ajjan

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12916 2025-07-18 cs.CV cs.AI 73%

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models

Yifan Xu, Chao Zhang, Hanqi Jiang, Xiaoyan Wang, Ruifei Ma, Yiwei Li, Zihao Wu, Zeju Li, Xiangde Liu

机构 * School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) Beijing Digital Native Digital City Research Center(北京数字原生数字城市研究院) School of Computing, The University of Georgia(佐治亚大学计算机学院) School of Computer and Communication Engineering, University of Science and Technology Beijing(北京科技大学计算机与通信工程学院) Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学(深圳)计算机科学与工程系)

专题命中 其他多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

Comments Accepted by TNNLS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09531 2025-07-15 cs.CV cs.AI cs.LG 73%

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization

Son Nguyen, Giang Nguyen, Hung Dao, Thao Do, Daeyoung Kim

机构 * KAIST, South Korea(韩国加尔文科学技术院) Auburn University, US(美国阿肯色大学)

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01133 2024-09-04 cs.CV cs.AI 73%

Large Language Models Can Understanding Depth from Monocular Images

Zhongyi Xia, Tianzhao Wu

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07989 2024-08-16 cs.CV cs.AI 73%

IIU: Independent Inference Units for Knowledge-based Visual Question Answering

Yili Li, Jing Yu, Keke Gai, Gang Xiong

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02717 2024-07-11 cs.CV cs.AI 73%

Complementary Information Mutual Learning for Multimodality Medical Image Segmentation

Chuyun Shen, Wenhao Li, Haoqing Chen, Xiaoling Wang, Fengping Zhu, Yuxin Li, Xiangfeng Wang, Bo Jin

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 35 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03701 2023-09-29 cs.CV cs.AI 73%

LMEye: An Interactive Perception Network for Large Language Models

Yunxin Li, Baotian Hu, Xinyu Chen, Lin Ma, Yong Xu, Min Zhang

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.08965 2021-08-23 cs.CV cs.CL 73%

Localize, Group, and Select: Boosting Text-VQA by Scene Text Modeling

Xiaopeng Lu, Zhen Fan, Yansen Wang, Jean Oh, Carolyn P. Rose

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.02388 2021-08-12 cs.CV cs.AI 73%

TransRefer3D: Entity-and-Relation Aware Transformer for Fine-Grained 3D Visual Grounding

Dailan He, Yusheng Zhao, Junyu Luo, Tianrui Hui, Shaofei Huang, Aixi Zhang, Si Liu

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments ACM MM2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.00515 2020-10-06 cs.CV cs.CL 73%

Linguistic Structure Guided Context Modeling for Referring Image Segmentation

Tianrui Hui, Si Liu, Shaofei Huang, Guanbin Li, Sansi Yu, Faxi Zhang, Jizhong Han

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by ECCV 2020. Code is available at https://github.com/spyflying/LSCM-Refseg

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06094 2026-08-11 cs.LG 版本更新 71%

Modeling Normal Is All You Need: Joint Latent Clustering for Anomaly Detection in Multimodal Cyber-Physical Systems

你所需要的只是对正常情况建模:多模态网络物理系统中异常检测的联合潜在聚类

Alexander Apartsin, Yehudit Aperstein

机构 * Holon Institute of Technology (HIT)(霍隆技术学院) Afeka Academic College of Engineering(阿法卡工程学院)

专题命中 其他多模态 :multimodal(title)

AI总结 研究多模态网络物理系统异常检测,提出联合潜在聚类方法,通过MIIM假设集、公平协议及潜在评分建模正常行为,在三个真实数据集上表现优异,优于其他深度检测器。

Comments 17 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22046 2026-06-23 cs.RO 新提交 71%

A Multimodal Tiltwing Framework for Bioinspired Aerial Robots

一种用于仿生空中机器人的多模态倾转机翼框架

Krispin C. V. Broers, Sophie F. Armanini

机构 * Imperial College London(伦敦帝国理工学院)

专题命中 其他多模态 :multimodal(title)

AI总结 提出一种可切换悬停、高速前飞和高效滑翔的多模态倾转机翼框架,通过双独立扑翼推力矢量控制增强机动性,并采用混合苏格兰轭扑动机构实现宽扑动角度以利用拍合效应。

Comments 18 pages, 23 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1108.1042 2026-06-03 math.NA cs.NA math.OC 71%

On strong homogeneity of two global optimization algorithms based on statistical models of multimodal objective functions

基于多模态目标函数统计模型的两种全局优化算法的强齐次性

Antanas Zilinskas

专题命中 其他多模态 :multimodal(title)

AI总结 本文提出利用无穷算术实现全局优化算法,并引入强齐次性概念,证明P算法和一步贝叶斯算法具有该性质。

Comments 11 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
1108.2126 2026-06-03 cs.RO cs.SY eess.SY math.OC 71%

Multi-Modal Local Sensing and Communication for Collective Underwater Systems

多模态本地感知与通信用于集体水下系统

Serge Kernbach, Tobias Dipper, Donny Sutantyo

专题命中 其他多模态 :multi-modal(title)

AI总结 本文研究集体水下系统中用于网络和集群模式的本地感知与通信,通过模态和子模态通信的特定组合实现多AUV间的专用协作。

Journal ref Proceedings of the 11th International Conference on Mobile Robots and Competitions, Robotica 2011, Lisbon, pp.96-101, 2011

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01784 2026-06-02 eess.IV 71%

MoRE: A Mixture-of-Experts-Based Task-Adaptive End-to-End Network for Multimodal MRI Reconstruction

MoRE:一种基于混合专家模型的任务自适应端到端网络用于多模态MRI重建

Yuyang Li, Yipin Deng, Wenlei Shang, Juncen Wu, Xin Bai, Zijian Zhou, Peng Hu

专题命中 其他多模态 :multimodal(title)

AI总结 提出MoRE,将稀疏激活的混合专家模块集成到端到端变分网络中,通过样本级无监督路由激活最小专家子集并保持物理一致性,在fastMRI多线圈脑部和膝关节数据集上实现高稳定SSIM和PSNR,且路由嵌入的t-SNE可视化揭示可解释的模态感知专家专业化。

Comments Accepted at the 2026 48th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC 2026), Toronto, Canada, July 26-30, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏