arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2507.19253 2025-07-28 cs.CV 79%

BridgeNet: A Unified Multimodal Framework for Bridging 2D and 3D Industrial Anomaly Detection

An Xiang, Zixuan Huang, Xitong Gao, Kejiang Ye, Cheng-zhong Xu

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Shenzhen University of Advanced Technology(深圳先进技术大学) State Key Lab of IOTSC, Department of CIS, University of Macau(物联网科学国家重点实验室,澳门大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19052 2025-07-28 cs.CV 79%

Probing Multimodal Fusion in the Brain: The Dominance of Audiovisual Streams in Naturalistic Encoding

Hamid Abdollahi, Amir Hossein Mansouri Majoumerd, Amir Hossein Bagheri Baboukani, Amir Abolfazl Suratgar, Mohammad Bagher Menhaj

机构 * Distributed and Intelligent Optimization Research Laboratory(分布式智能优化研究实验室) Electrical Engineering Department(电气工程系) Amirkabir University of Technology(阿米尔卡比尔技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10203 2025-07-22 cs.CV 79%

Improving Multimodal Learning via Imbalanced Learning

Shicai Wei, Chunbo Luo, Yang Luo

机构 * University of Electronic Science and Technology of China(电子科技大学) National and Local Joint Engineering Research Center for Cloud Operating System(云计算操作系统国家级地方联合工程研究中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14175 2025-07-22 cs.LG cs.AI stat.AP 79%

Latent Space Data Fusion Outperforms Early Fusion in Multimodal Mental Health Digital Phenotyping Data

Youcef Barkat, Dylan Hamitouche, Deven Parekh, Ivy Guo, David Benrimoh

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13026 2025-07-17 cs.CV 79%

HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model

Tao Wang, Changxu Cheng, Lingfeng Wang, Senda Chen, Wuyue Zhao

机构 * Uni-Ubi Zhejiang University(浙江大学) Tongji University(同济大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025; the code is at https://github.com/yayafengzi/LMM-HiMTok

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11003 2025-07-16 cs.CV 79%

Bridge Feature Matching and Cross-Modal Alignment with Mutual-filtering for Zero-shot Anomaly Detection

Yuhu Bai, Jiangning Zhang, Yunkang Cao, Guangyuan Lu, Qingdong He, Xiangtai Li, Guanzhong Tian

机构 * Zhejiang University(浙江大学) YouTu Lab, Tencent(腾讯YouTu实验室) Huazhong University of Science and Technology(华中科技大学) Peking University(北京大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10620 2025-07-16 cs.LG cs.AI 79%

LLMs Meet Cross-Modal Time Series Analytics: Overview and Directions

Chenxi Liu, Hao Miao, Cheng Long, Yan Zhao, Ziyue Li, Panos Kalnis

机构 * Nanyang Technological University(南洋理工大学) The Hong Kong Polytechnic University(香港理工大学) University of Electronic Science and Technology of China(电子科学与技术大学) Technical University of Munich(慕尼黑技术大学) King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.AI

Comments Accepted at SSTD 2025 (Tutorial). arXiv admin note: text overlap with arXiv:2505.02583

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09334 2025-07-15 cs.CV 79%

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding

Wencan Huang, Daizong Liu, Wei Hu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10715 2025-07-14 cs.CV 79%

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection

Yongjin Lee, Hyeon-Mun Jeong, Yurim Jeon, Sanghyun Kim

机构 * ThorDrive Co., Ltd(ThorDrive公司) Seoul National University(首尔国立大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08000 2025-07-11 cs.CV cs.LG 79%

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models

Helen Qu, Sang Michael Xie

机构 * Flatiron Institute(Flatiron研究所) Stanford University(斯坦福大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15804 2025-07-11 cs.CV 79%

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Zongzhao Li, Zongyang Ma, Mingze Li, Songyou Li, Yu Rong, Tingyang Xu, Ziqi Zhang, Deli Zhao, Wenbing Huang

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) DAMO Academy, Alibaba Group, Hangzhou, China(阿里云达摩院) Hupan Lab, Hangzhou, China(杭州华普实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15285 2025-07-11 cs.CV 79%

EEPNet-V2: Patch-to-Pixel Solution for Efficient Cross-Modal Registration between LiDAR Point Cloud and Camera Image

Yuanchao Yue, Hui Yuan, Zhengxin Li, Shuai Li, Wei Zhang

机构 * School of Control Science and Engineering, Shandong University(控制科学与工程学院,山东大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17968 2025-07-10 cs.CV 79%

A Multimodal Fusion Framework for Bridge Defect Detection with Cross-Verification

Ravi Datta Rachuri, Duoduo Liao, Samhita Sarikonda, Datha Vaishnavi Kondur

机构 * School of Computing George Mason University Fairfax, VA(计算学院乔治·马歇尔大学弗劳伊德堡)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Big Data 2024

Journal ref 2024 IEEE International Conference on Big Data (BigData)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04945 2025-07-09 cs.CL cs.LG eess.SP 79%

MEIT: Multimodal Electrocardiogram Instruction Tuning on Large Language Models for Report Generation

Zhongwei Wan, Che Liu, Xin Wang, Chaofan Tao, Hui Shen, Jing Xiong, Rossella Arcucci, Huaxiu Yao, Mi Zhang

机构 * The Ohio State University(俄亥俄州立大学) Imperial College London(伦敦帝国理工学院) The University of Hong Kong(香港大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05165 2025-07-08 cs.CV 79%

Differential Attention for Multimodal Crisis Event Analysis

Nusrat Munia, Junfeng Zhu, Olfa Nasraoui, Abdullah-Al-Zubaer Imran

机构 * University of Kentucky(肯塔基大学) Kentucky Geological Survey(肯塔基地质调查局) University of Louisville(路易斯维尔大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Presented at CVPRw 2025, MMFM3

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04891 2025-07-08 eess.IV cs.CV 79%

MurreNet: Modeling Holistic Multimodal Interactions Between Histopathology and Genomic Profiles for Survival Prediction

Mingxin Liu, Chengfei Cai, Jun Li, Pengbo Xu, Jinze Li, Jiquan Ma, Jun Xu

机构 * Jiangsu Key Laboratory of Intelligent Medical Image Computing(江苏省智能医学图像计算重点实验室) School of Artificial Intelligence(人工智能学院) Nanjing University of Information Science and Technology(南京信息工程大学) College of Information Engineering(信息工程学院) Taizhou University(泰州大学) College of Bioinformatics Science and Technology(生物信息科学与技术学院) Harbin Medical University(哈尔滨医科大学) School of Computer and Big Data(计算机与大数据学院) Heilongjiang University(黑龙江大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 2 figures, Accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04369 2025-07-08 cs.CV 79%

MambaFusion: Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection

Hanshi Wang, Jin Gao, Weiming Hu, Zhipeng Zhang

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA(多模态人工智能系统国家重点实验室(MAIS),中国科学院自动化所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) Anyverse Intelligence Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information(北京多模态信息超级智能安全重点实验室) School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03917 2025-07-08 cs.LG cs.CV 79%

Consistency-Aware Padding for Incomplete Multi-Modal Alignment Clustering Based on Self-Repellent Greedy Anchor Search

Shubin Ma, Liang Zhao, Mingdong Lu, Yifan Guo, Bo Xu

机构 * School of Software Technology, Dalian University of Technology(大连理工大学软件学院)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments Accepted at IJCAI 2025. 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03304 2025-07-08 cs.CV 79%

Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations

Hai Huang, Yan Xia, Sashuai Zhou, Hanting Wang, Shulei Wang, Zhou Zhao

机构 * Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02908 2025-07-08 cs.LG cs.AI 79%

Hyperbolic Kernel Graph Neural Networks for Neurocognitive Decline Analysis from Multimodal Brain Imaging

Meimei Yang, Yongheng Sun, Qianqian Wang, Andrea Bozoki, Maureen Kohi, Mingxia Liu

机构 * Department of Radiology and Biomedical Research Imaging Center (BRIC), University of North Carolina at Chapel Hill(放射科与生物医学研究成像中心(BRIC)、北卡罗来纳大学教堂山分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 14 pages, 5 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02205 2025-07-08 cs.CV 79%

Team RAS in 9th ABAW Competition: Multimodal Compound Expression Recognition Approach

Elena Ryumina, Maxim Markitantov, Alexandr Axyonov, Dmitry Ryumin, Mikhail Dolgushin, Alexey Karpov

机构 * St. Petersburg Federal Research Center of the Russian Academy of Sciences(俄罗斯科学院圣彼得堡联邦研究中心) ITMO University(ITMO大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 7

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05945 2025-07-04 cs.CV 79%

MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection

Zitian Wang, Zehao Huang, Yulu Gao, Naiyan Wang, Si Liu

机构 * Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02279 2025-07-04 cs.CV 79%

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

Juntao Liu, Liqiang Niu, Wenchao Chen, Jie Zhou, Fandong Meng

机构 * Pattern Recognition Center, WeChat AI, Tencent Inc(模式识别中心、微信AI、腾讯公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01984 2025-07-04 cs.LG cs.CL cs.SI 79%

Multimodal Misinformation Detection Using Early Fusion of Linguistic, Visual, and Social Features

Gautam Kishore Shahi

机构 * University of Duisburg-Essen(杜伊斯堡-埃森大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01054 2025-07-03 cs.LG cond-mat.mtrl-sci cs.AI 79%

XxaCT-NN: Structure Agnostic Multimodal Learning for Materials Science

Jithendaraa Subramanian, Linda Hung, Daniel Schweigert, Santosh Suram, Weike Ye

机构 * Toyota Research Institute(丰田研究机构)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04353 2025-07-03 cs.CV cs.LG 79%

DeFusion: An Effective Decoupling Fusion Network for Multi-Modal Pregnancy Prediction

Xueqiang Ouyang, Jia Wei, Wenjie Huo, Xiaocong Wang, Rui Li, Jianlong Zhou

机构 * School of Computer Science and Engineering, South China University of Technology(南方科技大学计算机科学与工程学院) Department of Obstetrics and Gynecology, Nanfang Hospital, Southern Medical University(南方医科大学 obstetrics and gynecology 部门) Golisano College of Computing and Information Sciences, Rochester Institute of Technology(罗切斯特理工学院 computing and information sciences 学院) UTS Data Science Institute, University of Technology Sydney(悉尼大学技术科学研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00849 2025-07-02 cs.CV 79%

UAVD-Mamba: Deformable Token Fusion Vision Mamba for Multimodal UAV Detection

Wei Li, Jiaman Tang, Yang Li, Beihao Xia, Ligang Tan, Hongmao Qin

机构 * College of Mechanical and Vehicle Engineering, Hunan University(湖南大学机械与车辆工程学院) School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments The paper was accepted by the 36th IEEE Intelligent Vehicles Symposium (IEEE IV 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00506 2025-07-02 cs.CV 79%

SCING:Towards More Efficient and Robust Person Re-Identification through Selective Cross-modal Prompt Tuning

Yunfei Xie, Yuxuan Cheng, Juncheng Wu, Haoyu Zhang, Yuyin Zhou, Shoudong Han

机构 * Huazhong University of Science and Technology(华中科技大学) Huazhong Agricultural University(华中农业大学) University of California, Santa Cruz(加州大学圣克ruz分校) City University of Hong Kong (Dongguan)(香港城市大学(东莞))

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16297 2025-07-01 cs.CV 79%

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

Renshan Zhang, Rui Shao, Gongwei Chen, Miao Zhang, Kaiwen Zhou, Weili Guan, Liqiang Nie

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to the IEEE/CVF International Conference on Computer Vision (ICCV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10557 2025-07-01 cs.CL 79%

MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models

Jianhong Tu, Zhuohao Ni, Nicholas Crispino, Zihao Yu, Michael Bendersky, Beliz Gunel, Ruoxi Jia, Xin Liu, Lingjuan Lyu, Dawn Song, Chenguang Wang

机构 * Washington University in St. Louis(华盛顿大学圣路易斯分校) The University of British Columbia(不列颠哥伦比亚大学) Google Research(谷歌研究) Virginia Tech(弗吉尼亚理工大学) University of California, Davis(加州大学戴维斯分校) Sony AI(索尼人工智能) University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏