arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6878 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6878 篇

1908.11216 2019-09-11 cs.CL cs.AI cs.IR 81%

From the Token to the Review: A Hierarchical Multimodal approach to Opinion Mining

Alexandre Garcia, Pierre Colombo, Slim Essid, Florence d'Alché-Buc, Chloé Clavel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP) and 9th International Joint Conference on Natural Language Processing (IJCNLP)

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.06008 2019-08-19 cs.LG cs.AI cs.CL stat.ML 81%

Variational Fusion for Multimodal Sentiment Analysis

Navonil Majumder, Soujanya Poria, Gangeshwar Krishnamurthy, Niyati Chhaya, Rada Mihalcea, Alexander Gelbukh

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.05067 2019-08-15 cs.CL cs.CV 81%

Reactive Multi-Stage Feature Fusion for Multimodal Dialogue Modeling

Yi-Ting Yeh, Tzu-Chuan Lin, Hsiao-Hua Cheng, Yu-Hsuan Deng, Shang-Yu Su, Yun-Nung Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted for a poster session at the DSTC7 workshop at AAAI 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.03196 2019-07-09 cs.CV eess.AS eess.IV 81%

Multimodal Fusion with Deep Neural Networks for Audio-Video Emotion Recognition

Juan D. S. Ortega, Mohammed Senoussaoui, Eric Granger, Marco Pedersoli, Patrick Cardinal, Alessandro L. Koerich

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.00378 2019-06-04 cs.CL cs.CV 81%

Unsupervised Bilingual Lexicon Induction from Mono-lingual Multimodal Data

Shizhe Chen, Qin Jin, Alexander Hauptmann

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by AAAI 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.02627 2018-11-08 cs.CV cs.AI 81%

Vehicle Tracking Using Surveillance with Multimodal Data Fusion

Yue Zhang, Bin Song, Xiaojiang Du, Mohsen Guizani

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages,6 figures,33 conferences

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.10565 2018-10-26 cs.CV cs.AI cs.HC 81%

Multimodal Polynomial Fusion for Detecting Driver Distraction

Yulun Du, Chirag Raman, Alan W Black, Louis-Philippe Morency, Maxine Eskenazi

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments INTERSPEECH 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.09779 2018-10-15 cs.CV cs.AI cs.LG q-bio.QM q-bio.TO 81%

Deep Multi-Modal Classification of Intraductal Papillary Mucinous Neoplasms (IPMN) with Canonical Correlation Analysis

Sarfaraz Hussein, Pujan Kandel, Juan E. Corral, Candice W. Bolan, Michael B. Wallace, Ulas Bagci

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted for publication in IEEE International Symposium on Biomedical Imaging (ISBI) 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.06225 2018-09-18 cs.CV cs.AI 81%

Investigation of Multimodal Features, Classifiers and Fusion Methods for Emotion Recognition

Zheng Lian, Ya Li, Jianhua Tao, Jian Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages, 11 figures and 4 Tables. EmotiW2018 challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.04456 2018-09-17 cs.LG cs.AI cs.CV stat.ML 81%

Multimodal Deep Neural Networks using Both Engineered and Learned Representations for Biodegradability Prediction

Garrett B. Goh, Khushmeen Sakloth, Charles Siegel, Abhinav Vishnu, Jim Pfaendtner

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Submitted to a peer-reviewed ML conference

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.06228 2018-06-19 cs.CL cs.CV 81%

Multimodal Sentiment Analysis using Hierarchical Fusion with Context Modeling

N. Majumder, D. Hazarika, A. Gelbukh, E. Cambria, S. Poria

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted for publication at Knowledge Based Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.11264 2018-05-30 stat.ML cs.CL cs.LG cs.SD eess.AS 81%

Disentangling by Partitioning: A Representation Learning Framework for Multimodal Sensory Data

Wei-Ning Hsu, James Glass

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.05106 2018-04-27 cs.MM cs.CV cs.LG 81%

CM-GANs: Cross-modal Generative Adversarial Networks for Common Representation Learning

Yuxin Peng, Jinwei Qi, Yuxin Yuan

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.09816 2018-02-28 cs.CV cs.AI cs.LG cs.NE stat.ML 81%

Coarse to fine non-rigid registration: a chain of scale-specific neural networks for multimodal image alignment with application to remote sensing

Armand Zampieri, Guillaume Charpiat, Yuliya Tarabalka

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.00924 2018-02-06 cs.LG cs.AI cs.CL stat.ML 81%

Multimodal Sentiment Analysis with Word-Level Fusion and Reinforcement Learning

Minghai Chen, Sen Wang, Paul Pu Liang, Tadas Baltrušaitis, Amir Zadeh, Louis-Philippe Morency

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments ICMI 2017 Oral Presentation, Honorable Mention Award

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.06228 2017-12-19 cs.CV cs.AI cs.LG 81%

Visual Explanations from Hadamard Product in Multimodal Deep Networks

Jin-Hwa Kim, Byoung-Tak Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 5 figures, including appendix, NIPS 2017 Workshop on Visually-Grounded Interaction and Language (ViGIL)

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.02251 2017-09-08 cs.CV cs.LG cs.MM 81%

Multi-modal Conditional Attention Fusion for Dimensional Emotion Prediction

Shizhe Chen, Qin Jin

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments Appeared at ACM Multimedia 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.00153 2017-06-27 cs.MM cs.CV cs.LG 81%

Cross-modal Common Representation Learning by Hybrid Transfer Network

Xin Huang, Yuxin Peng, Mingkuan Yuan

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments To appear in the proceedings of 26th International Joint Conference on Artificial Intelligence (IJCAI), Melbourne, Australia, Aug. 19-25, 2017. 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0703091 2016-08-14 cs.AI cs.MM 81%

Multimodal Meaning Representation for Generic Dialogue Systems Architectures

Frédéric Landragin, Alexandre Denis, Annalisa Ricci, Laurent Romary

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI、cs.MM

Journal ref Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC 2004) (2004) 521-524

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22158 2026-03-24 cs.LG cs.AI 80%

Multimodal Survival Analysis with Locally Deployable Large Language Models

多模态生存分析与可本地部署的大语言模型

Moritz Gögl, Christopher Yau

机构 * University of Oxford(牛津大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI;multi-modal(comments)

AI总结 本文提出利用可本地部署的大语言模型进行多模态生存分析,结合临床文本、表格数据和基因组数据,通过教师-学生蒸馏和原理化的多模态融合,实现校准的生存概率估计和简洁的诊断文本生成,优于标准基线并在隐私和准确性方面表现更优。

Comments NeurIPS 2025 Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18652 2026-03-17 cs.CL 80%

PolyFrame at MWE-2026 AdMIRe 2: When Words Are Not Enough: Multimodal Idiom Disambiguation

PolyFrame在MWE-2026 AdMIRe 2:当词语不够时:多模态成语消歧

Nina Hosseini-Kivanani

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出PolyFrame系统,通过统一管道处理图像+文本排名和纯文本描述排名任务,利用轻量模块提升多语言成语消歧性能,无需微调大模型。

Comments Accepted at AdMIRe 2 shared task (Advancing Multimodal Idiomaticity Representation) colocated with 22nd Workshop on Multiword Expressions (MWE 2026) @EACL2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03622 2026-02-04 cs.CV physics.med-ph 80%

Quasi-multimodal-based pathophysiological feature learning for retinal disease diagnosis

基于准多模态的病理特征学习用于视网膜疾病诊断

Lu Zhang, Huizhen Yu, Zuowei Wang, Fu Gui, Yatu Guo, Wei Zhang, Mengyu Jia

机构 * Tianjin University(天津大学) Tianjin Key Laboratory of Ophthalmology and Visual Science(天津眼科学与视觉科学重点实验室) Tianjin Eye Institute(天津眼科研究院) Tianjin Eye Hospital(天津眼科医院) Clinical College of Ophthalmology, Tianjin Medical University(天津医科大学临床医学院) Department of Ophthalmology, The Second Affiliated Hospital of Nanchang University(南昌大学第二附属医院眼科部) Nankai University Affiliated Eye Hospital(南开大学附属眼科医院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于准多模态的视网膜疾病诊断方法,通过多模态数据合成与融合提升分类和分级的准确性。

Journal ref Zhang, L., Yu, H., Wang, Z., Gui, F., Guo, Y., Zhang, W., Jia, M., 2026. Quasi-multimodal-based pathophysiological feature learning for retinal disease diagnosis. Medical Image Analysis 109, 103886

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09470 2026-01-15 physics.ed-ph cs.AI 80%

Personalized Multimodal Feedback Using Multiple External Representations: Strategy Profiles and Learning in High School Physics

基于多种外部表征的个性化反馈:策略配置与高中物理学习中的学习

Natalia Revenga-Lozano, Karina E. Avila, Steffen Steinert, Matthias Schweinberger, Clara E. Gómez-Pérez, Jochen Kuhn, Stefan Küchemann

机构 * Chair of Physics Education, Faculty of Physics, Ludwig-Maximilians-Universität München (LMU Munich)(物理教育系主任,物理学院,慕尼黑路易斯-马克西姆利安大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文研究了多种外部表征与个性化反馈在高中物理学习中的整合效果,发现详细多表征反馈对学习成绩有积极影响,且学习者根据表征能力选择不同反馈策略。

Comments Keywords: Adaptive Feedback, Multimodal Learning, Multiple External Representations, Physics Education, Science Education, Representational Competences, Intelligent Tutoring Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15986 2025-11-25 cs.CV cs.CY cs.LG 80%

Fairness in Multi-modal Medical Diagnosis with Demonstration Selection

多模态医学诊断中的公平性与演示选择

Dawei Li, Zijian Gu, Peng Wang, Chuhan Song, Zhen Tan, Mohan Zhang, Tianlong Chen, Yu Tian, Song Wang

机构 * Arizona State University(亚利桑那州立大学) University of Rochester(罗切斯特大学) University of Virginia(弗吉尼亚大学) UCL(伦敦大学学院) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Central Florida(佛罗里达中央大学)

专题命中 多模态训练与对齐 :multi-modal(title,comments);multimodal(abstract);分类 cs.CV

AI总结 本文提出FADS方法,通过基于聚类的采样提升多模态医学影像诊断的公平性,减少性别、种族和族裔相关差异,同时保持高准确性。

Comments 10 pages (including 2 pages of references), 4 figures. This work explores fairness in multi-modal medical image reasoning using in-context learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05177 2025-10-29 cs.CV 80%

Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy

Yunhang Shen, Chaoyou Fu, Shaoqi Dong, Xiong Wang, Yi-Fan Zhang, Peixian Chen, Mengdan Zhang, Haoyu Cao, Ke Li, Shaohui Lin, Xiawu Zheng, Yan Zhang, Yiyi Zhou, Ran He, Caifeng Shan, Rongrong Ji, Xing Sun

机构 * Tencent Youtu Lab(腾讯云图实验室) Nanjing University(南京大学) East China Normal University(华东师范大学) Xiamen University(厦门大学) CASIA(中国科学院自动化研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV;MLLM(comments)

Comments https://github.com/VITA-MLLM/Long-VITA

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10353 2025-04-21 cs.CV cs.CR cs.LG 80%

Robust image classification with multi-modal large language models

Francesco Villani, Igor Maljkovic, Dario Lazzaro, Angelo Sotgiu, Antonio Emanuele Cinà, Fabio Roli

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV;multimodal(comments)

Comments Paper accepted at Pattern Recognition Letters journal Keywords: adversarial examples, rejection defense, multimodal-informed systems, machine learning security

Journal ref Pattern Recognition Letters 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14504 2025-03-25 cs.CV 80%

Aligning Multimodal LLM with Human Preference: A Survey

Tao Yu, Yi-Fan Zhang, Chaoyou Fu, Junkang Wu, Jinda Lu, Kun Wang, Xingyu Lu, Yunhang Shen, Guibin Zhang, Dingjie Song, Yibo Yan, Tianlong Xu, Qingsong Wen, Zhang Zhang, Yan Huang, Liang Wang, Tieniu Tan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Project page: https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.02884 2024-06-12 cs.CV 80%

Vision+X: A Survey on Multimodal Learning in the Light of Data

Ye Zhu, Yu Wu, Nicu Sebe, Yan Yan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Survey paper on multimodal learning and generation, to appear at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13765 2024-05-21 cs.AI 80%

Towards ethical multimodal systems

Alexis Roger, Esma Aïmeur, Irina Rish

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 5 pages, multimodal ethical dataset building, accepted in the NeurIPS 2023 MP2 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04125 2023-11-01 cs.LG cs.CL cs.HC 80%

Multimodal Fusion Interactions: A Study of Human and Automatic Quantification

Paul Pu Liang, Yun Cheng, Ruslan Salakhutdinov, Louis-Philippe Morency

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments International Conference on Multimodal Interaction (ICMI '23), Code available at: https://github.com/pliang279/PID. arXiv admin note: text overlap with arXiv:2302.12247

详情

展开后加载摘要…

URL PDF HTML 收藏