arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4868 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4868 篇

2108.02278 2021-08-06 cs.CV cs.AI q-bio.GN q-bio.QM q-bio.TO 81%

Pan-Cancer Integrative Histology-Genomic Analysis via Interpretable Multimodal Deep Learning

Richard J. Chen, Ming Y. Lu, Drew F. K. Williamson, Tiffany Y. Chen, Jana Lipkova, Muhammad Shaban, Maha Shady, Mane Williams, Bumjin Joo, Zahra Noor, Faisal Mahmood

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Demo: http://pancancer.mahmoodlab.org

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.13884 2021-07-06 cs.CV cs.CL cs.LG 81%

Multimodal Few-Shot Learning with Frozen Language Models

Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami, Oriol Vinyals, Felix Hill

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.05903 2021-06-11 cs.CL cs.CV cs.CY cs.LG 81%

Deciphering Implicit Hate: Evaluating Automated Detection Algorithms for Multimodal Hate

Austin Botelho, Bertie Vidgen, Scott A. Hale

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Please note the paper contains examples of hateful content

Journal ref Findings of ACL, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.14462 2021-06-01 cs.CL cs.AI 81%

Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine Translation

Zhiyong Wu, Lingpeng Kong, Wei Bi, Xiang Li, Ben Kao

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments To appear at ACL 2021 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.10044 2021-04-22 cs.CL cs.CV 81%

Cross-lingual Visual Pre-training for Multimodal Machine Translation

Ozan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Madhyastha, Erkut Erdem, Aykut Erdem, Lucia Specia

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted to EACL 2021 (Camera-ready version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.11115 2021-02-23 cs.LG cs.CL cs.CV 81%

Probing Multimodal Embeddings for Linguistic Properties: the Visual-Semantic Case

Adam Dahlgren Lindström, Suna Bensch, Johanna Björklund, Frank Drewes

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Submitted July 1 2020, COLING 2020 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.06735 2020-12-15 cs.CV cs.MM 81%

Multimodal In-bed Pose and Shape Estimation under the Blankets

Yu Yin, Joseph P. Robinson, Yun Fu

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.04951 2020-12-10 stat.ML cs.CV cs.LG cs.SD eess.AS 81%

Conjugate Mixture Models for Clustering Multimodal Data

Vasil Khalidov, Florence Forbes, Radu Horaud

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、eess.AS

Journal ref Neural Computation, 23(2), 2011

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.06850 2020-11-16 cs.CV cs.AI 81%

Transductive Zero-Shot Learning using Cross-Modal CycleGAN

Patrick Bordes, Eloi Zablocki, Benjamin Piwowarski, Patrick Gallinari

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.03651 2020-09-09 cs.LG cs.AI cs.CV stat.ML 81%

Learning more expressive joint distributions in multimodal variational methods

Sasho Nedelkoski, Mihail Bogojeski, Odej Kao

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 12 pages, Accepted and presented at LOD 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.04809 2020-01-29 cs.HC cs.AI cs.CL 81%

Detecting depression in dyadic conversations with multimodal narratives and visualizations

Joshua Y. Kim, Greyson Y. Kim, Kalina Yacef

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 12 pages

Journal ref AI 2019: Advances in Artificial Intelligence. AI 2019 vol 11919

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.00776 2020-01-08 cs.CV cs.MM 81%

Cross-modal Subspace Learning via Kernel Correlation Maximization and Discriminative Structure Preserving

Jun Yu, Xiao-Jun Wu

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments The paper is under consideration at Multimedia Tools and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.05728 2019-10-15 cs.CV cs.CL cs.LG 81%

Granular Multimodal Attention Networks for Visual Dialog

Badri N. Patro, Shivansh Patel, Vinay P. Namboodiri

专题命中 其他多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL

Comments ICCV Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.12200 2019-10-03 eess.IV cs.AI cs.CV cs.LG stat.ML 81%

Missing MRI Pulse Sequence Synthesis using Multi-Modal Generative Adversarial Network

Anmol Sharma, Ghassan Hamarneh

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted for publication in IEEE Transactions on Medical Imaging

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.07826 2019-09-04 cs.CL cs.CV 81%

Unsupervised Discovery of Multimodal Links in Multi-image, Multi-sentence Documents

Jack Hessel, Lillian Lee, David Mimno

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Code and data available at http://www.cs.cornell.edu/~jhessel/multiretrieval/multiretrieval.html

Journal ref EMNLP 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.08979 2019-08-27 cs.LG cs.CL cs.SD eess.AS stat.ML 81%

Controlling for Confounders in Multimodal Emotion Classification via Adversarial Learning

Mimansa Jaiswal, Zakaria Aldeneh, Emily Mower Provost

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、eess.AS

Comments 10 pages, ICMI 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.06371 2019-07-04 cs.CL cs.AI 81%

Multimodal Grounding for Language Processing

Lisa Beinborn, Teresa Botschen, Iryna Gurevych

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments The paper has been published in the Proceedings of the 27 Conference of Computational Linguistics. Please refer to this version for citations: https://www.aclweb.org/anthology/papers/C/C18/C18-1197/

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.11683 2019-05-31 cs.CV cs.CL cs.LG eess.IV 81%

Multi-level Multimodal Common Semantic Space for Image-Phrase Grounding

Hassan Akbari, Svebor Karaman, Surabhi Bhargava, Brian Chen, Carl Vondrick, Shih-Fu Chang

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted in CVPR 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.00839 2019-04-03 cs.CV cs.CL 81%

Improving Referring Expression Grounding with Cross-modal Attention-guided Erasing

Xihui Liu, Zihao Wang, Jing Shao, Xiaogang Wang, Hongsheng Li

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by CVPR 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.05485 2019-03-14 cs.AI cs.CL 81%

MMKG: Multi-Modal Knowledge Graphs

Ye Liu, Hui Li, Alberto Garcia-Duran, Mathias Niepert, Daniel Onoro-Rubio, David S. Rosenblum

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments ESWC 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.08010 2018-09-05 cs.AI cs.CV 81%

Multi-Modal Coreference Resolution with the Correlation between Space Structures

Qibin Zheng, Xingchun Diao, Jianjun Cao, Xiaolei Zhou, Yi Liu, Hongmei Li

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.03361 2018-07-11 cs.CV cs.AI cs.LG cs.NE 81%

Weakly-Supervised Convolutional Neural Networks for Multimodal Image Registration

Yipeng Hu, Marc Modat, Eli Gibson, Wenqi Li, Nooshin Ghavami, Ester Bonmati, Guotai Wang, Steven Bandula, Caroline M. Moore, Mark Emberton, Sébastien Ourselin, J. Alison Noble, Dean C. Barratt, Tom Vercauteren

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted manuscript in Medical Image Analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
1612.01895 2017-12-01 cs.CV cs.AI 81%

Multimodal Transfer: A Hierarchical Deep Convolutional Neural Network for Fast Artistic Style Transfer

Xin Wang, Geoffrey Oxholm, Da Zhang, Yuan-Fang Wang

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by CVPR 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.07177 2017-10-20 cs.CL cs.CV 81%

Findings of the Second Shared Task on Multimodal Machine Translation and Multilingual Image Description

Desmond Elliott, Stella Frank, Loïc Barrault, Fethi Bougares, Lucia Specia

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Journal ref Proceedings of the Second Conference on Machine Translation, 2017, pp. 215--233

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.03167 2017-07-12 cs.CV cs.AI cs.LG cs.RO 81%

RegNet: Multimodal Sensor Registration Using Deep Neural Networks

Nick Schneider, Florian Piewak, Christoph Stiller, Uwe Franke

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments published in IEEE Intelligent Vehicles Symposium, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.03714 2017-03-13 cs.CL cs.AI cs.HC cs.RO 81%

Applying the Wizard-of-Oz Technique to Multimodal Human-Robot Dialogue

Matthew Marge, Claire Bonial, Brendan Byrne, Taylor Cassidy, A. William Evans, Susan G. Hill, Clare Voss

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Presented at the 2016 IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN), Interactive Session, August 26-31, 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.07492 2017-02-27 cs.RO cs.AI cs.CV stat.ML 81%

Robot gains Social Intelligence through Multimodal Deep Reinforcement Learning

Ahmed Hussain Qureshi, Yutaka Nakamura, Yuichiro Yoshikawa, Hiroshi Ishiguro

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments The paper is published in IEEE-RAS International Conference on Humanoid Robots (Humanoids) 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02955 2026-08-04 cs.CV 80%

Multimodal image registration for effective thermographic fever screening

C. Y. N. Dwith, Pejhman Ghassemi, Joshua Pfefer, Jon Casamento, Quanzeng Wang

专题命中 其他多模态 :multimodal(title,journal_ref);multi-modal(abstract);分类 cs.CV

Journal ref Proceedings Volume 10057, Multimodal Biomedical Imaging XII 100570S, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09076 2026-04-13 cs.CV 80%

Cross-Modal Knowledge Distillation from Spatial Transcriptomics to Histology

从空间转录组到组织学的跨模态知识蒸馏

Arbel Hizmi, Artemii Bakulin, Shai Bagon, Nir Yosef

机构 * Weizmann Institute of Science(魏茨曼科学研究所) Reichman University(赖希曼大学)

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出利用空间转录组与H&E数据进行跨模态蒸馏,将转录组学的 niches 结构转移到仅依赖组织学的模型中,提升与转录组学 niches 结构的一致性,并通过细胞类型分析验证生物意义的邻近组成。

Comments Accepted to the CVMI Workshop at CVPR 2026. Project page: https://cross-modal-distillation.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14879 2025-10-31 cs.HC cs.AI 80%

Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-In-The-Loop LLM

Jiachen Li, Xiwen Li, Justin Steinberg, Akshat Choube, Bingsheng Yao, Xuhai Xu, Dakuo Wang, Elizabeth Mynatt, Varun Mishra

机构 * Northeastern University(东北大学) Columbia University(哥伦比亚大学)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Jiachen Li, Xiwen Li, Justin Steinberg, Akshat Choube, Bingsheng Yao, Xuhai Xu, Dakuo Wang, Elizabeth Mynatt, and Varun Mishra. 2025. Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-in-the-Loop LLM. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 9, 3, Article 101 (September 2025), 37 pages

Journal ref Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 9 (2025) 101:1-37

详情

展开后加载摘要…

URL PDF HTML 收藏