arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9119 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9119 篇

2209.15214 2023-03-21 cs.AI cs.CL cs.IR cs.LG 81%

Construction and Applications of Billion-Scale Pre-Trained Multimodal Business Knowledge Graph

Shumin Deng, Chengming Wang, Zhoubo Li, Ningyu Zhang, Zelin Dai, Hehong Chen, Feiyu Xiong, Ming Yan, Qiang Chen, Mosha Chen, Jiaoyan Chen, Jeff Z. Pan, Bryan Hooi, Huajun Chen

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments OpenBG. Accepted by ICDE 2023. The project is released at https://github.com/OpenBGBenchmark/OpenBG . Website: https://kg.alibaba.com/ , Leaderboard: https://tianchi.aliyun.com/dataset/dataDetail?dataId=122271

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.05556 2023-03-10 cs.CV cs.CL 81%

ViLPAct: A Benchmark for Compositional Generalization on Multimodal Human Activities

Terry Yue Zhuo, Yaqing Liao, Yuecheng Lei, Lizhen Qu, Gerard de Melo, Xiaojun Chang, Yazhou Ren, Zenglin Xu

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted at EACL2023 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11039 2023-03-10 cs.CV cs.CL 81%

Flat Multi-modal Interaction Transformer for Named Entity Recognition

Junyu Lu, Dixiang Zhang, Pingjian Zhang

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by COLING 2022, oral paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.09856 2023-03-08 cs.CL cs.SD eess.AS 81%

Knowledge-aware Bayesian Co-attention for Multimodal Emotion Recognition

Zihan Zhao, Yu Wang, Yanfeng Wang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、eess.AS

Comments Accepted to IEEE ICASSP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04613 2023-03-07 cs.CV cs.AI 81%

Enhancing Fine-Grained 3D Object Recognition using Hybrid Multi-Modal Vision Transformer-CNN Models

Songsong Xiong, Georgios Tziafas, Hamidreza Kasaei

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06565 2023-02-27 cs.LG cs.AI cs.CV eess.IV 81%

That's the Wrong Lung! Evaluating and Improving the Interpretability of Unsupervised Multimodal Encoders for Medical Data

Denis Jered McInerney, Geoffrey Young, Jan-Willem van de Meent, Byron C. Wallace

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.06560 2023-02-14 cs.CL cs.MM 81%

Large Scale Multi-Lingual Multi-Modal Summarization Dataset

Yash Verma, Anubhav Jangra, Raghvendra Kumar, Sriparna Saha

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04135 2023-02-09 cs.CV cs.AI 81%

Multi-Modal Evaluation Approach for Medical Image Segmentation

Seyed M. R. Modaresi, Aomar Osmani, Mohammadreza Razzazi, Abdelghani Chibani

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08279 2022-12-19 cs.LG cs.CL cs.CV 81%

Werewolf Among Us: A Multimodal Dataset for Modeling Persuasion Behaviors in Social Deduction Games

Bolin Lai, Hongxin Zhang, Miao Liu, Aryan Pariani, Fiona Ryan, Wenqi Jia, Shirley Anugrah Hayati, James M. Rehg, Diyi Yang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.04828 2022-12-09 cs.CV cs.AI 81%

FLAME: Facial Landmark Heatmap Activated Multimodal Gaze Estimation

Neelabh Sinha, Michal Balazia, Francois Bremond

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Preprint. Final paper accepted at the 17th IEEE International Conference on Advanced Video and Signal-based Surveillance (AVSS), virtual, November 2021. 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.02896 2022-12-07 cs.CV cs.AI 81%

Multimodal Tree Decoder for Table of Contents Extraction in Document Images

Pengfei Hu, Zhenrong Zhang, Jianshu Zhang, Jun Du, Jiajia Wu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ICPR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15425 2022-11-29 cs.CV cs.AI 81%

FAF: A novel multimodal emotion recognition approach integrating face, body and text

Zhongyu Fang, Aoyun He, Qihui Yu, Baopeng Gao, Weiping Ding, Tong Zhang, Lei Ma

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14451 2022-11-29 cs.CV cs.CL cs.LG 81%

GLAMI-1M: A Multilingual Image-Text Fashion Dataset

Vaclav Kosar, Antonín Hoskovec, Milan Šulc, Radek Bartyzal

专题命中 多模态评测 :image-text(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11049 2022-11-23 cs.CL cs.AI 81%

Explaining (Sarcastic) Utterances to Enhance Affect Understanding in Multimodal Dialogues

Shivani Kumar, Ishani Mondal, Md Shad Akhtar, Tanmoy Chakraborty

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted at AAAI 2023. 11 Pages; 14 Tables; 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.15359 2022-10-28 cs.CV cs.AI 81%

Exploiting modality-invariant feature for robust multimodal emotion recognition with missing modalities

Haolin Zuo, Rui Liu, Jinming Zhao, Guanglai Gao, Haizhou Li

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 5 pages, 3 figures, 1 table. Submitted to ICASSP 2023. We release the code at: https://github.com/ZhuoYulang/IF-MMIN

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.13134 2022-10-25 cs.CL cs.CV 81%

Multilingual Multimodal Learning with Machine Translated Text

Chen Qiu, Dan Oneata, Emanuele Bugliarello, Stella Frank, Desmond Elliott

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12775 2022-10-25 cs.CL cs.AI 81%

McQueen: a Benchmark for Multimodal Conversational Query Rewrite

Yifei Yuan, Chen Shi, Runze Wang, Liyi Chen, Feijun Jiang, Yuan You, Wai Lam

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP22

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.11896 2022-09-27 eess.IV cs.CV eess.AS 81%

Unsupervised active speaker detection in media content using cross-modal information

Rahul Sharma, Shrikanth Narayanan

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、eess.AS

Comments Under review at IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.08893 2022-09-01 cs.LG cs.AI cs.CL cs.SI 81%

Multimodal Learning on Graphs for Disease Relation Extraction

Yucong Lin, Keming Lu, Sheng Yu, Tianxi Cai, Marinka Zitnik

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.13626 2022-08-30 cs.AI cs.CV cs.LG cs.MA cs.RO 81%

CH-MARL: A Multimodal Benchmark for Cooperative, Heterogeneous Multi-Agent Reinforcement Learning

Vasu Sharma, Prasoon Goyal, Kaixiang Lin, Govind Thattai, Qiaozi Gao, Gaurav S. Sukhatme

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.10745 2022-05-24 cs.CV cs.AI 81%

Classification of Quasars, Galaxies, and Stars in the Mapping of the Universe Multi-modal Deep Learning

Sabeesh Ethiraj, Bharath Kumar Bolla

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Presented at Deep Learning Developers Conference, 2021, Bangalore

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.08007 2022-05-18 cs.MM cs.SD eess.AS eess.IV 81%

Perceptual Evaluation on Audio-visual Dataset of 360 Content

Randy F Fela, Andréas Pastor, Patrick Le Callet, Nick Zacharov, Toinon Vigier, Søren Forchhammer

专题命中 多模态评测 :audio-visual(title);multimodal(abstract);分类 cs.MM、eess.AS

Comments 6 pages, 5 figures, International Conference on Multimedia and Expo 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.01133 2022-05-09 cs.CL cs.CV cs.LG 81%

Hausa Visual Genome: A Dataset for Multi-Modal English to Hausa Machine Translation

Idris Abdulmumin, Satya Ranjan Dash, Musa Abdullahi Dawud, Shantipriya Parida, Shamsuddeen Hassan Muhammad, Ibrahim Sa'id Ahmad, Subhadarshi Panda, Ondřej Bojar, Bashir Shehu Galadanci, Bello Shehu Bello

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at Language Resources and Evaluation Conference 2022 (LREC2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.01850 2022-05-05 cs.CL cs.CV 81%

Visual Commonsense in Pretrained Unimodal and Multimodal Models

Chenyu Zhang, Benjamin Van Durme, Zhuowan Li, Elias Stengel-Eskin

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments To appear in NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.01879 2022-05-05 cs.CV cs.AI 81%

Learning Two-Stream CNN for Multi-Modal Age-related Macular Degeneration Categorization

Weisen Wang, Xirong Li, Zhiyan Xu, Weihong Yu, Jianchun Zhao, Dayong Ding, Youxin Chen

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by IEEE Journal of Biomedical and Health Informatics (J-BHI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.07154 2022-05-03 cs.CL cs.CV 81%

MMChat: Multi-Modal Chat Dataset on Social Media

Yinhe Zheng, Guanyi Chen, Xin Liu, Jian Sun

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by LREC2022. Dataset available in https://github.com/silverriver/MMChat

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09788 2022-04-22 cs.CV cs.AI 81%

SELMA: SEmantic Large-scale Multimodal Acquisitions in Variable Weather, Daytime and Viewpoints

Paolo Testolina, Francesco Barbato, Umberto Michieli, Marco Giordani, Pietro Zanuttigh, Michele Zorzi

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 14 figures, 14 tables. This paper has been submitted to IEEE. Copyright may change without notice

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15125 2022-04-06 cs.CV cs.CL cs.LG 81%

Text2Pos: Text-to-Point-Cloud Cross-Modal Localization

Manuel Kolmet, Qunjie Zhou, Aljosa Osep, Laura Leal-Taixe

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments CVPR2022 Camera Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00061 2022-03-22 cs.CV cs.CL cs.CY cs.LG 81%

Open-Domain, Content-based, Multi-modal Fact-checking of Out-of-Context Images via Online Resources

Sahar Abdelnabi, Rakibul Hasan, Mario Fritz

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments CVPR'22

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.09824 2022-03-21 cs.CV cs.LG eess.AS 81%

Cross-Modal Perceptionist: Can Face Geometry be Gleaned from Voices?

Cho-Ying Wu, Chin-Cheng Hsu, Ulrich Neumann

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、eess.AS

Comments Accepted to CVPR 2022. Project page: https://choyingw.github.io/works/Voice2Mesh/index.html. This version supersedes arXiv:2104.10299

详情

展开后加载摘要…

URL PDF HTML 收藏