arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9119 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9119 篇

2304.03897 2023-10-03 cs.CL cs.CV 81%

Factify 2: A Multimodal Fake News and Satire News Dataset

S Suryavardan, Shreyash Mishra, Parth Patwa, Megha Chakraborty, Anku Rani, Aishwarya Reganti, Aman Chadha, Amitava Das, Amit Sheth, Manoj Chinnakotla, Asif Ekbal, Srijan Kumar

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL

Comments Defactify2 @AAAI2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.06495 2023-09-14 cs.CL cs.AI cs.PF 81%

AGIBench: A Multi-granularity, Multimodal, Human-referenced, Auto-scoring Benchmark for Large Language Models

Fei Tang, Wanling Gao, Luzhou Peng, Jianfeng Zhan

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01383 2023-09-06 cs.CV cs.AI 81%

LoRA-like Calibration for Multimodal Deception Detection using ATSFace Data

Shun-Wen Hsiao, Cheng-Yuan Sun

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12156 2023-08-24 cs.CV cs.AI 81%

Multimodal Latent Emotion Recognition from Micro-expression and Physiological Signals

Liangfei Zhang, Yifei Qian, Ognjen Arandjelovic, Anthony Zhu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07686 2023-08-16 cs.CV cs.AI 81%

Boosting Multi-modal Model Performance with Adaptive Gradient Modulation

Hong Li, Xingyu Li, Pengbo Hu, Yinuo Lei, Chunxiao Li, Yi Zhou

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02508 2023-08-08 eess.IV cs.AI cs.CV 81%

A Multimodal Supervised Machine Learning Approach for Satellite-based Wildfire Identification in Europe

Angelica Urbanelli, Luca Barco, Edoardo Arnaudo, Claudio Rossi

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at IGARSS 2023, short paper (4 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.00628 2023-08-08 cs.CV cs.AI cs.LG 81%

Human-M3: A Multi-view Multi-modal Dataset for 3D Human Pose Estimation in Outdoor Scenes

Bohao Fan, Siqi Wang, Wenxuan Guo, Wenzhao Zheng, Jianjiang Feng, Jie Zhou

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Code and data will be released on https://github.com/soullessrobot/Human-M3-Dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.16125 2023-08-03 cs.CL cs.CV 81%

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, Ying Shan

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Technical Report; Project released at: https://github.com/AILab-CVC/SEED-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.11952 2023-07-25 cs.CV cs.AI 81%

Pathology-and-genomics Multimodal Transformer for Survival Outcome Prediction

Kexin Ding, Mu Zhou, Dimitris N. Metaxas, Shaoting Zhang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to MICCAI2023 (Top14%)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02716 2023-07-07 cs.CL cs.CV 81%

CFSum: A Coarse-to-Fine Contribution Network for Multimodal Summarization

Min Xiao, Junnan Zhu, Haitao Lin, Yu Zhou, Chengqing Zong

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments acl2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02499 2023-07-07 cs.CL cs.AI 81%

mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Yuhao Dan, Chenlin Zhao, Guohai Xu, Chenliang Li, Junfeng Tian, Qian Qi, Ji Zhang, Fei Huang

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18326 2023-07-04 cs.CV cs.AI 81%

BigVideo: A Large-scale Video Subtitle Translation Dataset for Multimodal Machine Translation

Liyan Kang, Luyang Huang, Ningxin Peng, Peihao Zhu, Zewei Sun, Shanbo Cheng, Mingxuan Wang, Degen Huang, Jinsong Su

专题命中 多模态评测 :multimodal(title);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ACL 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15977 2023-06-29 cs.CV cs.AI 81%

A Dimensional Structure based Knowledge Distillation Method for Cross-Modal Learning

Lingyu Si, Hongwei Dong, Wenwen Qiang, Junzhi Yu, Wenlong Zhai, Changwen Zheng, Fanjiang Xu, Fuchun Sun

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13968 2023-06-27 cs.CL cs.AI 81%

Fusing Multimodal Signals on Hyper-complex Space for Extreme Abstractive Text Summarization (TL;DR) of Scientific Contents

Yash Kumar Atri, Vikram Goyal, Tanmoy Chakraborty

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to ADS-SIGKDD2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.05493 2023-06-12 cs.CV cs.AI cs.LG 81%

Multi-Modal Classifiers for Open-Vocabulary Object Detection

Prannay Kaul, Weidi Xie, Andrew Zisserman

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments ICML 2023, project page: https://www.robots.ox.ac.uk/vgg/research/mm-ovod/

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04738 2023-06-09 cs.CV cs.AI 81%

MultiEarth 2023 -- Multimodal Learning for Earth and Environment Workshop and Challenge

Miriam Cha, Gregory Angelides, Mark Hamilton, Andy Soszynski, Brandon Swenson, Nathaniel Maidel, Phillip Isola, Taylor Perron, Bill Freeman

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04021 2023-06-08 cs.CV cs.AI cs.LG cs.RO 81%

Energy-Based Models for Cross-Modal Localization using Convolutional Transformers

Alan Wu, Michael S. Ryoo

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments ICRA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18641 2023-05-31 cs.CL cs.CV 81%

Enhanced Chart Understanding in Vision and Language Task via Cross-modal Pre-training on Plot Table Pairs

Mingyang Zhou, Yi R. Fung, Long Chen, Christopher Thomas, Heng Ji, Shih-Fu Chang

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by Findings of ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07044 2023-05-30 cs.CV cs.AI 81%

SSL4EO-S12: A Large-Scale Multi-Modal, Multi-Temporal Dataset for Self-Supervised Learning in Earth Observation

Yi Wang, Nassim Ait Ali Braham, Zhitong Xiong, Chenying Liu, Conrad M Albrecht, Xiao Xiang Zhu

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by IEEE Geoscience and Remote Sensing Magazine. 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17388 2023-05-30 cs.CL cs.CV 81%

MPCHAT: Towards Multimodal Persona-Grounded Conversation

Jaewoo Ahn, Yeda Song, Sangdoo Yun, Gunhee Kim

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14142 2023-05-24 cs.CV cs.AI 81%

A multimodal method based on cross-attention and convolution for postoperative infection diagnosis

Xianjie Liu, Hongwei Shi

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07167 2023-05-15 cs.CV cs.CL cs.LG eess.IV 81%

OneCAD: One Classifier for All image Datasets using multimodal learning

Shakti N. Wadekar, Eugenio Culurciello

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.07672 2023-05-05 cs.CV cs.CL 81%

Summary-Oriented Vision Modeling for Multimodal Abstractive Summarization

Yunlong Liang, Fandong Meng, Jinan Xu, Jiaan Wang, Yufeng Chen, Jie Zhou

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at ACL 2023 as a long paper of the main conference. Data and Code: https://github.com/XL2248/SOV-MAS

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14501 2023-05-01 cs.CV cs.AI cs.HC 81%

Read My Mind: A Multi-Modal Dataset for Human Belief Prediction

Jiafei Duan, Samson Yu, Nicholas Tan, Yi Ru Wang, Cheston Tan

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to ICRA 2023 Communicating Robot Learning Across Human-Robot Interaction Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10794 2023-04-28 cs.LG cs.CL cs.MM q-bio.QM 81%

PheME: A deep ensemble framework for improving phenotype prediction from multi-modal data

Shenghan Zhang, Haoxuan Li, Ruixiang Tang, Sirui Ding, Laila Rasmy, Degui Zhi, Na Zou, Xia Hu

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00422 2023-04-07 cs.IR cs.AI cs.CV 81%

Personalized Showcases: Generating Multi-Modal Explanations for Recommendations

An Yan, Zhankui He, Jiacheng Li, Tianyang Zhang, Julian McAuley

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to SIGIR-23, with additional dataset details. Code and data: https://github.com/zzxslp/Gest

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08646 2023-03-31 cs.LG cs.AI cs.CV 81%

AutoFed: Heterogeneity-Aware Federated Multimodal Learning for Robust Autonomous Driving

Tianyue Zheng, Ang Li, Zhe Chen, Hongbo Wang, Jun Luo

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12734 2023-03-23 cs.CV cs.CL cs.LG 81%

MultiModal Bias: Introducing a Framework for Stereotypical Bias Assessment beyond Gender and Race in Vision Language Models

Sepehr Janghorbani, Gerard de Melo

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10971 2023-03-21 cs.CV cs.AI cs.CG 81%

Self-Supervised Learning for Multimodal Non-Rigid 3D Shape Matching

Dongliang Cao, Florian Bernard

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10372 2023-03-21 cs.CV cs.MM eess.IV 81%

Just Noticeable Visual Redundancy Forecasting: A Deep Multimodal-driven Approach

Wuyuan Xie, Shukang Wang, Sukun Tian, Lirong Huang, Ye Liu, Miaohui Wang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

Journal ref AAAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏