arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4882 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4882 篇

2403.09063 2024-03-15 cs.CV cs.AI 62%

Distribution and Depth-Aware Transformers for 3D Human Mesh Recovery

Jerrin Bright, Bavesh Balaji, Harish Prakash, Yuhao Chen, David A Clausi, John Zelek

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Submitted to 21st International Conference on Robots and Vision (CRV'24), Guelph, Ontario, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01118 2024-03-05 cs.CV cs.AI 62%

Adversarial Testing for Visual Grounding via Image-Aware Property Reduction

Zhiyuan Chang, Mingyang Li, Junjie Wang, Cheng Li, Boyu Wu, Fanjiang Xu, Qing Wang

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 14pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13641 2024-02-28 cs.CV cs.AI cs.CY cs.LG 62%

How Good is ChatGPT at Face Biometrics? A First Look into Recognition, Soft Biometrics, and Explainability

Ivan DeAndres-Tame, Ruben Tolosana, Ruben Vera-Rodriguez, Aythami Morales, Julian Fierrez, Javier Ortega-Garcia

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Journal ref IEEE Access, February 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16318 2024-02-27 cs.CV cs.MM 62%

Gradient-Guided Modality Decoupling for Missing-Modality Robustness

Hao Wang, Shengda Luo, Guosheng Hu, Jianguo Zhang

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.MM

Comments AAAI24

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10580 2024-02-19 cs.CV cs.AI cs.LG 62%

Efficient Multi-task Uncertainties for Joint Semantic Segmentation and Monocular Depth Estimation

Steven Landgraf, Markus Hillemann, Theodor Kapler, Markus Ulrich

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 17 pages, 5 figures, 10 tables, submitted to peer-reviewed journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06792 2024-01-18 cs.CL cs.AI 62%

LightHouse: A Survey of AGI Hallucination

Feng Wang

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02523 2024-01-08 cs.CV cs.AI cs.LG cs.SY eess.SY 62%

Image-based Deep Learning for Smart Digital Twins: a Review

Md Ruman Islam, Mahadevan Subramaniam, Pei-Chi Huang

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 2 figures, and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.10249 2023-12-27 cs.CL cs.AI 62%

Large Generative AI Models for Telecom: The Next Big Thing?

Lina Bariah, Qiyang Zhao, Hang Zou, Yu Tian, Faouzi Bader, Merouane Debbah

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07638 2023-12-14 cs.HC cs.AI cs.CV cs.RO 62%

Teaching Unknown Objects by Leveraging Human Gaze and Augmented Reality in Human-Robot Interaction

Daniel Weber

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments PhD Thesis, University of Tübingen

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06428 2023-12-12 cs.CV cs.AI cs.IR cs.LG 62%

VisionTraj: A Noise-Robust Trajectory Recovery Framework based on Large-scale Camera Network

Zhishuai Li, Ziyue Li, Xiaoru Hu, Guoqing Du, Yunhao Nie, Feng Zhu, Lei Bai, Rui Zhao

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08622 2023-11-16 cs.CV cs.CL cs.LG 62%

Multiple-Question Multiple-Answer Text-VQA

Peng Tang, Srikar Appalaraju, R. Manmatha, Yusheng Xie, Vijay Mahadevan

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04262 2023-11-09 cs.CV cs.AI cs.DL cs.LG 62%

ETDPC: A Multimodality Framework for Classifying Pages in Electronic Theses and Dissertations

Muntabir Hasan Choudhury, Lamia Salsabil, William A. Ingram, Edward A. Fox, Jian Wu

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 3 figures, accepted to Innovative Applications of Artificial Intelligence (IAAI-24)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13343 2023-10-23 cs.CL cs.AI 62%

Challenges and Contributing Factors in the Utilization of Large Language Models (LLMs)

Xiaoliang Chen, Liangbin Li, Le Chang, Yunhe Huang, Yuxuan Zhao, Yuxiao Zhang, Dinuo Li

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11751 2023-10-17 cs.CV cs.AI cs.CR cs.LG 62%

How Robust is Google's Bard to Adversarial Image Attacks?

Yinpeng Dong, Huanran Chen, Jiawei Chen, Zhengwei Fang, Xiao Yang, Yichi Zhang, Yu Tian, Hang Su, Jun Zhu

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09979 2023-09-29 cs.RO cs.AI cs.CV cs.LG 62%

General In-Hand Object Rotation with Vision and Touch

Haozhi Qi, Brent Yi, Sudharshan Suresh, Mike Lambeta, Yi Ma, Roberto Calandra, Jitendra Malik

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments CoRL 2023; Website: https://haozhi.io/rotateit/

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04643 2023-09-18 cs.LG cs.AI cs.CV q-bio.NC 62%

Critical Learning Periods for Multisensory Integration in Deep Networks

Michael Kleinman, Alessandro Achille, Stefano Soatto

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments CVPR 2023 (Highlighted Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15016 2023-08-31 cs.CV cs.AI cs.LG 62%

How Good is Google Bard's Visual Understanding? An Empirical Study on Open Challenges

Haotong Qin, Ge-Peng Ji, Salman Khan, Deng-Ping Fan, Fahad Shahbaz Khan, Luc Van Gool

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Journal ref Machine Intelligence Research. 20(5), October 2023, 605-613

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11513 2023-08-23 cs.CV cs.AI cs.LG 62%

TrackFlow: Multi-Object Tracking with Normalizing Flows

Gianluca Mancusi, Aniello Panariello, Angelo Porrello, Matteo Fabbri, Simone Calderara, Rita Cucchiara

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07957 2023-08-21 cs.CV cs.AI cs.LG cs.RO 62%

Hidden Biases of End-to-End Driving Models

Bernhard Jaeger, Kashyap Chitta, Andreas Geiger

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at ICCV 2023. Camera ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.12662 2023-07-25 cs.CL cs.AI 62%

XTQA: Span-Level Explanations of the Textbook Question Answering

Jie Ma, Qi Chai, Jun Liu, Qingyu Yin, Pinghui Wang, Qinghua Zheng

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted by IEEE TNNLS

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16410 2023-06-29 cs.CL cs.CV 62%

Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language

William Berrios, Gautam Mittal, Tristan Thrush, Douwe Kiela, Amanpreet Singh

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.06144 2023-06-22 cs.CV cs.CL cs.IR cs.LG 62%

Combining Visual and Textual Features for Semantic Segmentation of Historical Newspapers

Raphaël Barman, Maud Ehrmann, Simon Clematide, Sofia Ares Oliveira, Frédéric Kaplan

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Journal ref Journal of Data Mining & Digital Humanities, HistoInformatics, HistoInformatics (January 19, 2021) jdmdh:6107

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08213 2023-06-16 cs.CV cs.AI 62%

SMC-UDA: Structure-Modal Constraint for Unsupervised Cross-Domain Renal Segmentation

Zhusi Zhong, Jie Li, Lulu Bi, Li Yang, Ihab Kamel, Rama Chellappa, Xinbo Gao, Harrison Bai, Zhicheng Jiao

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06379 2023-06-09 cs.CV cs.CL 62%

One does not fit all! On the Complementarity of Vision Encoders for Vision and Language Tasks

Gregor Geigle, Chen Cecilia Liu, Jonas Pfeiffer, Iryna Gurevych

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Repl4NLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14167 2023-05-25 cs.CV cs.AI 62%

DetGPT: Detect What You Need via Reasoning

Renjie Pi, Jiahui Gao, Shizhe Diao, Rui Pan, Hanze Dong, Jipeng Zhang, Lewei Yao, Jianhua Han, Hang Xu, Lingpeng Kong, Tong Zhang

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06472 2023-05-15 cs.LG cs.AI cs.CL 62%

ChatGPT-Like Large-Scale Foundation Models for Prognostics and Health Management: A Survey and Roadmaps

Yan-Fu Li, Huan Wang, Muxia Sun

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments 55 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06407 2023-05-12 cs.CV cs.AI 62%

Combo of Thinking and Observing for Outside-Knowledge VQA

Qingyi Si, Yuchen Mo, Zheng Lin, Huishan Ji, Weiping Wang

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments ACL-23, Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06314 2023-05-11 cs.CV cs.AI cs.LG 62%

Scan2LoD3: Reconstructing semantic 3D building models at LoD3 using ray casting and Bayesian networks

Olaf Wysocki, Yan Xia, Magdalena Wysocki, Eleonora Grilli, Ludwig Hoegner, Daniel Cremers, Uwe Stilla

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted for Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04460 2023-05-09 cs.CL cs.AI 62%

Language Independent Neuro-Symbolic Semantic Parsing for Form Understanding

Bhanu Prakash Voutharoja, Lizhen Qu, Fatemeh Shiri

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted to ICDAR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14936 2023-05-01 cs.CL cs.AI 62%

Information Redundancy and Biases in Public Document Information Extraction Benchmarks

Seif Laatiri, Pirashanth Ratnamogan, Joel Tang, Laurent Lam, William Vanhuffel, Fabien Caspani

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments 15 pages, ICDAR 2023 (17th International Conference on Document Analysis and Recognition)

详情

展开后加载摘要…

URL PDF HTML 收藏