arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4884 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4884 篇

2409.05275 2024-09-10 cs.CL 57%

RexUniNLU: Recursive Method with Explicit Schema Instructor for Universal NLU

Chengyuan Liu, Shihang Wang, Fubang Zhao, Kun Kuang, Yangyang Kang, Weiming Lu, Changlong Sun, Fei Wu

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL

Comments arXiv admin note: substantial text overlap with arXiv:2304.14770

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18967 2024-08-29 cs.CV 57%

Structural Attention: Rethinking Transformer for Unpaired Medical Image Synthesis

Vu Minh Hieu Phan, Yutong Xie, Bowen Zhang, Yuankai Qi, Zhibin Liao, Antonios Perperidis, Son Lam Phung, Johan W. Verjans, Minh-Son To

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments MICCAI version before camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09382 2024-08-20 cs.HC cs.AI cs.ET 57%

VRCopilot: Authoring 3D Layouts with Generative AI Models in VR

Lei Zhang, Jin Pan, Jacob Gettig, Steve Oney, Anhong Guo

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments UIST 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08321 2024-08-19 cs.HC cs.CV 57%

Can ChatGPT assist visually impaired people with micro-navigation?

Junxian He, Shrinivas Pundlik, Gang Luo

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04868 2024-08-12 cs.CV 57%

ChatGPT Meets Iris Biometrics

Parisa Farmanifard, Arun Ross

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Published at IJCB 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02464 2024-08-06 cs.CV 57%

Fairness and Bias Mitigation in Computer Vision: A Survey

Sepehr Dehdashtian, Ruozhen He, Yi Li, Guha Balakrishnan, Nuno Vasconcelos, Vicente Ordonez, Vishnu Naresh Boddeti

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments 20 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11129 2024-08-06 cs.CV 57%

Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales

Minghe Gao, Shuang Chen, Liang Pang, Yuan Yao, Jisheng Dang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Yueting Zhuang, Tat-Seng Chua

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19518 2024-07-30 cs.RO cs.CV 57%

Solving Short-Term Relocalization Problems In Monocular Keyframe Visual SLAM Using Spatial And Semantic Data

Azmyin Md. Kamal, Nenyi K. N. Dadson, Donovan Gegg, Corina Barbalata

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments 8 pages, Keywords: VSLAM, Localization, Semantics. Presented in 2024 IEEE/ASME International Conference on Advanced Intelligent Mechatronics (AIM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19492 2024-07-30 cs.HC cs.AI cs.ET 57%

Heads Up eXperience (HUX): Always-On AI Companion for Human Computer Environment Interaction

Sukanth K, Sudhiksha Kandavel Rajan, Rajashekhar V S, Gowdham Prabhakar

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments 48 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04418 2024-07-25 cs.HC cs.AI cs.LG 57%

Enabling On-Device LLMs Personalization with Smartphone Sensing

Shiquan Zhang, Ying Ma, Le Fang, Hong Jia, Simon D'Alfonso, Vassilis Kostakos

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 5 pages, 3 figures, conference demo paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19580 2024-07-24 cs.CV 57%

OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation

Zhenyu Wang, Yali Li, Taichi Liu, Hengshuang Zhao, Shengjin Wang

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ECCV2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17591 2024-07-23 cs.CV 57%

DocParseNet: Advanced Semantic Segmentation and OCR Embeddings for Efficient Scanned Document Annotation

Ahmad Mohammadshirazi, Ali Nosrati Firoozsalari, Mengxi Zhou, Dheeraj Kulshrestha, Rajiv Ramnath

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13408 2024-07-19 cs.HC cs.AI 57%

DISCOVER: A Data-driven Interactive System for Comprehensive Observation, Visualization, and ExploRation of Human Behaviour

Dominik Schiller, Tobias Hallmen, Daksitha Withanage Don, Elisabeth André, Tobias Baur

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12339 2024-07-18 cs.CV 57%

Exploring Deeper! Segment Anything Model with Depth Perception for Camouflaged Object Detection

Zhenni Yu, Xiaoqin Zhang, Li Zhao, Yi Bin, Guobao Xiao

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10743 2024-07-16 cs.RO cs.AI 57%

Scaling 3D Reasoning with LMMs to Large Robot Mission Environments Using Datagraphs

W. J. Meijer, A. C. Kemmeren, E. H. J. Riemens, J. E. Fransman, M. van Bekkum, G. J. Burghouts, J. D. van Mil

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted to the RSS Workshop on Semantics for Robotics: From Environment Understanding and Reasoning to Safe Interaction 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09913 2024-07-16 cs.CV 57%

Emotion Detection through Body Gesture and Face

Haoyang Liu

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments 25 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04362 2024-07-08 cs.CV cs.HC 57%

Towards Context-aware Support for Color Vision Deficiency: An Approach Integrating LLM and AR

Shogo Morita, Yan Zhang, Takuto Yamauchi, Sinan Chen, Jialong Li, Kenji Tei

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00579 2024-07-08 cs.IR cs.AI 57%

A Review of Modern Recommender Systems Using Generative Models (Gen-RecSys)

Yashar Deldjoo, Zhankui He, Julian McAuley, Anton Korikov, Scott Sanner, Arnau Ramisa, René Vidal, Maheswaran Sathiamoorthy, Atoosa Kasirzadeh, Silvia Milano

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments This survey accompanies a tutorial presented at ACM KDD'24

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19302 2024-06-28 cs.CV cs.LG 57%

Mapping Land Naturalness from Sentinel-2 using Deep Contextual and Geographical Priors

Burak Ekim, Michael Schmitt

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments 6 pages, 3 figures, ICLR 2024 Tackling Climate Change with Machine Learning Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19297 2024-06-28 cs.CV 57%

Enhancing Continual Learning in Visual Question Answering with Modality-Aware Feature Distillation

Malvina Nikandrou, Georgios Pantazopoulos, Ioannis Konstas, Alessandro Suglia

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19054 2024-06-28 cs.LG cs.AI cs.HC 57%

A look under the hood of the Interactive Deep Learning Enterprise (No-IDLE)

Daniel Sonntag, Michael Barz, Thiago Gouvêa

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments DFKI Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01720 2024-06-26 cs.LG cs.AI 57%

Perceiver-based CDF Modeling for Time Series Forecasting

Cat P. Le, Chris Cannella, Ali Hasan, Yuting Ng, Vahid Tarokh

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted in Winter Simulation Conference 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16455 2024-06-25 cs.AI 57%

Guardrails for avoiding harmful medical product recommendations and off-label promotion in generative AI models

Daniel Lopez-Martinez

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments CVPR 2024 Responsible Generative AI (ReGenAI) workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09940 2024-06-17 q-bio.NC cs.AI cs.NE 57%

Implementing engrams from a machine learning perspective: XOR as a basic motif

Jesus Marco de Lucas, Maria Peña Fernandez, Lara Lloret Iglesias

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 9 pages, short comment

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04334 2024-06-07 cs.CV 57%

DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

Lingchen Meng, Jianwei Yang, Rui Tian, Xiyang Dai, Zuxuan Wu, Jianfeng Gao, Yu-Gang Jiang

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Project Page: https://deepstack-vl.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07839 2024-06-04 cs.LG cs.AI stat.ML 57%

Constrained Exploration via Reflected Replica Exchange Stochastic Gradient Langevin Dynamics

Haoyang Zheng, Hengrong Du, Qi Feng, Wei Deng, Guang Lin

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments 28 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12923 2024-05-22 cs.IR cs.AI cs.HC 57%

Panmodal Information Interaction

Chirag Shah, Ryen W. White

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12759 2024-05-22 cs.CV 57%

Cross-spectral Gated-RGB Stereo Depth Estimation

Samuel Brucker, Stefanie Walz, Mario Bijelic, Felix Heide

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10890 2024-05-20 astro-ph.IM astro-ph.GA cs.AI 57%

A Versatile Framework for Analyzing Galaxy Image Data by Implanting Human-in-the-loop on a Large Vision Model

Mingxiang Fu, Yu Song, Jiameng Lv, Liang Cao, Peng Jia, Nan Li, Xiangru Li, Jifeng Liu, A-Li Luo, Bo Qiu, Shiyin Shen, Liangping Tu, Lili Wang, Shoulin Wei, Haifeng Yang, Zhenping Yi, Zhiqiang Zou

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 26 pages, 10 figures, to be published on Chinese Physics C

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.06299 2024-05-13 eess.SP cs.AI 57%

Cross-domain Learning Framework for Tracking Users in RIS-aided Multi-band ISAC Systems with Sparse Labeled Data

Jingzhi Hu, Dusit Niyato, Jun Luo

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏