arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4868 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4868 篇

2509.05671 2025-09-09 cs.LG cs.AI cs.CR stat.ML 79%

GraMFedDHAR: Graph Based Multimodal Differentially Private Federated HAR

Labani Halder, Tanmay Sen, Sarbani Palit

机构 * Indian Statistical Institute Kolkata(印度统计研究所科塔加特)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20460 2025-08-29 cs.CL 79%

Prediction of mortality and resource utilization in critical care: a deep learning approach using multimodal electronic health records with natural language processing techniques

Yucheng Ruan, Xiang Lan, Daniel J. Tan, Hairil Rizal Abdullah, Mengling Feng

机构 * Saw Swee Hock School of Public Health(Saw Swee Hock 公共卫生学院) National University of Singapore(新加坡国立大学) Institute of Data Science(数据科学研究所) Department of Anaesthesiology(麻醉科部门)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08249 2025-08-27 cs.NI cs.AI 79%

GeNet: A Multimodal LLM-Based Co-Pilot for Network Topology and Configuration

Beni Ifland, Elad Duani, Rubin Krief, Miro Ohana, Aviram Zilberman, Andres Murillo, Ofir Manor, Ortal Lavi, Hikichi Kenji, Asaf Shabtai, Yuval Elovici, Rami Puzis

机构 * 1 Ben Gurion University of the Negev, Software

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18249 2025-08-26 cs.RO cs.CV 79%

Scene-Agnostic Traversability Labeling and Estimation via a Multimodal Self-supervised Framework

Zipeng Fang, Yanbo Wang, Lei Zhao, Weidong Chen

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17976 2025-08-26 cs.CV eess.IV 79%

Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization

Keyang Zhang, Chenqi Kong, Hui Liu, Bo Ding, Xinghao Jiang, Haoliang Li

机构 * Department of Electrical Engineering, City University of Hong Kong(香港城市大学电子工程系) Rapid-Rich Object Search (ROSE) Lab, School of Electrical and Electronic Engineering, Nanyang Technology University(南洋理工大学电子与电气工程学院快速丰富对象搜索(ROSE)实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 其他多模态 :MLLM(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13357 2025-08-26 cs.CL 79%

Adaptive Linguistic Prompting (ALP) Enhances Phishing Webpage Detection in Multimodal Large Language Models

Atharva Bhargude, Ishan Gonehal, Dave Yoon, Kaustubh Vinnakota, Chandler Haney, Aaron Sandoval, Kevin Zhu

机构 * Algoverse AI Research(Algoverse AI研究院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

Comments Published at ACL 2025 SRW, 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16296 2025-08-26 cs.CV cs.GR eess.IV 79%

LiDAR-3DGS: LiDAR Reinforced 3D Gaussian Splatting for Multimodal Radiance Field Rendering

Hansol Lim, Hanbeom Chang, Jongseong Brad Choi, Chul Min Yeum

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16748 2025-08-26 cs.LG cs.AI 79%

FAIRWELL: Fair Multimodal Self-Supervised Learning for Wellbeing Prediction

Jiaee Cheong, Abtin Mogharabin, Paul Liang, Hatice Gunes, Sinan Kalkan

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.00528 2025-08-08 cs.CL 79%

Unimodal Intermediate Training for Multimodal Meme Sentiment Classification

Muzhaffar Hazman, Susan McKeever, Josephine Griffith

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted for Publication at RANLP2023

Journal ref https://aclanthology.org/2023.ranlp-1.55/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02243 2025-08-05 cs.CV cs.IR 79%

I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking

Ziyan Liu, Junwen Li, Kaiwen Li, Tong Ruan, Chao Wang, Xinyan He, Zongyu Wang, Xuezhi Cao, Jingping Liu

机构 * East China University of Science and Technology(东华大学) South China University of Technology(华南理工大学) Shanghai University(上海大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 6 figures, accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01465 2025-08-05 cs.CV 79%

EfficientGFormer: Multimodal Brain Tumor Segmentation via Pruned Graph-Augmented Transformer

Fatemeh Ziaeetabar

机构 * School of Mathematics, Statistics and Computer Science, College of Science, University of Tehran(数学、统计与计算机科学学院,科学学院,德黑兰大学) University of Tehran(德黑兰大学)

专题命中 其他多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00665 2025-08-04 cs.AI cs.HC cs.LG 79%

Transparent Adaptive Learning via Data-Centric Multimodal Explainable AI

Maryam Mosleh, Marie Devlin, Ellis Solaiman

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21378 2025-07-30 cs.HC cs.AI 79%

ProMemAssist: Exploring Timely Proactive Assistance Through Working Memory Modeling in Multi-Modal Wearable Devices

Kevin Pu, Ting Zhang, Naveen Sendhilnathan, Sebastian Freitag, Raj Sodhi, Tanya Jonker

机构 * University of Toronto(多伦多大学) Meta Reality Labs(Meta现实实验室)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted to UIST'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15652 2025-07-22 cs.CV 79%

Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models

Haoran Zhou, Zihan Zhang, Hao Chen

机构 * Southeast University(东南大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20090 2025-07-22 cs.CV 79%

Transfer Attack for Bad and Good: Explain and Boost Adversarial Transferability across Multimodal Large Language Models

Hao Cheng, Erjia Xiao, Jiayan Yang, Jinhao Duan, Yichi Wang, Jiahang Cao, Qiang Zhang, Le Yang, Kaidi Xu, Jindong Gu, Renjing Xu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Oxford(牛津大学) Drexel University(德雷塞尔大学) Beijing University of Technology(北京理工大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Xi’an Jiaotong University(西安交通大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments This paper is accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11001 2025-07-16 cs.RO cs.CV 79%

Learning to Tune Like an Expert: Interpretable and Scene-Aware Navigation via MLLM Reasoning and CVAE-Based Adaptation

Yanbo Wang, Zipeng Fang, Lei Zhao, Weidong Chen

专题命中 其他多模态 :MLLM(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06143 2025-07-14 physics.ed-ph cs.AI 79%

Multilingual Performance of a Multimodal Artificial Intelligence System on Multisubject Physics Concept Inventories

Gerd Kortemeyer, Marina Babayeva, Giulia Polverini, Ralf Widenhorn, Bor Gregorcic

机构 * AI Center, ETH Zurich(ETH Zurich人工智能中心) Michigan State University(密歇根州立大学) Charles University(查理大学) Uppsala University(乌普萨拉大学) Portland State University(波特兰州立大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Journal ref Phys. Rev. Phys. Educ. Res. 21, 020101 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05983 2025-07-09 cs.LG cs.AI 79%

Longitudinal Ensemble Integration for sequential classification with multimodal data

Aviad Susman, Rupak Krishnamurthy, Yan Chak Li, Mohammad Olaimat, Serdar Bozdag, Bino Varghese, Nasim Sheikh-Bahaei, Gaurav Pandey

机构 * Department of Genetics and Genomic Sciences, Icahn School of Medicine at Mount Sinai(遗传学与基因组科学系,伊坎医学院-西奈山医院) Department of Computer Science and Engineering, University of North Texas(计算机科学与工程系,北德克萨斯大学) Department of Radiology, Keck School of Medicine, University of Southern California(放射学系,凯克医学院,南加州大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to IEEE ICDH 2025. This is the author's accepted manuscript (AAM). The final version will appear in the IEEE ICDH 2025 proceedings on IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19593 2025-06-25 cs.CV cs.SY eess.SY 79%

Implementing blind navigation through multi-modal sensing and gait guidance

Feifan Yan, Tianle Zeng, Meixi He

机构 * China University of Mining and Technology - Beijing(中国矿业大学(北京))

专题命中 其他多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18716 2025-06-24 cs.LG cs.CL 79%

Multi-modal Anchor Gated Transformer with Knowledge Distillation for Emotion Recognition in Conversation

Jie Li, Shifei Ding, Lili Guo, Xuan Li

机构 * School of Computer Science and Technology, China University of Mining and Technology(计算机科学与技术学院,中国矿业大学) Mine Digitization Engineering Research Center of Ministry of Education, China University of Mining and Technology(教育部矿山数字化工程研究中心,中国矿业大学)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CL

Comments This paper has been accepted by IJCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17248 2025-06-24 cs.LG cs.AI stat.ML 79%

Efficient Quantification of Multimodal Interaction at Sample Level

Zequn Yang, Hongfa Wang, Di Hu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China(中国人民大学语言学院) Tencent Data Platform(腾讯数据平台) Tsinghua University, Beijing, China(清华大学) Beijing Key Laboratory of Research on Large Models(北京大型模型研究关键实验室) Engineering Research Center of Next-Generation Intelligent Search(下一代智能搜索工程研究中心)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07575 2025-06-10 cs.CV cs.LG 79%

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models

Ruiyang Zhang, Hu Zhang, Hao Fei, Zhedong Zheng

机构 * FST and ICI, University of Macau, China(澳门大学先进科技研究院和国际研究中心) CSIRO Data61, Australia(澳大利亚CSIRO数据61研究所) National University of Singapore(新加坡国立大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments Project page: https://uncertainty-o.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07055 2025-06-10 cs.CV 79%

A Layered Self-Supervised Knowledge Distillation Framework for Efficient Multimodal Learning on the Edge

Tarique Dahri, Zulfiqar Ali Memon, Zhenyu Yu, Mohd. Yamani Idna Idris, Sheheryar Khan, Sadiq Ahmad, Maged Shoman, Saddam Aziz, Rizwan Qureshi

机构 * Fast School of Computing, National University of Computer and Emerging Sciences(国家计算机与新兴科学大学快速计算学院) Universiti Malaya(马来亚大学) School of Professional Education and Executive Development, The Hong Kong Polytechnic University(香港理工大学专业教育与执行发展学院) COMSATS University Islamabad, Wah Campus(伊斯兰堡COMSATS大学瓦校区) Intelligent Transportation Systems University of Tennessee-Oak Ridge Innovation Institute’s Energy Storage and Transportation Convergent Research Initiative(智能交通系统田纳西大学奥克岭创新研究所能源存储与交通融合研究计划) Center for research in Computer Vision, University of Central Florida(计算机视觉研究中心,佛罗里达中央大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07441 2025-06-10 cs.HC cs.AI 79%

SensPS: Sensing Personal Space Comfortable Distance between Human-Human Using Multimodal Sensors

Ko Watanabe, Nico Förster, Shoya Ishimaru

机构 * German Research Center for Artifcial Intelligence(德国人工智能研究中心) RPTU Kaiserslautern-Landau(科隆-拉登堡应用技术大学) Osaka Metropolitan University(大阪市立大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11726 2025-06-03 cs.CL 79%

Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures

Shun Inadumi, Nobuhiro Ueda, Koichiro Yoshino

机构 * Nara Institute of Science and Technology(奈良科学技術研究所) Guardian Robot Project(守護機器人計劃) RIKEN(理化学研究所) Kyoto University(京都大學) Institute of Science Tokyo(東京科學研究院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

Comments ACL2025 main. Code available at https://github.com/SInadumi/mmrr

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24427 2025-06-02 cs.CL 79%

Donate or Create? Comparing Data Collection Strategies for Emotion-labeled Multimodal Social Media Posts

Christopher Bagdon, Aidan Combs, Carina Silberer, Roman Klinger

机构 * Fundamentals of Natural Language Processing, University of Bamberg(自然语言处理基础,巴姆伯格大学) Department of Sociology, The Ohio State University(社会学系,俄亥俄州立大学) Institut für Maschinelle Sprachverarbeitung, University of Stuttgart(机械语言处理研究所,斯图加特大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

Comments Published at ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20960 2025-05-30 cs.CL cs.CY cs.LG 79%

Multi-Modal Framing Analysis of News

Arnav Arora, Srishti Yadav, Maria Antoniak, Serge Belongie, Isabelle Augenstein

机构 * University of Copenhagen(哥本哈根大学)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21792 2025-05-29 cs.LG cs.AI 79%

Multimodal Federated Learning: A Survey through the Lens of Different FL Paradigms

Yuanzhe Peng, Jieming Bian, Lei Wang, Yin Huang, Jie Xu

机构 * University of Florida(佛罗里达大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21553 2025-05-29 cs.NI cs.AI cs.LG 79%

MetaSTNet: Multimodal Meta-learning for Cellular Traffic Conformal Prediction

Hui Ma, Kai Yang

机构 * College of Electronic and Information Engineering, Tongji University(电子信息工程学院,同济大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21419 2025-05-29 cs.AI cs.OS 79%

Diagnosing and Resolving Cloud Platform Instability with Multi-modal RAG LLMs

Yifan Wang, Kenneth P. Birman

机构 * Computer Science Department, Cornell University(康奈尔大学计算机科学系)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Published in EuroMLSys2025

Journal ref 2025, Association for Computing Machinery, Proceedings of the 5th Workshop on Machine Learning and Systems, 139-147, 9, series EuroMLSys '25

详情

展开后加载摘要…

URL PDF HTML 收藏