arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4884 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4884 篇

2509.18132 2025-09-24 cs.AI 57%

Position Paper: Integrating Explainability and Uncertainty Estimation in Medical AI

Xiuyi Fan

机构 * Lee Kong Chian School of Medicine, College of Computing Data Science, Nanyang Technological University, Singapore

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted at the International Joint Conference on Neural Networks, IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20673 2025-09-23 cs.DC cs.AI 57%

ClusterRCA: An End-to-End Approach for Network Fault Localization and Classification for HPC System

Yongqian Sun, Xijie Pan, Xiao Xiong, Lei Tao, Jiaju Wang, Shenglin Zhang, Yuan Yuan, Yuqi Li, Kunlin Jian

机构 * National University of Defense Technology(国防科技大学) Huawei(华为) Tianjin Key Laboratory of Software Experience(天津软件体验与人机交互重点实验室) Haihe Laboratory of Information Technology Application Innovation(海河信息科技应用创新实验室)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11960 2025-09-19 cs.CL cs.LG 57%

Fast Multipole Attention: A Scalable Multilevel Attention Mechanism for Text and Images

Yanming Kang, Giang Tran, Hans De Sterck

机构 * University of Waterloo(滑铁卢大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02358 2025-09-18 cs.LG cs.AI 57%

Empowering Time Series Analysis with Foundation Models: A Comprehensive Survey

Jiexia Ye, Yongzi Yu, Weiqi Zhang, Le Wang, Jia Li, Fugee Tsung

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学) Shanghai University of Finance and Economics(上海财经大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 10 figures, 5 tables, 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02141 2025-09-16 cs.CV 57%

Bayesian Unsupervised Disentanglement of Anatomy and Geometry for Deep Groupwise Image Registration

Xinzhe Luo, Xin Wang, Linda Shapiro, Chun Yuan, Jianfeng Feng, Xiahai Zhuang

机构 * School of Data Science, Fudan University(复旦大学数据科学学院) Department of Electrical and Electronic Engineering and I-X, Imperial College London(帝国理工学院电子与电气工程系) Department of Electrical and Computer Engineering, University of Washington(华盛顿大学电气与计算机工程系) Paul G. Allen School of Computer Science and Engineering, University of Washington(华盛顿大学保罗·G·艾伦计算机科学与工程学院) Department of Radiology and Imaging Sciences, University of Utah(犹他大学放射科与成像科学系) Department of Radiology, University of Washington(华盛顿大学放射科) Institute of Science and Technology for Brain-Inspired Intellengence, Fudan University(复旦大学脑启发智能科学技术研究院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10537 2025-09-16 cs.LG cs.AI cs.DC 57%

On Using Large-Batches in Federated Learning

Sahil Tyagi

机构 * Indiana University(印第安纳大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07577 2025-09-11 cs.AI 57%

Towards explainable decision support using hybrid neural models for logistic terminal automation

Riccardo D'Elia, Alberto Termine, Francesco Flammini

机构 * University of Applied Sciences and Arts of Southern Switzerland(应用科学与艺术大学(南瑞士)) Dalle Molle Institute for Artificial Intelligence(达勒莫勒人工智能研究所) University of Florence(佛罗伦萨大学) Department of Mathematics and Computer Science Ulisse Dini(数学与计算机科学系(乌利塞·迪尼))

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07445 2025-09-10 cs.RO cs.AI 57%

Text2Touch: Tactile In-Hand Manipulation with LLM-Designed Reward Functions

Harrison Field, Max Yang, Yijiong Lin, Efi Psomopoulou, David Barton, Nathan F. Lepora

机构 * School of Computer Science, University of Bristol, 2 Bristol Robotics Laboratory 3 School of Engineering Mathematics and Technology, University of Bristol(1 计算机科学学院,布里斯托尔大学 2 布里斯托尔机器人实验室 3 工程数学与技术学院,布里斯托尔大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted at CoRL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19636 2025-08-28 cs.AI cs.NE 57%

Fitness Landscape of Large Language Model-Assisted Automated Algorithm Search

Fei Liu, Qingfu Zhang, Jialong Shi, Xialiang Tong, Kun Mao, Mingxuan Yuan

机构 * Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) School of Mathematics and Statistics, Xi’ an Jiaotong University(西安交通大学数学与统计学学院) HUAWEI Noah’s Ark Lab(华为诺亚实验室) Huawei Cloud EI Service Product Department(华为云EI服务产品部)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07410 2025-08-27 physics.ao-ph cs.AI 57%

Leveraging GNN to Enhance MEF Method in Predicting ENSO

Saghar Ganji, Ahmad Reza Labibzadeh, Alireza Hassani, Mohammad Naisipour

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 17 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06911 2025-08-22 cs.LG cs.AI 57%

MMiC: Mitigating Modality Incompleteness in Clustered Federated Learning

Lishan Yang, Wei Emma Zhang, Quan Z. Sheng, Lina Yao, Weitong Chen, Ali Shakeri

机构 * The University of Adelaide(阿德莱德大学) Macquarie University(麦考瑞大学) CSIRO’s Data 61 and The University of New South Wales(CSIRO的数据61与新南威尔士大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04742 2025-08-22 cs.CY cs.AI 57%

A Case for Specialisation in Non-Human Entities

El-Mahdi El-Mhamdi, Lê-Nguyên Hoang, Mariame Tighanimine

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments Accepted to AAAI/ACM AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17976 2025-08-21 cs.SE cs.AI 57%

The importance of visual modelling languages in generative software engineering

Roberto Rossi

机构 * Business School, University of Edinburgh(爱丁堡大学商学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 9 pages, working paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14190 2025-08-21 cs.CR cs.CL cs.LG 57%

Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text

Zixin Rao, Youssef Mohamed, Shang Liu, Zeyan Liu

机构 * University of Georgia(佐治亚大学) Egypt-Japan University of Science and Technology(埃及-日本科学技术大学) University of Louisville(路易斯维尔大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CL

Comments Securecomm 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07970 2025-08-19 cs.LG cs.AI 57%

WeChat-YATT: A Scalable, Simple, Efficient, and Production Ready Training Library

Junyu Wu, Weiming Chang, Xiaotao Liu, Guanyou He, Tingfeng Xian, Haoqiang Hong, Boqi Chen, Hongtao Tian, Tao Yang, Yunsheng Shi, Feng Lin, Ting Yao, Jiatao Xu

机构 * Tencent(腾讯)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2507.22789

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09630 2025-08-18 cs.LG cs.AI 57%

TimeMKG: Knowledge-Infused Causal Reasoning for Multivariate Time Series Modeling

Yifei Sun, Junming Liu, Yirong Chen, Xuefeng Yan, Ding Wang

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09785 2025-08-14 cs.CV 57%

DSS-Prompt: Dynamic-Static Synergistic Prompting for Few-Shot Class-Incremental Learning

Linpu He, Yanan Li, Bingze Li, Elvis Han Cui, Donghui Wang

机构 * Department of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术系) Research Center for Frontier Fundamental Studies, Zhejiang Lab(浙江实验室前沿基础研究中心) Department of Neurology, University of California, Irvine(加州大学伊市医学院神经科)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17659 2025-08-14 cs.CV 57%

See the Forest and the Trees: A Synergistic Reasoning Framework for Knowledge-Based Visual Question Answering

Junjie Wang, Yunhan Tang, Yijie Wang, Zhihao Yuan, Huan Wang, Yangfan He, Bin Li

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments We are withdrawing this preprint because it is undergoing a major revision and restructuring. We feel that the current version does not convey our core contributions and methodology with sufficient clarity and accuracy

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05884 2025-08-14 cs.IT cs.AI math.IT 57%

User-Intent-Driven Semantic Communication via Adaptive Deep Understanding

Peigen Ye, Jingpu Duan, Hongyang Du, Yulan Guo

机构 * Sun Yat-sen University(中山大学) Pengcheng Laboratory(鹏城实验室) The University of Hong Kong(香港大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments 300 *^_^* IEEE Globecom 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07329 2025-08-12 cs.LG cs.AI 57%

Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative

Tuo Zhang, Ning Li, Xin Yuan, Wenchao Xu, Quan Chen, Song Guo, Haijun Zhang

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06966 2025-08-12 cs.LG cs.AI 57%

Can Multitask Learning Enhance Model Explainability?

Hiba Najjar, Bushra Alshbib, Andreas Dengel

机构 * University of Kaiserslautern-Landau(凯撒斯劳滕-兰道大学) German Research Center for Artificial Intelligence(德国人工智能研究中心)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted at GCPR 2025, Special Track "Photogrammetry and remote sensing"

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02498 2025-08-05 cs.CL 57%

Monsoon Uprising in Bangladesh: How Facebook Shaped Collective Identity

Md Tasin Abir, Arpita Chowdhury, Ashfia Rahman

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02113 2025-08-05 cs.CV eess.IV 57%

DeflareMamba: Hierarchical Vision Mamba for Contextually Consistent Lens Flare Removal

Yihang Huang, Yuanfei Huang, Junhui Lin, Hua Huang

机构 * School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院) Engineering Research Center of Intelligent Technology and Educational Application, Ministry of Education(教育部智能技术与教育应用工程研究中心)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ACMMM 2025

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27--31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22464 2025-07-31 cs.LG cs.AI cs.MA stat.AP 57%

Towards Interpretable Renal Health Decline Forecasting via Multi-LMM Collaborative Reasoning Framework

Peng-Yi Wu, Pei-Cing Huang, Ting-Yu Chen, Chantung Ku, Ming-Yen Lin, Yihuang Kang

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20300 2025-07-31 cs.HC cs.MM 57%

Talking-to-Build: How LLM-Assisted Interface Shapes Player Performance and Experience in Minecraft

Xin Sun, Lei Wang, Yue Li, Jie Li, Massimo Poesio, Julian Frommel, Koen Hinriks, Jiahuan Pei

专题命中 其他多模态 :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11261 2025-07-29 cs.CV 57%

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition

Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He

机构 * South China University of Technology(华南理工大学) Guangdong Engineering Center for Large Model and GenAI Technology(广东省大模型与生成式人工智能技术工程中心) State Key Laboratory of Subtropical Building and Urban Science(亚热带建筑科学国家重点实验室) Ministry of Education Key Laboratory of Big Data and Intelligent Robot(教育部大数据与智能机器人重点实验室) Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information(广东省计算智能与网络信息重点实验室) Singapore Management University(新加坡国立大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07994 2025-07-29 cs.CV 57%

Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint Detection

Subhajit Maity, Ayan Kumar Bhunia, Subhadeep Koley, Pinaki Nath Chowdhury, Aneeshan Sain, Yi-Zhe Song

机构 * Department of Computer Science, University of Central Florida(中央佛罗里达大学计算机科学系) SketchX, CVSSP, University of Surrey(SketchX、CVSSP、塞雷尔大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted at ICCV 2025. Project Page: https://subhajitmaity.me/DYKp

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16815 2025-07-29 cs.CV 57%

FREE-Merging: Fourier Transform for Efficient Model Merging

Shenghe Zheng, Hongzhi Wang

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09880 2025-07-28 cs.CV 57%

Information Extraction from Unstructured data using Augmented-AI and Computer Vision

Aditya Parikh

机构 * Department of Electronics and Telecommunication Engineering(电子与电信工程系) Vishwakarma Institute of Technology(维斯瓦克arma技术学院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00286 2025-07-22 cs.HC cs.AI cs.ET 57%

"Before, I Asked My Mom, Now I Ask ChatGPT": Visual Privacy Management with Generative AI for Blind and Low-Vision People

Tanusree Sharma, Yu-Yun Tseng, Lotus Zhang, Ayae Ide, Kelly Avery Mack, Leah Findlater, Danna Gurari, Yang Wang

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Computer Science, University of Colorado(计算机科学,科罗拉多大学) Human Centered Design and Engineering, University of Washington(以人为核心的设计与工程,华盛顿大学) Information Sciences, University of Illinois at Urbana-Champaign(信息科学,伊利诺伊大学厄巴纳-香槟分校)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏