arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4878 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4878 篇

2407.15086 2024-07-23 cs.RO cs.AI 70%

MaxMI: A Maximal Mutual Information Criterion for Manipulation Concept Discovery

Pei Zhou, Yanchao Yang

专题命中 其他多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16996 2024-05-28 cs.CV 70%

Mitigating Noisy Correspondence by Geometrical Structure Consistency Learning

Zihua Zhao, Mengxi Chen, Tianjie Dai, Jiangchao Yao, Bo han, Ya Zhang, Yanfeng Wang

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 10 pages, 5 figures, received by IEEE/CVF Computer Science and Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08272 2024-05-15 cs.CV 70%

VS-Assistant: Versatile Surgery Assistant on the Demand of Surgeons

Zhen Chen, Xingjian Luo, Jinlin Wu, Danny T. M. Chan, Zhen Lei, Jinqiao Wang, Sebastien Ourselin, Hongbin Liu

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08224 2024-03-14 cs.CV 70%

REPAIR: Rank Correlation and Noisy Pair Half-replacing with Memory for Noisy Correspondence

Ruochen Zheng, Jiahao Hong, Changxin Gao, Nong Sang

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00698 2023-10-03 cs.CV cs.HC 70%

Comics for Everyone: Generating Accessible Text Descriptions for Comic Strips

Reshma Ramaprasad

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted at CLVL: 5th Workshop On Closing The Loop Between Vision And Language (ICCV 2023 Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12224 2023-09-22 cs.CL 70%

Towards Answering Health-related Questions from Medical Videos: Datasets and Approaches

Deepak Gupta, Kush Attal, Dina Demner-Fushman

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CL

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.01322 2022-07-12 cs.CV cs.AI cs.CL cs.MM 70%

Recent, rapid advancement in visual question answering architecture: a review

Venkat Kodali, Daniel Berleant

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 8 pages. Accepted to EIT2022 conference and posted on ArXiv in accordance with IEEE policy

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.06977 2022-04-22 cs.CV 70%

Discrete Cosine Transform Network for Guided Depth Map Super-Resolution

Zixiang Zhao, Jiangshe Zhang, Shuang Xu, Zudi Lin, Hanspeter Pfister

专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2022 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.08291 2021-11-17 cs.LG cs.AI eess.SP 70%

Switching Recurrent Kalman Networks

Giao Nguyen-Quynh, Philipp Becker, Chen Qiu, Maja Rudolph, Gerhard Neumann

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

Comments Machine Learning for Autonomous Driving Workshop at the 35th Conference on Neural Information Processing Systems (NeurIPS 2021), Sydney, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.00602 2021-09-03 cs.CL 70%

Point-of-Interest Type Prediction using Text and Images

Danae Sánchez Villegas, Nikolaos Aletras

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted at EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.11486 2021-05-26 eess.IV cs.CV cs.LG 70%

Experimenting with Knowledge Distillation techniques for performing Brain Tumor Segmentation

Ashwin Nalwade, Jackie Kisa

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.00696 2020-10-08 cs.CL 70%

Robust and Interpretable Grounding of Spatial References with Relation Networks

Tsung-Yen Yang, Andrew S. Lan, Karthik Narasimhan

专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CL

Comments Findings of Empirical Methods in Natural Language Processing (EMNLP) 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.10858 2019-12-24 cs.CL cs.LG q-fin.ST stat.ML 70%

"The Squawk Bot": Joint Learning of Time Series and Text Data Modalities for Automated Financial Information Filtering

Xuan-Hong Dang, Syed Yousaf Shah, Petros Zerfos

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23960 2026-08-17 physics.optics 版本更新 67%

Single-Shot Realization of 10000-Mode Octave-Spanning Artificial Gauge Fields

单次实现10000模式倍频程跨度人工规范场

Lida Xu, Apurva Padhye, Supratik Sarkar, Alireza Parhizkar, Christopher J. Flower, Gregory Moille, Kartik Srinivasan, Mohammad Hafezi, Mahmoud Jalali Mehrabad

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

AI总结 提出超宽带多模态色散校正人工规范场理论框架,利用集成光子学实现首个光子整数量子霍尔模型的频率梳,覆盖超过10000个模式,通过克尔非线性实现单次控制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04033 2026-08-06 cs.CR 新提交 67%

SoK: How Frontier AI Reshapes System-Level Security Risk Dynamics in Critical Infrastructure

SoK:前沿人工智能如何重塑关键基础设施中的系统级安全风险动态

Chandra Thapa, Mohan Baruwal Chhetri, Marthie Grobler, Shahroz Tariq, Tooba Aamir

专题命中 其他多模态 :multimodal(abstract);multimodal foundation model(abstract)

AI总结 本研究提出五维风险动态框架,分析前沿人工智能重塑关键基础设施系统级安全风险的机制,指出学术研究与运营约束的错配,推动系统级安全保证研究。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29404 2026-08-03 math.DS 新提交 67%

Decomposition-based Energy-based Dual-Phase Dynamics Identification for Nonlinear MDOF Systems

基于分解的非线性多自由度系统能量双相动力学辨识方法

Sayantan Ghosh, Cristian López, Aryan Singh, Keegan J. Moore

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

AI总结 本研究将能量双相动力学辨识(EDDI)扩展至多自由度系统,提出基于分解的EDDI方法,经两层塔结构实验验证,可有效辨识复杂多模态非线性结构动力学。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16639 2026-08-03 cs.LG 版本更新 67%

SPICE: Synergy and Partial Information Based Curriculum Evolution

SPICE: 基于协同与部分信息的课程演化

Ankush Pratap Singh, Houwei Cao, Yong Liu

机构 * New York Institute of Technology(纽约理工学院) New York University(纽约大学)

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract)

AI总结 提出SPICE框架,利用部分信息分解理论动态量化样本复杂度,设计渐进式课程使模型从学习共享跨模态线索过渡到模态特定模式再到复杂协同交互,在多个多模态基准上取得一致改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06570 2026-07-21 cs.RO 版本更新 67%

Unsupervised Discovery of Failure Taxonomies from Deployment Logs

无监督发现部署日志中的故障分类

Aryaman Gupta, Yusuf Umut Ciftci, Somil Bansal

机构 * Stanford University(斯坦福大学) University of Southern California(南加州大学)

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract)

AI总结 本文提出无监督发现部署日志中的故障分类方法,通过视觉-语言推理和语义聚类,提升机器人系统鲁棒性和故障监控能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06649 2026-07-09 cs.CR cs.LG 新提交 67%

POPS: Recovering Unlearned Multi-Modality Knowledge in MLLMs with Prompt-Optimized Parameter Shaking

POPS:通过提示优化参数抖动恢复多模态大语言模型中未学习的多模态知识

Zhangheng LI, Jianing Zhu, Junyuan Hong, Sungmin Eum, Shuowen Hu, Suya You, Zhangyang Wang

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract)

AI总结 研究针对多模态大语言模型中隐私敏感信息擦除不彻底问题,提出POPS对抗策略,通过提示后缀优化生成潜在私人示例并微调模型,实验揭示现有MMU算法弱点,POPS能近乎完全恢复敏感信息,暴露隐私保护漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21788 2026-06-09 cs.DC cs.LG 版本更新 67%

Efficient Scaling of LLM Training with Flexible Context Parallelism

利用灵活上下文并行实现LLM训练的高效扩展

Yifan Niu, Han Xiao, Dongyi Liu, Wei Zhou, Jia Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 其他多模态 :MLLM(abstract,abstract_cn)

AI总结 针对数据异构导致负载不均和通信冗余问题,提出自适应重配置通信组和上下文并行度的FCP策略,实现近线性加速比,最高达1.46倍吞吐提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05242 2026-05-28 cs.CL cs.AI cs.CV cs.LG 67%

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

超越外部监控:增强大型语言模型的透明度以便于监控

Guanxu Chen, Jing Shao, Tao Luo, Lijie Hu, Qihao Lin, Dongrui Liu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ICISEE, Shanghai Jiao Tong University(上海交通大学ICISEE) School of Mathematical Sciences, Institute of Natural Sciences, MOE-LSC, CMA-Shanghai, Shanghai Jiao Tong University(上海交通大学数学科学学院) King Abdullah University of Science and Technology(卡塔尔国王 Abdullah 科学与技术大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出TELLME方法,通过改进大型语言模型的内部表征透明度,帮助监控者识别不当和敏感行为,并在去毒化任务中验证其有效性。

Comments 28 pages,8 figures,15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10054 2026-05-26 cs.LG cs.AI cs.CL cs.CV 67%

Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs

Uni-DPO:大语言模型动态偏好优化的统一范式

Shangpin Peng, Weinong Wang, Zhuotao Tian, Senqiao Yang, Xing Wu, Haotian Xu, Chengquan Zhang, Takashi Isobe, Baotian Hu, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Xi’an Jiaotong University(西安交通大学) The Chinese University of Hong Kong(香港中文大学) University of Chinese Academy of Sciences(中国科学院大学) Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 针对现有DPO方法忽略数据质量和学习难度差异的问题,提出Uni-DPO统一框架,通过自适应重加权偏好对实现更有效的数据利用和更优性能。

Comments Accepted by ICLR 2026. Code & models: https://github.com/pspdada/Uni-DPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10501 2026-05-12 cs.DC 67%

Accelerating Compound LLM Training Workloads with Maestro

通过Maestro加速复合大语言模型训练工作负载

Xiulong Yuan, Hongqing Chen, Jiaxuan Peng, Fan Zhou, Zhixiang Ruan, Zekun Wang, Bo Zheng, Rui Men, Haiquan Wang, Zhipeng Zhang, Langshi Chen, Man Yuan, Jiaqi Gao, Zhengping Qian, Junyang Lin, Yong Li, Wei Lin, Junhua Wang, Jingren Zhou

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract)

AI总结 本文提出Maestro框架,针对复合LLM训练中的静态异构性和动态不规则性,通过分段图重构和波前调度算法提升硬件利用率,减少40%的GPU消耗。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07353 2026-04-20 cs.CL cs.AI cs.CV cs.CY cs.SI 67%

Toxic Memes: A Survey of Computational Perspectives on the Detection and Explanation of Meme Toxicities

有毒的迷因:计算视角下对迷因毒性的检测与解释的综述

Delfina Sol Martinez Pandiani, Erik Tjong Kim Sang, Davide Ceolin

机构 * organization= Centrum Wiskunde \& Informatica , addressline= Science Park 123 , city= Amsterdam , postcode= 1098 XG , country= The Netherlands organization= Netherlands eScience center , addressline= Science Park 402 , city= Amsterdam , postcode= 1098 XH , country= The Netherlands

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文综述了计算视角下对迷因毒性的检测与解释的研究,识别了30多个数据集及毒性分类方法,提出了新的毒性分类体系,并探讨了跨模态推理、低资源语言处理等挑战。

Comments 39 pages, 12 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09250 2026-04-13 astro-ph.IM physics.ed-ph physics.soc-ph 67%

Unseen Astronomy

未见天文学

Dr James W. Trayford

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

AI总结 本文探讨多模态科学在天文学中的应用,介绍其创新方法和潜在影响,强调其在教育、沟通和研究中的跨领域价值。

Comments Published in Astronomy and Geophysics, Volume 67, Issue 2, April 2026, Pages 2.22-2.26. This is the authors' accepted version of the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01424 2026-01-06 cs.LG q-bio.NC 67%

Unveiling the Heart-Brain Connection: An Analysis of ECG in Cognitive Performance

揭示心脑联系:对认知表现中ECG的分析

Akshay Sasi, Malavika Pradeep, Nusaibah Farrukh, Rahul Venugopal, Elizabeth Sherly

机构 * Digital University Kerala(数字大学凯拉尔) Centre for Consciousness Studies, NIMHANS(意识研究中心,NIMHANS)

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract)

AI总结 本文通过ECG信号分析认知负荷,提出跨模态XGBoost框架实现EEG代表性认知空间投影,验证ECG在日常认知监测中的可行性。

Comments 6 pages, 6 figures. Code available at https://github.com/AkshaySasi/Unveiling-the-Heart-Brain-Connection-An-Analysis-of-ECG-in-Cognitive-Performance. Presented at AIHC (not published)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07666 2026-01-01 cs.LG cs.AI cs.CL cs.CV 67%

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities

在大语言模型、多模态大语言模型及更广泛的领域中进行模型融合:方法、理论、应用与机遇

Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie Zhang, Dacheng Tao

机构 * Shenzhen Campus of Sun Yat-sen University, China(中山大学深圳校区) Northeastern University China(东北大学) Shenzhen Campus of Sun Yat-sen University China(中山大学深圳校区) Nanyang Technological University Singapore(南洋理工大学) Northeastern University(东北大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Nanyang Technological University(南洋理工大学) Institute for Clarity in Documentation Dublin Ohio USA(文档清晰研究所) Inria Paris-Rocquencourt Rocquencourt France(巴黎-罗quentourt研究所) Rajiv Gandhi University Doimukh Arunachal Pradesh India(拉贾·甘地大学) Tsinghua University Haidian Qu Beijing Shi China(清华大学) Palmer Research Laboratories San Antonio Texas USA(帕勒研究中心) Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-罗quentourt研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒研究中心)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文综述了模型融合的方法、理论、应用及未来方向,提出新的分类方法并探讨其在多个机器学习领域的应用及挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06850 2025-09-16 cs.NE 67%

Visual Evolutionary Optimization on Graph-Structured Combinatorial Problems with MLLMs: A Case Study of Influence Maximization

Jie Zhao, Kang Hao Cheong

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05182 2025-08-22 cs.IR 67%

On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools

Shivani Upadhyay, Messiah Ataey, Syed Shariyar Murtaza, Yifan Nie, Jimmy Lin

专题命中 其他多模态 :multi-modal(abstract);MLLM(abstract)

Comments 15 pages, 5 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20448 2025-07-22 cs.LG 67%

Knockout: A simple way to handle missing inputs

Minh Nguyen, Batuhan K. Karaman, Heejong Kim, Alan Q. Wang, Fengbei Liu, Mert R. Sabuncu

机构 * Cornell University(康奈尔大学) Weill Cornell Medicine(韦尔医学院)

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

Comments Accepted at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏