arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-06-10 至 2026-06-10 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 13 篇

2606.09966 2026-06-10 cs.SD 新提交 89%

RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification

RespiraMFM:一种用于呼吸道疾病识别的对比音频-语言对齐多模态基础模型

Shakhrul Iman Siam, Tiantian Feng, Jiankun Zhang, Shrikanth Narayanan, Mi Zhang

机构 * The Ohio State University(俄亥俄州立大学) University of Southern California(南加州大学) University of Chicago(芝加哥大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract)

AI总结 提出RespiraMFM多模态基础模型,通过对比音频-文本对齐策略整合呼吸音与临床信息,在监督和零样本任务中分别提升AUROC 9.15%和20.98%。

Comments ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09859 2026-06-10 cs.LG cs.AI 新提交 87%

Mitigating Manifold Departure: Uncertainty-Aware Subspace Rectification for Trustworthy MLLM Decoding

缓解流形偏离:面向可信MLLM解码的不确定性感知子空间校正

Yingxuan Zhuang, Jingxiao Yang, Miao Pan, Cheng Tan, Yuxiang Cai, Siwei Tan, Chen Zhi, Xuhong Zhang, Jianwei Yin, Jintao Chen

机构 * Nanyang Technological University(南洋理工大学)

专题命中 多模态训练与对齐 :MLLM(title,title_cn);multimodal(abstract);分类 cs.AI

AI总结 提出MGAP方法,通过SVD构建语言先验子空间并自适应衰减投影分量,在抑制幻觉的同时保持语义结构,优于现有解码基线。

Comments ICML 2026 regular

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10412 2026-06-10 cs.AI 新提交 86%

A Unified Multi-Modal Framework for Intelligent Financial Systems: Integrating Reinforcement Learning, High-Frequency Trading, and Game-Theoretic Approaches with Cross-Modal Sentiment Analysis

面向智能金融系统的统一多模态框架:整合强化学习、高频交易和博弈论方法与跨模态情感分析

Fanrong Liu, Zhang Yuwei, Mingni Luo

机构 * Henan University, International Eurasia College(河南大学,国际欧亚学院) City University of Hong Kong, College of Business(香港城市大学,商学院) Northeastern University, School of Electronic and Information Engineering(东北大学,电子与信息工程学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(title);分类 cs.AI

AI总结 提出统一框架整合PPO、高频预测、上下文学习、博弈论和跨模态情感分析,在多个金融任务上平均提升20%以上性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10504 2026-06-10 cs.AI 新提交 85%

Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and Algorithm

无配对数据的跨模态知识蒸馏:理论基础与算法

Trong Khiem Tran, Anh Duc Chu, Quang Hung Pham, Phi Le Nguyen, Trong Nghia Hoang

机构 * School of Information and Communications Technology, Hanoi University of Science and Technology, Hanoi, Vietnam(信息与通信技术学院,河内科学技术大学,越南河内) School of Electrical Engineering and Computer Science, Washington State University, Pullman, US(电气工程与计算机科学学院,华盛顿州立大学,华盛顿州普尔曼)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.AI

AI总结 提出无配对数据下的跨模态知识蒸馏框架,通过特征对齐和标签对齐两种分布对齐机制,实现跨模态知识迁移,理论保证且实验效果显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11188 2026-06-10 cs.CV 新提交 79%

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations

ARM: 一种具有统一离散表示的自回归大型多模态模型

Junke Wang, Xiao Wang, Jiacheng Pan, Xuefeng Hu, Feng Li, Jingxiang Sun, Chaorui Deng, Zilong Chen, Yunpeng Chen, Kaibin Tian, Matthew Gwilliam, Hao Chen, Danhui Guan, Kun Xu, Weilin Huang, Zuxuan Wu, Haoqi Fan, Yu-Gang Jiang, Zhenheng Yang

机构 * Shanghai Key Lab of Intelligent Information Processing, Fudan University(复旦大学上海智能信息处理重点实验室) School of Computer Science, Fudan University(复旦大学计算机科学技术学院) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) Youtu Lab, Tencent(腾讯优图实验室) Meta AI Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 提出ARM模型,通过离散语义视觉分词器将图像映射为紧凑token序列,结合自回归建模和强化学习,统一实现图像理解、生成和编辑,并提升任务性能与跨任务协同。

Comments technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10572 2026-06-10 cs.AI 新提交 79%

One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA

每个多模态证据一个令牌:面向资源受限问答的潜在记忆

Zhi Zheng, Ziqiao Meng, Hao Luan, Wei Liu, Wee Sun Lee

机构 * School of Computing, National University of Singapore(新加坡国立大学计算机学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 提出潜在记忆范式,将每个证据压缩为单个高维潜在令牌,通过统一训练实现高效检索与生成,在资源受限场景下以3-10倍令牌节省达到竞争性问答性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10488 2026-06-10 cs.CV 新提交 79%

5% > 100%: Flatness Preference is All You Need for Multimodal Parameter-Efficient Fine-Tuning

5% > 100%: 平坦性偏好是您进行多模态参数高效微调所需的一切

Yifan Zhu, Can Lin, Hangjie Yuan, Zixiang Zhao, Pengfei Zhang, Tao Feng, Zhonghong Ou

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Zhejiang University(浙江大学) ETH Zürich(苏黎世联邦理工学院) Anhui University of Science and Technology(安徽理工大学) Tsinghua University(清华大学) State Key Laboratory of Networking and Switching Technology(网络与交换技术国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 揭示参数高效微调方法中普遍存在的平坦性偏好,即少量尖锐维度主导泛化,并提出FlatPO方法优化这些维度以提升泛化性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03063 2026-06-10 cs.SI 78%

Unsupervised Multimodal Graph-based Model for Geo-social Analysis

无监督多模态图模型用于地理社交分析

Ehsaneddin Jalilian, Bernd Resch

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出无监督多模态图模型,整合语义与地理信息,提升灾难管理和舆论监控中的内容分析效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10166 2026-06-10 cs.CV 新提交 57%

Fusing Satellite Imagery and Planimetric Maps for Cross-View Localization

融合卫星图像与平面地图的跨视角定位

Quang Long Ho Ngo, Zimin Xia, Alexandre Alahi

机构 * École Polytechnique Fédérale de Lausanne(瑞士联邦理工学院洛桑校区)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 提出一种融合卫星图像与平面地图的模块,通过跨模态条件化和补丁级融合规则,将定位误差降低30.13%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00809 2026-06-10 cs.CV 版本更新 57%

Let ViT Speak: Generative Language-Image Pre-training

让ViT说话:生成式语言-图像预训练

Yan Fang, Mengcheng Lan, Zilong Huang, Weixian Lei, Yunqing Zhao, Yujie Zhong, Yingchen Yu, Qi She, Yao Zhao, Yunchao Wei

机构 * Beijing Jiaotong University(北京交通大学) ByteDance(字节跳动) Nanyang Technological University(南洋理工大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 提出GenLIP框架,通过语言建模目标直接训练ViT从视觉token预测语言token,无需对比学习或额外文本解码器,实现简单、可扩展且性能优异的视觉编码器。

Comments 27 pages, 11 figures. Code and models are available at https://github.com/YanFangCS/GenLIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10365 2026-06-10 cs.SD 新提交 50%

KFC-KWS: Keyframe Fusion with CTC for User-Defined Keyword Spotting

KFC-KWS: 基于CTC的关键帧融合用于用户自定义关键词唤醒

Jin Li, Wenbin Jiang, Ji Hu

机构 * School of Electronics and Information Engineering, Hangzhou Dianzi University(杭州电子科技大学电子信息学院) School of Communication Engineering, Hangzhou Dianzi University(杭州电子科技大学通信工程学院)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 提出KFC-KWS多模态框架,利用CTC引导的关键帧选择对齐音频、音素和文本模态,通过交叉注意力融合关键帧与全句表示,在LibriPhrase上达到98.73% AUC,困难子集上97.65% AUC和7.75% EER,有效区分易混淆关键词。

Comments Accepted by Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10227 2026-06-10 cs.LG 新提交 50%

Spatiotemporal Graph Transformer for 3D Neighborhood Interaction and Quality Prediction in Metal Additive Manufacturing

时空图Transformer用于金属增材制造中的3D邻域交互与质量预测

Joyce Karen Pelaez, Siqi Zhang, Hoo Sang Ko

机构 * Department of Industrial Engineering(工业工程系)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 提出一种时空图Transformer,通过加权网络表示和双注意力机制建模3D邻域交互,显著提升金属增材制造质量预测性能。

Comments Submitted to Journal of Intelligent Manufacturing, 23 pages, 10 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22017 2026-06-10 cs.LG 版本更新 50%

Domain Adapted Large Language Models for Additive Manufacturing

面向增材制造的领域自适应大语言模型

Peter Pak, Amir Barati Farimani

机构 * Department of Mechanical Engineering, Carnegie Mellon University(机械工程系,卡内基梅隆大学)

专题命中 多模态训练与对齐 :multi-modal(abstract)

AI总结 本文通过约5000万token的小型数据集对开源大语言模型进行领域自适应预训练和指令微调,构建多模态领域自适应模型,在增材制造基准测试中达到90%以上准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏