arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2307.06235 2023-10-10 cs.LG q-bio.BM 71%

Multimodal Molecular Pretraining via Modality Blending

Qiying Yu, Yudi Zhang, Yuyan Ni, Shikun Feng, Yanyan Lan, Hao Zhou, Jingjing Liu

专题命中 多模态训练与对齐 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.00003 2023-04-04 eess.IV 71%

Multimodal Information Fusion For The Diagnosis Of Diabetic Retinopathy

Yihao Li, Hassan Al Hajj, Pierre-Henri Conze, Mostafa EI Habib Daho, Sophie Bonnin, Hugang Ren, Niranchana Manivannan, Stephanie Magazzeni, Ramin Tadayoni, Mathieu Lamard, Gwenole Quellec

专题命中 多模态训练与对齐 :multimodal(title)

Comments Abstract

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.01775 2023-02-09 cs.LG eess.SP 71%

Multimodal Representation Learning using Deep Multiset Canonical Correlation

Krishna Somandepalli, Naveen Kumar, Ruchir Travadi, Shrikanth Narayanan

专题命中 多模态训练与对齐 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.04185 2022-12-05 cs.RO 71%

Coupled Modeling and Fusion Control for a Multi-modal Deformable Land-air Robot

Xinyu Zhang, Yuanhao Huang, Kangyao Huang, Ziqi Zhao, Jingwei Li, Huaping Liu, Jun Li

专题命中 多模态训练与对齐 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10365 2022-10-20 cs.RO 71%

A sensor-to-pattern calibration framework for multi-modal industrial collaborative cells

Daniela Rato, Miguel Oliveira, Vítor Santos, Manuel Gomes, Angel Sappa

专题命中 多模态训练与对齐 :multi-modal(title)

Comments Journal of Manufacturing Systems

Journal ref Journal of Manufacturing Systems 64 (2022) 497-507

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.08982 2022-05-19 cs.IR 71%

Attention-based Multimodal Feature Representation Model for Micro-video Recommendation

Mohan Hasama, Jing Li

专题命中 多模态训练与对齐 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.02259 2021-10-07 cs.CR 71%

Multi-Modal Attack Detection for Cyber-Physical Additive Manufacturing

Shih-Yuan Yu, Arnav Vaibhav Malawade, Mohammad Abdullah Al Faruque

专题命中 多模态训练与对齐 :multi-modal(title)

Journal ref TC-CPS Newsletter Volume 05, Issue 01 (Mar. 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06452 2021-08-17 cs.LG cs.SI 71%

AdaGNN: A multi-modal latent representation meta-learner for GNNs based on AdaBoosting

Qinyi Zhu, Yiou Xiao

专题命中 多模态训练与对齐 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06437 2021-08-17 cs.HC 71%

VR Sickness Prediction from Integrated HMD's Sensors using Multimodal Deep Fusion Network

Rifatul Islam, Kevin Desai, John Quarles

专题命中 多模态训练与对齐 :multimodal(title)

Comments Manuscript to be published at the 2021 IEEE International Symposium on Mixed and Augmented Reality (ISMAR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.05124 2021-07-13 cs.IR cs.LG cs.NE 71%

Transformers with multi-modal features and post-fusion context for e-commerce session-based recommendation

Gabriel de Souza P. Moreira, Sara Rabhi, Ronay Ak, Md Yasin Kabir, Even Oldridge

专题命中 多模态训练与对齐 :multi-modal(title)

Comments In Proceedings of SIGIR eCom'21 - SIGIR eCommerce Workshop Data Challenge 2021. https://sigir-ecom.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.02209 2021-05-05 gr-qc astro-ph.HE 71%

New effective precession spin for modeling multimodal gravitational waveforms in the strong-field regime

Lucy M. Thomas, Patricia Schmidt, Geraint Pratten

专题命中 多模态训练与对齐 :multimodal(title)

Comments Matches version published in Phys. Rev. D. 23 pages (incl. appendices and bibliography), 17 figures

Journal ref Phys. Rev. D 103, 083022 (2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.12241 2021-05-04 cs.RO 71%

Multimodal Data Fusion for Power-On-and-Go Robotic Systems in Retail

Shubham Sonawani, Kailas Maneparambil, Heni Ben Amor

专题命中 多模态训练与对齐 :multimodal(title)

Comments POGO Workshop, RSS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.12656 2021-03-24 cs.LG 71%

Bidirectional Representation Learning from Transformers using Multimodal Electronic Health Record Data to Predict Depression

Yiwen Meng, William Speier, Michael K. Ong, Corey W. Arnold

专题命中 多模态训练与对齐 :multimodal(title)

Comments in IEEE Journal of Biomedical and Health Informatics (2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.04078 2020-12-09 cs.HC cs.RO 71%

Supporting User Autonomy with Multimodal Fusion to Detect when a User Needs Assistance from a Social Robot

Alex Reneau, Jason R. Wilson

专题命中 多模态训练与对齐 :multimodal(title)

Comments 9 pages, 5 figures, 3 tables, AI-HRI FSS

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.14966 2020-12-01 cs.LG 71%

Depression Status Estimation by Deep Learning based Hybrid Multi-Modal Fusion Model

Hrithwik Shalu, Harikrishnan P, Hari Sankar CN, Akash Das, Saptarshi Majumder, Arnhav Datar, Subin Mathew MS, Anugyan Das, Juned Kadiwala

专题命中 多模态训练与对齐 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.10726 2020-08-26 cs.LG eess.SP 71%

Unsupervised Multi-Modal Representation Learning for Affective Computing with Multi-Corpus Wearable Data

Kyle Ross, Paul Hungler, Ali Etemad

专题命中 多模态训练与对齐 :multi-modal(title)

Comments 16 pages,5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.06879 2020-05-20 cs.LG eess.SP stat.ML 71%

Sensor Fusion using Backward Shortcut Connections for Sleep Apnea Detection in Multi-Modal Data

Tom Van Steenkiste, Dirk Deschrijver, Tom Dhaene

专题命中 多模态训练与对齐 :multi-modal(title)

Comments Paper presented at ML4H (Machine Learning for Health) workshop at NeurIPS 2019. https://ml4health.github.io/2019/

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.12081 2020-04-28 cs.HC eess.SP 71%

A novel multimodal approach for hybrid brain-computer interface

Zhe Sun, Zihao Huang, Feng Duan, Yu Liu

专题命中 多模态训练与对齐 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.14564 2019-11-25 math.PR math.OC 71%

Statistical Estimation of the Poincar{é} constant and Application to Sampling Multimodal Distributions

Loucas Pillaud-Vivien, Francis Bach, Tony Lelièvre, Alessandro Rudi, Gabriel Stoltz

专题命中 多模态训练与对齐 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.03306 2019-08-12 q-bio.BM 71%

Cross-Modal Fusion Between Data in SAXS and Cryo-EM for Biomolecular Structure Determination

Shengnan Lyu, Christian Wülker, Yuqing Pan, Amitesh S. Jayaraman, Jianhao Zheng, Yilin Cai, Gregory S. Chirikjian

专题命中 多模态训练与对齐 :cross-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.03684 2019-03-11 cs.CE 71%

Visual Attention Model for Cross-sectional Stock Return Prediction and End-to-End Multimodal Market Representation Learning

Ran Zhao, Yuntian Deng, Mark Dredze, Arun Verma, David Rosenberg, Amanda Stent

专题命中 多模态训练与对齐 :multimodal(title)

Comments Accepted as full paper in the 32nd International FLAIRS Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.12954 2018-11-01 q-bio.NC 71%

Alternating Diffusion Map Based Fusion of Multimodal Brain Connectivity Networks for IQ Prediction

Li Xiao, Julia M. Stephen, Tony W. Wilson, Vince D. Calhoun, Yu-Ping Wang

专题命中 多模态训练与对齐 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
1505.07717 2016-08-24 cs.IT math.IT 71%

Exploring multimodal data fusion through joint decompositions with flexible couplings

Rodrigo Cabral Farias, Jeremy Emile Cohen, Pierre Comon

专题命中 多模态训练与对齐 :multimodal(title)

Comments 15 pages, 7 figures, revised version

详情

展开后加载摘要…

URL PDF HTML 收藏
1409.5050 2015-06-22 physics.optics cond-mat.mtrl-sci 71%

Multimodal Plasmonics in Fused Colloidal Networks

Alexandre Teulle, M. Bosman, C. Girard, Kargal L. Gurunatha, Mei Li, Stephen Mann, Erik Dujardin

专题命中 多模态训练与对齐 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19598 2026-08-21 cs.CV cs.AI cs.CL cs.MM 新提交 70%

PEA-DPO: Perception-Enhanced Alignment Direct Preference Optimization for MLLMs Alignment

PEA-DPO:用于多模态大语言模型对齐的感知增强型直接偏好优化

Jiawei Feng, Jiancan Wu, Xingyu Zhu, Junkang Wu, Xiang Wang, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 针对多模态大语言模型偏好优化存在的视觉不敏感性问题,提出PEA-DPO框架,利用视觉偏好信号缓解该问题,提升多模态对齐效果并减少幻觉。

Journal ref Proceedings of the 34th ACM International Conference on Multimedia (MM '26), November 10--14, 2026, Rio de Janeiro, Brazil

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17414 2026-08-19 cs.CV cs.PL 新提交 70%

REChart: Reasoning-Efficient Chart Editing with Large Reasoning Models

REChart:基于大型推理模型的推理高效图表编辑方法

Yuanbang Liu, Chenxi Ruan, Yihan Hou, Qiong Luo, Wei Zeng

机构 * HKUST(GZ)(香港科技大学(广州))

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract_cn);分类 cs.CV

AI总结 REChart是两阶段训练框架,通过过程级监督提升图表编辑的保真度与推理效率,在两个基准上实现同规模开源模型最优性能,推理token用量降低79.0%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06458 2026-08-19 cs.IR cs.CV cs.LG 70%

PixRec: Leveraging Visual Context for Next-Item Prediction in Sequential Recommendation

PixRec: 利用视觉上下文进行序列推荐中的下一项预测

Sayak Chakrabarty, Souradip Pal

机构 * Department of Computer Science, Northwestern University(北western大学计算机科学系) Elmore Family School of Electrical and Computer Engineering(电气与计算机工程学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 PixRec通过整合视觉信息提升序列推荐的准确性,利用视觉-语言模型在电子商务中实现更精准的下一项预测。

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14829 2026-08-18 eess.IV cs.CV 新提交 70%

Modality-Invariant Coarse-to-Fine Retinal Image Registration

模态无关的由粗到细视网膜图像配准

Bo Wen, Nehal Nailesh Mehta, Melanie Tran, Dirk-Uwe Bartsch, William Freeman, Truong Nguyen

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 针对现有视网膜配准方法依赖模态的局限,提出可泛化的两阶段模态无关框架,含稀疏特征匹配模型与MI-RAFT网络,在多模态视网膜图像配准中性能优于现有方法。

Comments This paper is a submission to IEEE Transactions on Image Processing (TIP-40498-2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14662 2026-08-18 eess.SP cs.AI cs.LG 新提交 70%

Does the Heart Show Your Pain? Tackling the X-ITE Pain Challenge with Self-Supervised ECG Representation Learning

心脏能反映你的疼痛吗?用自监督心电表示学习应对X-ITE疼痛挑战赛

Dominika Kunc, Przemysław Kazienko, Stanisław Saganowski

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本研究针对X-ITE疼痛挑战赛,结合自监督心电表示学习与多模态预训练,分析疼痛识别中的心电信号特性,为可穿戴疼痛监测提供了基础。

Comments 5 pages, 3 Figures, 1 Table, appear in the Proceedings of the 13th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14198 2026-08-17 cs.LG cs.CL 新提交 70%

MINT: A Universal Zero-Shot Predictor for Transaction Data

MINT:一种用于交易数据的通用零样本预测器

Parameswaran Kamalaruban, Viktor Drobnyi, Maeve Madigan, Julia Rozanova, David Sutton, Stuart Burrell

机构 * Visa Inc.(维萨公司)

专题命中 多模态训练与对齐 :multimodal(abstract,abstract_cn);分类 cs.CL

AI总结 该研究提出通用零样本预测框架MINT,通过轻量级嵌入注入等技术连接交易序列编码器与LLM,在交易预测问答任务中性能领先且资源消耗更低,证实紧凑交易嵌入更具优势。

详情

展开后加载摘要…

URL PDF HTML 收藏