arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6847 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6847 篇

2607.16076 2026-07-20 cs.CV cs.AI cs.CL 新提交 89%

HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection

HCIG:用于多模态讽刺和网络欺凌检测的分层跨模态不协调图网络

Bhavana Verma, Priyanka Meel, Dinesh Kumar Vishwakarma

机构 * Delhi Technological University(德里理工大学) Multimodal Data Analytics Research Laboratory(多模态数据分析研究实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 针对多模态讽刺和网络欺凌检测难题,提出HCIG框架,利用图注意力网络在不同层面建模跨模态不协调并整合,引入GCCN辅助推理。实验表明该方法在相关数据集上表现出色,分层多粒度建模比传统策略更有效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13678 2026-07-16 eess.SP 新提交 89%

M3F-UAV: A Missing-Modality Multimodal Foundation Model for Low-Altitude Wireless Sensing

M3F-UAV:一种用于低空无线传感的缺失模态多模态基础模型

Pengxuan Gao, Kai Ying, Botao Wu, Jianhua Mo, Qingsong Wen

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract)

AI总结 针对复杂环境下单模态模型可靠性低的问题,提出M3F-UAV模型,通过特定模态预训练特征提取器、跨模态融合及缺失模态感知预训练,从多观测中学习统一表示,在LAMBDA数据集实验中性能优于单模态基线且在缺失模态下稳健。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09966 2026-06-10 cs.SD 新提交 89%

RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification

RespiraMFM:一种用于呼吸道疾病识别的对比音频-语言对齐多模态基础模型

Shakhrul Iman Siam, Tiantian Feng, Jiankun Zhang, Shrikanth Narayanan, Mi Zhang

机构 * The Ohio State University(俄亥俄州立大学) University of Southern California(南加州大学) University of Chicago(芝加哥大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract)

AI总结 提出RespiraMFM多模态基础模型,通过对比音频-文本对齐策略整合呼吸音与临床信息,在监督和零样本任务中分别提升AUROC 9.15%和20.98%。

Comments ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07572 2026-03-10 cs.LG 89%

TS-MLLM: A Multi-Modal Large Language Model-based Framework for Industrial Time-Series Big Data Analysis

TS-MLLM:一种基于多模态大语言模型的工业时间序列大数据分析框架

Haiteng Wang, Yikang Li, Yunfei Zhu, Jingheng Yan, Lei Ren, Laurence T. Yang

机构 * School of Automation Science and Electrical Engineering, Beihang University(北京航空航天大学自动化科学与电气工程学院) School of Software, Beihang University(北京航空航天大学软件学院) Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院) State Key Laboratory of Intelligent Manufacturing System Technology(智能制造系统技术国家重点实验室) School of Computer and Artificial Intelligence, Zhengzhou University(郑州大学计算机与人工智能学院) Department of Computer Science, St. Francis Xavier University(圣弗朗西斯科大学计算机科学系)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(title,abstract);cross-modal(abstract)

AI总结 TS-MLLM通过多模态大语言模型联合建模时序信号、频域图像和文本知识,提升工业时间序列预测的鲁棒性、效率和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13715 2026-02-17 cs.IR 89%

DMESR: Dual-view MLLM-based Enhancing Framework for Multimodal Sequential Recommendation

DMESR: 基于双视角的多模态序列推荐增强框架

Mingyao Huang, Qidong Liu, Wenxuan Yang, Moranxin Wang, Yuqi Sun, Haiping Zhu, Feng Tian, Yan Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(title,abstract);cross-modal(abstract)

AI总结 DMESR通过双视角机制解决多模态序列推荐中的语义对齐和细粒度语义丢失问题,提升推荐效果。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03666 2026-01-12 cs.CL cs.AI cs.CV 89%

e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings

e5-omni: 显式跨模态对齐用于多模态嵌入

Haonan Chen, Sicheng Gao, Radu Timofte, Tetsuya Sakai, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) University of Würzburg(乌尔姆大学) Waseda University(早稻田大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);omni-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 e5-omni通过显式对齐方法改进多模态嵌入,解决相似性尺度不一致、负样本效果下降和跨模态统计不匹配问题。

Comments https://huggingface.co/Haon-Chen/e5-omni-7B

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10307 2025-09-16 cs.IR 89%

CROSSAN: Towards Efficient and Effective Adaptation of Multiple Multimodal Foundation Models for Sequential Recommendation

Junchen Fu, Yongxin Ni, Joemon M. Jose, Ioannis Arapakis, Kaiwen Zheng, Youhua Li, Xuri Ge

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21395 2025-07-30 cs.MM cs.AI cs.SD eess.AS 89%

Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion

Zeyu Deng, Yanhui Lu, Jiashu Liao, Shuang Wu, Chongfeng Wei

机构 * James Watt School of Engineering, University of Glasgow(格拉斯哥大学詹姆斯·瓦特工程学院) University of Bristol(布里斯托大学) School of Engineering Mathematics and Technology, University of Bristol(布里斯托大学工程数学与技术学院) School of Computing Science, University of Glasgow(格拉斯哥大学计算科学学院) Department of Civil, Environmental & Geomatic Engineering, University College London (UCL)(伦敦大学学院(UCL)土木、环境与测绘工程系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10838 2024-04-18 cs.CV cs.CL cs.MM 89%

Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning

Zhengyang Liang, Meiyu Liang, Wei Huang, Yawen Li, Zhe Xue

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.02172 2021-11-04 cs.CV cs.CL cs.MM 89%

A cross-modal fusion network based on self-attention and residual structure for multimodal emotion recognition

Ziwang Fu, Feng Liu, Hanyang Wang, Jiayin Qi, Xiangling Fu, Aimin Zhou, Zhibin Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments 5 pages, 1 figure, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01331 2024-06-12 cs.CL cs.AI 89%

LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model

Musashi Hinck, Matthew L. Olson, David Cobbley, Shao-Yen Tseng, Vasudev Lal

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CL、cs.AI

Comments CVPR 2024, MMFM workshop. Authors 1 and 2 contributed equally. Models available at https://huggingface.co/intel/llava-gemma-2b/ and https://huggingface.co/intel/llava-gemma-7b/ Training code at https://github.com/IntelLabs/multimodal_cognitive_ai/tree/main/LLaVA-Gemma

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07535 2026-08-11 cs.LG cs.AI cs.CY 新提交 89%

Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

多模态大语言模型的演进安全态势:新兴威胁与防护措施综述

Xi Li, Shu Zhao, Xiaohan Zou, Fei Zhao, Fuxiao Liu, Yusen Zhang, Cheng Han, Yushun Dong, Jiaqi Wang

机构 * University of Alabama at Birmingham(阿拉巴马大学伯明翰分校) NVIDIA(英伟达公司) Penn State University(宾夕法尼亚州立大学) Columbia University(哥伦比亚大学) University of Missouri-Kansas City(密苏里大学堪萨斯分校) Florida State University(佛罗里达州立大学) Auburn University(奥本大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract)

AI总结 本综述针对多模态大语言模型(MLLMs)的新型安全威胁,提出多模态安全威胁分类法,梳理安全策略进展,探讨未来安全机制的研究方向。

Comments Accepted at the ICLR 2026 Workshop on Principled Design for Trustworthy AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03660 2026-08-05 cs.AI 新提交 89%

Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Multimodal Post-Training

驯服隐式:面向持续多模态后训练的双通道风险感知强化微调

Yibei Liu, Jiajun Chen, Qianle Zhang, Tangyue Jin, Mengying Zhu, Meng Xi, Yangyang Wu

专题命中 多模态训练与对齐 :MLLM(summary_cn,abstract);multimodal(title,abstract);分类 cs.AI

AI总结 针对持续多模态后训练中RFT算法在任务分布偏移下遗忘加剧的问题,提出双通道风险感知强化微调框架RAPO,通过策略与数据通道的风险管控降低遗忘,在MLLM-CL基准上使遗忘减少79.8%且保留新任务竞争力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19120 2026-06-23 cs.LG cs.CV 新提交 89%

Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation

先看后思:解耦感知与推理以实现抗捷径的多模态在策略自蒸馏

Sihan Wang, Xiyao Liu, Lianqing Liu, Zhi Han

机构 * State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences(机器人与智能系统国家重点实验室,沈阳自动化研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :MLLM(summary_cn,abstract);multimodal(title,abstract);分类 cs.CV

AI总结 提出ViGOS框架,通过解耦感知和推理,在MLLM后训练中避免文本捷径,提升图像依赖行为。

Comments 29 pages, 5 figures, 8 tables; Project page: https://oedosoldier.github.io/ViGOS/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15308 2026-06-16 cs.AI 新提交 89%

Forced Deferral: Manipulating Routing Decisions in Multimodal LLM Cascades

强制延迟:在多模态大语言模型级联中操纵路由决策

Zhongye Liu, Yaopei Zeng, Yurui Chang, Lu Lin

机构 * Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 多模态训练与对齐 :MLLM(summary_cn,abstract);multimodal(title,abstract);分类 cs.AI

AI总结 提出强制延迟攻击(FDA),通过对抗性图像攻击降低弱模型置信度,迫使级联系统将查询路由到强模型,揭示了MLLM级联在计算分配上的安全漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06485 2026-06-05 cs.CV 89%

PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding

PAR3D: 一种用于场景理解的统一部件感知3D多模态大语言模型

Shaohui Dai, Yansong Qu, You Shen, Shengchuan Zhang, Liujuan Cao

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(教育部多媒体可信感知与高效计算重点实验室,厦门大学)

专题命中 多模态训练与对齐 :MLLM(title,summary_cn);multimodal(abstract);分类 cs.CV

AI总结 提出PAR3D框架,通过部件感知3D表示学习和层次化分割查询生成,解决现有3D-MLLM在细粒度部件理解上的不足,在部件级问答和指代分割任务上取得显著提升。

Comments Project page: https://atrovast.github.io/PAR3D/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17419 2026-05-08 cs.CV 89%

EAGLE: Expert-Augmented Attention Guidance for Tuning-Free Industrial Anomaly Detection in Multimodal Large Language Models

EAGLE:专家增强的注意力引导用于多模态大语言模型中的无调优工业异常检测

Xiaomeng Peng, Xilang Huang, Seon Han Choi

机构 * Ewha Womans University(峨山女子大学)

专题命中 多模态训练与对齐 :MLLM(summary_cn,abstract);multimodal(title,abstract);分类 cs.CV

AI总结 EAGLE通过整合专家异常检测器与冻结的MLLM,提出无调优框架,提升多模态大语言模型在工业异常检测中的准确率,且在MVTec-AD和VisA数据集上达到94.4%和88.1%的异常鉴别准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08592 2026-04-28 cs.CV 89%

Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations

通过几何参考的3D场景表示提升MLLM空间推理能力

Jiangye Yuan, Gowri Kumar, Baoyuan Wang

机构 * Zillow Group(智域集团)

专题命中 多模态训练与对齐 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CV

AI总结 本文提出几何参考3D场景表示GR3D,使MLLM能利用语言能力处理3D空间,无需额外训练,在空间推理基准中提升GPT-5性能9%和12%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20135 2026-04-23 cs.CL cs.IR 89%

AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce

AFMRL: 基于属性增强的电商细粒度多模态表示学习

Biao Zhang, Lixin Chen, Bin Zhang, Zongwei Wang, Tong Liu, Bo Zheng

机构 * Taobao & Tmall Group of Alibaba(阿里巴巴淘宝与天猫集团)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract,abstract_cn);multimodal(abstract);image-text(abstract)

AI总结 本文提出AFMRL,通过属性生成任务提升电商细粒度多模态表示学习,结合属性引导对比学习和检索感知属性强化,实现更精准的相似商品检索。

Comments Accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00214 2026-07-29 cs.CV cs.AI 88%

A Geometric Multimodal Foundation Model Integrating Bp-MRI and Clinical Reports in Prostate Cancer Classification

一种整合Bp-MRI和临床报告的几何多模态基础模型用于前列腺癌分类

Juan A. Olmos, Antoine Manzanera, Fabio Martínez

机构 * Biomedical Imaging, Vision and Learning Laboratory (BIVL$^2$ab), UIS, Colombia(生物医学成像、视觉与学习实验室(BIVL²ab), UIS,哥伦比亚) U2IS, ENSTA, Institut Polytechnique de Paris, France(U2IS, ENSTA,巴黎理工学院,法国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出MFM-Geom模型,通过整合Bp-MRI和临床报告,利用几何方法提升前列腺癌分类的准确性和鲁棒性。

Comments Accepted at IEEE International Symposium on Biomedical Imaging (ISBI) 2026

Journal ref 2026 IEEE 23rd International Symposium on Biomedical Imaging (ISBI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25034 2026-06-29 cs.CV cs.AI 新提交 88%

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety

Yuvion VL:面向对抗性内容和AI安全的多模态基础模型

Shikai Qiu, Xiaowen Xu, Benlei Cui, Ting Ma, Xiufeng Huang, Wenjing Jiang, Shaoxuan He, Haolei Xu, Chunyang Chai, Yujian Li, Yiliang Zhang, Guanghui Wang, Ziheng Wang, Ziwen Xu, Zhaoyu Fan, Jinhao Chen, Ruijie Jian, Hongxing Li, Chuxi Xiao, Xinyue Chen, Wenxuan Liu, Libin Dong, Yupeng Cao, Xiaoqian Xia, Jing Wang, Zhe Jiang, Zhenan Ye, Guang Yang, Bin Liu, Wei Peng, Ziqiang Zhu, Meihui Lian, Kaiwen Lv Kacuila, Haidong Ding, Dongjie Zhang, Yangfan Zhou, Bingyu Zhu, Yan Wang, Hai Zhao, Xuan Jin, Wei Zhao, Pengfei Sun, Huiming Zhang, Wei Wang, Xipeng Cao, Jialun Chen, Xiao Chen, Shaola Ren, Yunqing Hu, Bin Li, Chengwen Yao, Meng Huang, Xianfeng Li, Bin Tang, Chao Liu, Hui Xue, Longtao Huang, Haiwen Hong

机构 * Alibaba Security AGI Lab(阿里巴巴安全AGI实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 提出Yuvion VL系列多模态大语言模型,通过对抗性感知数据合成、三阶段训练和混淆对比微调,在内容和AI安全任务上达到行业领先性能,同时保持通用能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20268 2026-05-21 cs.LG cs.AI cs.CL 88%

Chronicle: A Multimodal Foundation Model for Joint Language and Time Series Understanding

Chronicle:一种用于联合语言和时间序列理解的多模态基础模型

Paul Quinlan, Jeremy Levasseur, Qingguo Li, Xiaodan Zhu

机构 * InertialAI Department of Electrical and Computer Engineering, Queen’s University(皇后大学电气与计算机工程系) Department of Mechanical and Materials Engineering, Queen’s University(皇后大学机械与材料工程系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title);cross-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出Chronicle,一种联合训练语言和时间序列的多模态基础模型,通过统一架构实现两者共享参数,从而在多个任务上取得了优异表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16806 2026-05-19 cs.LG cs.AI cs.CV 88%

Cross-modal Affinity-aligned Multimodal Learning Analytics for Predicting Student Collaboration Satisfaction in Game-Based Learning

跨模态亲和对齐的多模态学习分析用于预测基于游戏的学习中学生协作满意度

Wen-Hsin Tsai, Chia-Ming Lee, Yuk-Ying Tung

机构 * Institute of Education, National Cheng Kung University(国立成功大学教育研究所) Institute of Intelligent System, National Yang Ming Chiao Tung University(阳明交通大学智能系统研究所) Department of Computer Science, University at Albany, State University of New York(纽约州立大学水牛城分校计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种跨模态亲和对齐的多模态学习分析框架,通过建模模态间关系和对比学习来增强学生协作满意度预测的鲁棒性和可解释性。

Comments Accetped by CVPR 2026 CVxEdu Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10765 2026-05-12 cs.CV cs.AI cs.LG 88%

Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning

动态跨模态提示生成用于多模态持续指令微调

Tao Hu, Da-Wei Zhou

机构 * School of Artificial Intelligence, Nanjing University(南京大学人工智能学院) State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出DRAPE框架,通过生成连续实例特定的软提示提升多模态持续指令微调性能,采用跨模态注意力和投影梯度投影减少遗忘,实验显示优于现有基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24827 2025-10-30 cs.CV cs.MM 88%

MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition

Haoyang Zhang, Zhou Yang, Ke Sun, Yucai Pang, Guoliang Xu

机构 * Chongqing University of Posts and Telecommunications(重庆邮电大学) Xi’an Jiaotong University(西安交通大学) University of New South Wales(新南威尔士大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments The paper will be published in the MMAsia2025 conference proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17823 2025-10-28 cs.CV cs.AI cs.LG 88%

Robust Multimodal Learning via Cross-Modal Proxy Tokens

Md Kaykobad Reza, Ameya Patil, Mashhour Solh, M. Salman Asif

机构 * University of California Riverside(加州大学河滨分校) Amazon(亚马逊)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 28 Pages, 13 Figures, 11 Tables. Accepted by Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17205 2025-10-21 cs.CV cs.CL 88%

$\mathcal{V}isi\mathcal{P}runer$: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs

Yingqi Fan, Anhao Zhao, Jinlan Fu, Junlong Tong, Hui Su, Yijie Pan, Wei Zhang, Xiaoyu Shen

机构 * Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT, Ningbo(宁波空间智能与数字衍生关键实验室,数字孪生研究院,EIT,宁波) Shanghai Jiao Tong University(上海交通大学) Hong Kong Polytechnic University(香港理工大学) Meituan Inc.(美团公司) National University of Singapore(新加坡国立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08606 2025-10-13 cs.CL cs.AI 88%

Centering Emotion Hotspots: Multimodal Local-Global Fusion and Cross-Modal Alignment for Emotion Recognition in Conversations

Yu Liu, Hanlei Shi, Haoxun Li, Yuqing Sun, Yuxuan Ding, Linlin Gong, Leyuan Qu, Taihao Li

机构 * Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(杭州高等研究 institute,中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CL、cs.AI

Comments Under review for ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04653 2025-09-23 cs.CV cs.CL 88%

LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts

Yimu Wang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof Czarnecki

机构 * University of Waterloo(滑铁卢大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(title);MLLM(abstract);分类 cs.CV、cs.CL

Comments To appear at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16300 2025-08-25 cs.CV cs.AI 88%

A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension

Mohammad Zia Ur Rehman, Devraj Raghuvanshi, Umang Jain, Shubhi Bansal, Nagendra Kumar

机构 * Indian Institute of Technology Indore(印度印度理工学院Indore) Brown University(布朗大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments Published in Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏