Hybrid Multimodal Fusion for Humor Detection
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM
Comments 7 pages, 1 figure, to appear in MuSe 2022 (ACM MM2022 co-located workshop)
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM
Comments 7 pages, 1 figure, to appear in MuSe 2022 (ACM MM2022 co-located workshop)
专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL、cs.AI、cs.MM
Comments Accepted by COLING 2022
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、eess.AS
Comments 8 pages, 2 figures, to appear in MuSe 2022 (ACM MM2022 co-located workshop)
专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted by AAAI2022
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、eess.AS
Comments In INTERSPEECH 2021
Journal ref Proc. Interspeech 2021, 2381-2385
专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to ACM MM 2021
专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract)
Comments Under review
专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract)
Comments Accepted
Journal ref IEEE Transactions on Multimedia, 2021
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、eess.AS
Comments Camera-ready version for EACL 2021
专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Submitted to the 17th International Conference on Principles of Knowledge Representation and Reasoning (2020)
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)
Comments Code available at https://github.com/idearibosome/embracenet
Journal ref Information Fusion 51 (2019) 259-270
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments CVPR2019 accepted paper
专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract)
Comments To appear at NIPS 2018; 9 pages with supplement
专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract)
Comments Accepted in "2018 International Conference on Pattern Recognition"
专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Published at AAAI-18, 7 pages
专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
VAD:为多模态在线策略蒸馏中的目标重建归因视觉证据
机构 * Shanghai Jiao Tong University(上海交通大学) ; Xiaohongshu Inc.(小红书公司) ; The Chinese University of Hong Kong(香港中文大学) ; Zhejiang University(浙江大学) ; Southeast University(东南大学)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL
AI总结 该研究提出视觉归因蒸馏(VAD)算法,通过反事实目标重建分离教师校正中的视觉证据分量,在6个4B/9B规模的细粒度视觉基准上,性能优于现有蒸馏方法。
Comments The project is accessible at https://github.com/DeepExperience/VAD_Multimodal_OPD
停止思考,开始观察:通过无推理对齐实现多模态文档问答的高效训练后优化
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI
AI总结 研究多模态文档问答中高效训练后优化问题,提出感知 - RFT 框架,用组相对策略优化绕过推理令牌直接对齐视觉与基础输出,通过构建变体评估推理必要性,发现启用推理模型有优势,还识别基础差异,表明早期转换可减少训练数据并保持精度。
Comments Accepted at ICML 2026, Workshop on Efficient Multimodal Question Answering (EMM-QA)
理解与生成相冲突吗?统一多模态模型DPO的诊断研究
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI
AI总结 通过系统实验发现,在统一多模态模型上应用DPO时,生成质量难以对齐,主要原因是理解和生成梯度近乎正交且存在11-14倍的幅度不平衡,源于VQ token数量不对称。
Comments Experiments are inconclusive: The claim that architectures such as Chameleon or Emu would exhibit stronger gradient conflict is not supported by experiments or analysis, and all experiments are conducted on Janus-Pro without evaluation on other unified multimodal architectures
机构 * Department of Engineering Science, University of Oxford, Oxford, UK(工程科学系,牛津大学,牛津,英国) ; Nuffield Department of Clinical Neurosciences, University of Oxford, Oxford, UK(牛津大学临床神经科学系,牛津,英国)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments Accepted to MICCAI 2025, for the following workshop: ML-CDS 2025: Multimodal Learning and Fusion Across Scales for Clinical Decision Support
机构 * Southeast University(东南大学) ; Xi'an Jiaotong University(西安交通大学)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments Accepted by ACM MM 2025. First PTQ solution for Multimodal large language models applicable to 5 mainstream MLLMs
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments Keywords: Multimodal Fusion, Breast Cancer, Whole Slide Images, Deep Neural Network, Survival Prediction
Journal ref JBHI, 24 June 2024
专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、eess.AS;multimodal(comments)
Comments Accepted for publication in the 5th Multimodal Learning and Applications (MULA) Workshop at CVPR 2022
医学VQA中通过置信度-证据贝叶斯增益实现确定性幻觉检测
机构 * Department of Electrical Engineering, Stanford University, CA, USA(电气工程系,斯坦福大学) ; Department of Biology, Stanford University, CA, USA(生物学系,斯坦福大学) ; Division of Cardiology, Department of Medicine, Stanford University, CA, USA(心脏病学部,医学系,斯坦福大学) ; Department of Biomedical Data Science, Stanford University, CA, USA(生物医学数据科学系,斯坦福大学) ; Department of Computer Science, Stanford University, CA, USA(计算机科学系,斯坦福大学) ; Department of Psychiatry and Behavioral Sciences, Stanford University, CA, USA(精神病学与行为科学系,斯坦福大学)
专题命中 多模态训练与对齐 :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.AI
AI总结 本文提出CEBaG方法,利用模型自身log概率中的不一致置信度和弱视觉证据敏感性,实现无需随机采样和外部模型的确定性幻觉检测,在医疗MLLM和VQA基准测试中取得最佳AUC表现。
多语言和多模态大语言模型在野:为低资源语言构建
机构 * Qatar Computing Research Institute(卡塔尔计算研究所) ; HBKU(哈马德大学) ; York University(约克大学)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL;multimodal foundation model(comments)
AI总结 本文探讨了在有限数据和计算资源下构建多语言多模态大语言模型的方法,涵盖了低成本数据创建、三模态对齐适配器堆栈以及文化感知评估等核心技术和资源。
Comments Multimodal Foundation Models, Large Language Models, Native, Multilingual, Language Diversity, Low-resources-language
走出洞穴:JAM用于对齐独立训练的视觉和语言模型
机构 * Computation and Neural Systems(计算与神经系统) ; California Institute of Technology(加利福尼亚理工学院) ; Computation and Mathematical Sciences(计算与数学科学) ; Google DeepMind(谷歌DeepMind)
专题命中 多模态训练与对齐 :multimodal(abstract,abstract_cn);cross-modal(abstract,abstract_cn);分类 cs.CV
AI总结 本文提出JAM方法,通过联合训练模态特定的自编码器,优化视觉和语言模型的对齐,提升细粒度上下文区分能力。