Improving Cross-modal Alignment for Text-Guided Image Inpainting
专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL
Comments EACL 2023
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL
Comments EACL 2023
专题命中 多模态训练与对齐 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.CL
专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL
Comments Need to update the results
专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CL、cs.AI
专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CL、eess.AS
专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI
Comments 30 pages, 6 figures, 3 tables
专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI
专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI
Comments 8 pages, IJCAI 2021
专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL
Comments To be published at ICDAR 2021
专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI
Comments 11 pages, 6 figures
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI
Comments Published in ALVR 2020, a workshop in ACL 2020
Journal ref Proceedings of the First Workshop on Advances in Language and Vision Research 2020 (26-31)
专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL
Comments Accepted by CVPR 2020. Code is available at https://github.com/spyflying/CMPC-Refseg
专题命中 多模态训练与对齐 :image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL
Comments ECCV 2020
专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI
Comments 10 pages, 6 figures
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM
Comments 10 pages, 4 figures
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI
Comments EMNLP 2018
专题命中 多模态训练与对齐 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL
多模态大语言模型高效性中的标记压缩综述
机构 * Zhejiang University(浙江大学) ; Westlake University(西湖大学) ; Xiamen University(厦门大学) ; National University of Singapore(新加坡国立大学) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; University of Central Florida(佛罗里达大学) ; Salesforce AI Research(Salesforce AI研究) ; Rice University(德克萨斯大学)
专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
AI总结 本文综述了多模态大语言模型中标记压缩技术,分类讨论了图像、视频和音频三种模态的压缩方法及其机制,旨在推动该领域的发展。
Comments For ongoing updates and to track the latest advances in this promising area, we maintain a public repository: https://github.com/cokeshao/Awesome-Multimodal-Token-Compression
专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments This is the preprint version of the paper to appear in BMVC 2024. Please cite the final published version. Code is available at https://github.com/Mr-Monday/Multi-modal-Crowd-Counting-via-Modal-Emulation
专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI
Comments Multimodal Large Language Models Defense, 25 Pages
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments [TL;DR] we design and release the SNARE, the first large-scale multimodal alignment probing benchmark for current vision-language pretrained models
VGR:视觉基础推理
机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; ByteDance Inc.(字节跳动公司)
专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 本文提出VGR,一种增强视觉感知的多模态大语言模型,通过图像区域检测与回放提升多模态推理能力,在多个基准测试中表现优异。
Comments 9 pages, 4 figures
面向科技情报(STI)的高效多模态多语言观点抽取:一种基于QLoRA的微调方法
机构 * Beihang University(北京航空航天大学) ; Nanchang University(南昌大学) ; Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI
AI总结 本研究针对科技情报观点抽取的噪声过滤与结构化输出问题,提出基于QLoRA微调的多模态框架,在2194样本数据集上实现多语言观点抽取性能显著提升。
重新思考多模态零样本异常检测中的辅助模态:从语义融合到条件调制
专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV
AI总结 本研究针对现有多模态零样本异常检测方法的缺陷,提出即插即用的辅助条件增强框架,通过全局到局部的条件调制实现选择性多模态增强,在MVTec 3D-AD等数据集上提升了现有RGB零样本异常检测器的性能并达到最优。
GALA:面向文本到时间序列合成的生成感知跨模态对齐
专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CL
AI总结 本研究针对文本到时间序列合成中条件表示与信号模态不匹配的问题,提出GALA两阶段跨模态对齐方法,在TSFragment-600K数据集上实现SOTA,打破了生成器内部文本编码器的保真度与贴合度权衡。
Comments 21 pages, 6 figures
何时仅需任务向量?隐式多模态上下文学习的经验理论
机构 * Brown University(布朗大学)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV
AI总结 本文提出选择-实现假说,通过受控多模态任务和VQA基准研究发现,静态任务向量的成功取决于演示诱导变化的跨查询共享程度,为隐式多模态上下文学习的方法选择提供了统一经验理论。
Comments Accepted by Empirical Theory in Representation Learning @ ECCV 2026, Oral
EGM-Det:面向无人机RGB-IR目标检测的熵引导多模态自适应融合
机构 * School of Cybersecurity, Northwestern Polytechnical University(西北工业大学网络空间安全学院) ; School of Automation and Software Engineering, Shanxi University(山西大学自动化与软件学院)
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
AI总结 本文提出EGM-Det框架,通过熵引导多模态自适应融合解决无人机RGB-IR目标检测中模态可靠性被忽略的问题,在三个数据集上实现最优性能,VEDAI上较现有方法提升超10个百分点。
Comments 14 pages, 7 figures, 6 tables
TongGuOCR:面向中文历史文献的布局感知与 token 增强 OCR 框架
机构 * School of Electronic and Information Engineering, South China University of Technology(华南理工大学电子与信息工程学院) ; Huawei Technologies Co., Ltd.(华为技术有限公司)
专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.AI
AI总结 针对中文历史文献OCR的复杂布局、生僻字等挑战,提出TongGuOCR框架,通过布局感知预处理与token增强识别模块,在M5HisDoc等基准上取得优于同类模型的性能。
通过交叉注意力和门控融合进行多模态矛盾与犹豫识别
机构 * University of Paris 8(巴黎第八大学) ; LIASD Laboratory(LIASD实验室)
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
AI总结 为ECCV 2026的ABAW11挑战赛开发多模态框架,用预训练编码器提取多模态特征,建立单模态基线,在此基础上提出含交叉注意力和门控融合的多模态架构,验证集宏F1达0.7394,提升显著。
DAP-Pose:用于鲁棒姿态估计的深度时间对齐和物理感知跨模态传感器融合
机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院空天信息创新研究院) ; International Research Center of Big Data for Sustainable Development Goals(可持续发展大数据国际研究中心) ; University of Chinese Academy of Sciences(中国科学院大学) ; The University of Adelaide(阿德莱德大学)
专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV
AI总结 针对复杂环境下多模态传感器的姿态估计问题,提出DAP-Pose模型,通过双级跨模态融合模块捕捉线索,深度时间对齐模块处理异步流,结合物理感知约束,在KITTI数据集上达最优性能,平均平移误差1.31%,旋转误差0.46°,严重错位下也能准确估计。