GeRA: Label-Efficient Geometrically Regularized Alignment
专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL
Comments 9 pages
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL
Comments 9 pages
专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI
专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI
Comments 18 pages,
专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM
Comments Technical Report
专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI
Comments 9 pages, 12 figures
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL
Comments Accepted at ICLR 2023
专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL
Comments Fixed typos
专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL
Comments v2: (i) fix / update EVA IN-1K variants results. (ii) add / update EVA-CLIP results. (iii) add Appendix. (iv) release all the code and models at https://github.com/baaivision/EVA
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL
Comments Accepted by AAAI 2023. Code is available at https://github.com/MAEHCM/AET
专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
Comments 17 pages, 7 figures
专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI
Comments incomplete experiments
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL
Comments ACM Multimedia 2022
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
Comments This work has been accepted by IEEE Transactions on Neural Networks and Learning Systems
专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL
Comments CVPR 2022
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、eess.AS
Comments 10 pages, 1 figures, added references and an overview figure
专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL
Comments To appear in CVPR'2021
专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL
Comments AAAI 2021; Code is publicly available at: https://github.com/YehLi/TDEN
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL
Comments Accepted by EMNLP 2020
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL
Comments This paper is going to appear in TPAMI. Code is available at https://github.com/sibeiyang/sgmn/tree/master/lib/cmrin_models
无配对RGB-热成像高斯泼溅使用视觉几何变换器
机构 * Ecole Polytechnique Federale de Lausanne(瑞士联邦理工学院洛桑分校) ; Schindler EPFL Lab(施耐德EPFL实验室)
专题命中 多模态训练与对齐 :multi-modal(abstract,comments);cross-modal(abstract);分类 cs.CV
AI总结 提出一种无配对RGB-热成像新视角合成框架,利用VGGT估计各模态相机位姿并通过Procrustes对齐,结合多模态3D高斯泼溅实现联合重建,在保持RGB保真度的同时实现热成像视图合成。
Comments Accepted at ICRA 2026's Workshop MM-SpatialAI: Multi-Modal Spatial AI for Robust Navigation and Open-World Understanding
基于大语言模型的视觉编码器分层预训练
机构 * University of Cincinnati(辛辛那提大学) ; National Yang Ming Chiao Tung University(国立阳明交通大学)
专题命中 多模态训练与对齐 :multimodal(abstract,comments);分类 cs.CV、cs.CL、cs.AI;multimodal foundation model(comments)
AI总结 本文提出HIVE框架,通过引入视觉编码器与大语言模型间的分层交叉注意力机制,提升视觉语言对齐,改进特征融合与表征学习,实验表明其在图像分类和多模态任务中表现优异。
Comments 17 pages, 14 figures, accepted to Computer Vision and Pattern Recognition Conference (CVPR) Workshops 2026. 5th MMFM Workshop: What is Next in Multimodal Foundation Models?
Journal ref In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7415-7424) 2026
健康基础模型中的涌现符号结构:提取、对齐与跨模态迁移
机构 * Apple(苹果公司)
专题命中 多模态训练与对齐 :cross-modal(title)
AI总结 本文提出一种训练后框架,通过分解冻结嵌入以提取可解释的符号,用于对齐嵌入空间。在PPG和加速度计数据上验证,发现符号能选择性关联健康状况和生理属性,并支持跨模态迁移。
Comments 8 pages, Mechanistic Interpretability Workshop at the 43rd International Conference on Machine Learning, 4 main figures
Journal ref Mechanistic Interpretability Workshop at the 43 rd International Conference on Machine Learning, Seoul, South Korea, 2026
ChemFusion:用于反应产率预测的多模态交叉注意力网络
机构 * Duke University(杜克大学) ; Georgia Institute of Technology(佐治亚理工学院)
专题命中 多模态训练与对齐 :multimodal(title)
AI总结 研究过渡金属催化反应产率预测难题,提出ChemFusion多模态交叉注意力网络,融合电子特征与3D原子坐标,用交叉注意力机制,在交叉偶联库基准测试中性能出色,还能自主学习识别和惩罚空间位阻,提供物理可解释性。
Comments 10 pages, 4 figures, 2 tables
探索射电星系形态的图像-文本对齐
专题命中 多模态训练与对齐 :image-text(title)
AI总结 研究射电星系图像与文本描述的对齐,利用MiraBest数据集及SigLIP-2模型,经特定提示生成描述并评估。结果显示基于描述的星系分类与图像类似,微调改善局部连贯性,此为射电星系形态研究提供新方法。
Comments Accepted at the AI in Science (AIS) Conference 2025
用于连接受限环境中自主物联网节点的统一多模态传感与主动稳定框架:从周边哨兵到高价值货物保护
专题命中 多模态训练与对齐 :multi-modal(title)
AI总结 探讨连接受限环境下自主物联网节点需求,提出广义节点架构,扩展其传感与通信模态,通过复合风险指数耦合遥测与地理信息系统,统一两种部署模式并给出分析与实例。
SurgAM:用于机器人自主的多模态特征融合手术能力地图预测
机构 * Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程系) ; Department of Thoracic Surgery, Peking University People’s Hospital(北京大学人民医院胸外科) ; Thoracic Oncology Institute, Peking University People’s Hospital(北京大学人民医院胸部肿瘤研究所) ; Research Unit of Intelligence Diagnosis and Treatment in Early Non-small Cell Lung Cancer, Chinese Academy of Medical Sciences(中国医学科学院早期非小细胞肺癌智能诊断与治疗研究组) ; Institute of Advanced Clinical Medicine, Peking University(北京大学先进临床医学院) ; Beijing Key Laboratory of Innovative Application of Big Data in Lung Cancer, Peking University People’s Hospital(北京大学人民医院肺癌大数据创新应用北京市重点实验室)
专题命中 多模态训练与对齐 :multimodal(title)
AI总结 研究如何通过视觉数据识别手术可操作区域,提出自适应特征融合框架、分层提示学习机制和场景引导注意力解码器,建立新数据集验证,在真实模型上验证框架对下游自动化的适用性。
Comments ICRA 2026
传感器堆栈对非接触式床上身体位置的限制:20名受试者多模态雷达+热成像LOSO表征
专题命中 多模态训练与对齐 :multimodal(title)
AI总结 通过20名受试者的多模态雷达与热成像数据,研究非接触式床上体位推断的传感器表示限制,发现融合雷达与热成像的Logistic回归在床内/外分类中达到0.871中位平衡准确率,但俯卧检测性能不足(召回率0.50,精确率0.41),指出原始距离-FFT访问是下一步硬件实验方向。
Comments 11 pages, 3 figures, 6 tables
DriVerse:通过多模态轨迹提示和运动对齐实现驾驶模拟的导航世界模型
机构 * Baidu Inc.(百度公司)
专题命中 多模态训练与对齐 :multimodal(title)
AI总结 DriVerse通过多模态轨迹提示和运动对齐技术,实现从单张图像和未来轨迹生成导航驱动的驾驶场景,提升了动态对象的生成精度和时间一致性。
Comments 13 pages, 5 figures