Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification
像素与词语之间的桥梁:面向多模态媒体验证的掩码感知局部语义融合
机构 * State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学人工智能与机器人研究所人机混合增强智能全国重点实验室)
AI总结 本文提出MaLSF框架,通过主动双向验证和层次语义聚合模块,解决多模态虚假信息检测中的局部语义不一致问题,实现像素与词语的语义连接。
Comments Accepted by CVPR 2026