Self-Purification Mitigates Backdoors in Multimodal Diffusion Language Models
自净化缓解多模态扩散语言模型中的后门
机构 * National University of Singapore(新加坡国立大学)
专题命中 多模态生成 :multimodal(title,abstract)
AI总结 DiSP通过在推理过程中选择性屏蔽视觉标记,有效去除多模态扩散语言模型中的后门,无需额外模型或干净数据。
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
自净化缓解多模态扩散语言模型中的后门
机构 * National University of Singapore(新加坡国立大学)
专题命中 多模态生成 :multimodal(title,abstract)
AI总结 DiSP通过在推理过程中选择性屏蔽视觉标记,有效去除多模态扩散语言模型中的后门,无需额外模型或干净数据。
超越像素模拟:通过诊断语义标记和原型控制生成病理图像
机构 * College of Intelligent Robotics and Advanced Manufacturing(智能机器人与先进制造学院) ; Fudan University(复旦大学) ; Fysics Intelligence Technologies Co., Ltd.(Fysics智能科技有限公司) ; University of Science and Technology Beijing(北京科技大学) ; ByteDance(字节跳动) ; School of Computer Science and Technology(计算机科学与技术学院) ; Xi’an Jiaotong University(西安交通大学)
专题命中 多模态生成 :MLLM(abstract);image-text(abstract);分类 cs.CV
AI总结 UniPath通过诊断语义标记和原型控制实现可控的病理图像生成,取得SOTA性能,包括Patho-FID 80.9和98.7%的细粒度语义控制。
Comments accepted by CVPR 2026; 32 pages, 17 figures, and 6 tables
Bob的彩纸:音乐和视频生成中的语音记忆攻击
机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) ; University of California San Diego(加州大学圣地亚哥分校) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS
AI总结 研究揭示了生成音乐和视频AI系统中语音记忆攻击的漏洞,通过同音替代词绕过版权过滤,展示语音结构对跨模态检索的关键作用。
基于指令的图像编辑与规划、推理和生成
机构 * HKUST(香港科技大学)
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI
AI总结 本文提出一种多模态模型,通过链式思考规划、编辑区域推理和编辑,提升基于指令的图像编辑能力,以应对更复杂的真实场景。
Comments 10 pages, 7 figures
Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, Page 17506--17515
MammoWise:多模型本地RAG流水线用于乳腺X线摄影报告生成
机构 * University of California, Davis(加州大学戴维斯分校)
专题命中 多模态生成 :multimodal(abstract,comments);分类 cs.CV
AI总结 MammoWise是一种本地多模型流水线,通过多任务分类和检索增强生成技术,实现乳腺X线摄影报告的高准确度生成与分类。
Comments arXiv preprint (submitted 25 Feb 2026). Local multi-model pipeline for mammography report generation + classification using prompting, multimodal RAG (ChromaDB), and QLoRA fine-tuning; evaluates MedGemma, LLaVA-Med, Qwen2.5-VL on VinDr-Mammo and DMID; reports BERTScore/ROUGE-L and classification metrics
ThinkRL-Edit: 基于强化学习的图像编辑推理框架
机构 * Zhejiang University(浙江大学)
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV
AI总结 ThinkRL-Edit通过引入链式推理和无偏奖励策略,提升图像编辑的推理能力,实现更精准和稳定的编辑效果。