Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS
Comments Demo and implementation at https://auffusion.github.io
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS
Comments Demo and implementation at https://auffusion.github.io
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS
Comments ACL 2023. arXiv admin note: substantial text overlap with arXiv:2305.07760
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Journal ref Published at NeurIPS 2023
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to NeurIPS 2023
专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI、cs.MM、eess.AS
Comments Accepted by ICML 2023. Demo and implementation at https://audioldm.github.io. Evaluation toolbox at https://github.com/haoheliu/audioldm_eval
专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI、cs.MM
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted by ACM MM'23
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to ICCV 2023, Project Page: https://promptstyler.github.io/
专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI、cs.MM、eess.AS
Comments 16 pages, 3 figures, 2 tables, demo page: https://musicldm.github.io/
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Equal contribution: Bingshuai Liu and Longyue Wang. Work done while Bingshuai Liu and Chengyang Lyu were interning at Tencent AI Lab. Zhaopeng Tu is the corresponding author
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM
Comments Presented at AAAI23 CreativeAI workshop (Non-Archival). A later version is accepted to ACL23
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to Robotics: Science and Systems (RSS) 2023. The previous version appeared in CoRL Workshop on Language and Robot Learning 2022
专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI、cs.MM
Comments This paper has 29 pages with 22 figures, including rich supplementary information. Project page is at \url{https://classifier-as-generator.github.io/}
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments This paper is current under peer-review in IEEE TNNLS
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments To appear at Findings of ACL 2022
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.MM
Comments ACM MM 2021 (Video and Demo Track). Code: https://github.com/researchmm/generate-it
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM
Comments The first two authors contributed to this work equally
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted to TMM, an extended version of a paper published in ACM MM 2019. arXiv admin note: substantial text overlap with arXiv:1908.00999
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments ICML 2021 (15 pages, 4 figures, 14 tables)
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted for MICCAI 2019
专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract)
Comments 7 pages, 4 figures
Journal ref 19th IEEE International Conference on Bioinformatics and Bioengineering (IEEE BIBE, 2019)
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.MM
专题命中 多模态生成 :multi-modal(abstract,journal_ref);分类 cs.CV、cs.CL
Journal ref CVPR'2022 Workshop on Open-Domain Retrieval Under a Multi-Modal Setting
从语料到协同进化的能力:面向通用图像生成的以能力为中心的数据设计
机构 * Alibaba Group(阿里巴巴集团)
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI
AI总结 该研究提出能力驱动的数据基础设施,构建三类监督引擎,整理大规模图像语料,训练多模态扩散模型,在CPI-Bench等评估中展现良好生成与迁移能力。
Comments 19 pages, 10 figures
下一步编辑什么:对话系统中视觉对齐的图像编辑后续建议
机构 * Qwen Business Unit of Alibaba(阿里巴巴通义千问业务单元) ; Southeast University(东南大学)
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI
AI总结 该研究针对对话系统的图像创作场景,提出三阶段多模态推荐框架,经测试可降低视觉不一致率并显著提升推荐相关指标。
TRACE-Bench:多参考图像生成的分解与诊断
机构 * Shanghai Jiao Tong University(上海交通大学)
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI
AI总结 研究针对多参考图像生成基准的缺陷,构建TRACE-Bench基准,采用算子分解方法评估9个主流模型,发现解耦与属性绑定是核心瓶颈,最优模型属性保真度仅0.74
Comments Accepted to ACM Multimedia 2026 (ACM MM 2026)
Omni-LiveAvatar:分钟级实时流式视听联动数字人生成
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.MM
AI总结 本研究提出Omni-LiveAvatar框架,通过渐进式自回归蒸馏等技术,实现分钟级实时流式视听联动数字人生成,速度较LTX-2提升33倍且性能优于基线模型。
G-MAD:用于多视图RGB-T航空目标检测的基于游戏的数据生成框架
机构 * LIG Defense&Aerospace(LIG国防与航空航天公司) ; GIST(韩国科学技术院)
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI
AI总结 研究针对航空目标检测中数据集构建的问题,提出基于游戏的G-MAD数据生成框架,可解决视点控制、数据对齐及标注成本等问题,支持多视图相机放置等功能,还构建并发布了AMOD基准。
Comments ACM Multimedia 2026 (Supplementary Material Included)