MagicQuill: An Intelligent Interactive Image Editing System
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted to CVPR 2025. Code and demo available at https://magic-quill.github.io
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted to CVPR 2025. Code and demo available at https://magic-quill.github.io
专题命中 多模态生成 :cross-modal(abstract);image-text(abstract);分类 cs.CV
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments Project Page: https://opendatalab.github.io/LEGION
专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
Comments CVPR 2025
专题命中 多模态生成 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV
Comments CVPR2025 Accepted Paper
专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract);分类 cs.CV
Comments CVPR 2025. Project website: https://github.com/sssaury/HAM
专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract);分类 cs.CV
Comments Code is available at https://github.com/zhengchen1999/PromptSR
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments ICLR 2025. Project page: https://mc-lan.github.io/Text4Seg/
专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
Comments 11 pages, 17 figures
专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.AI
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments Project page: https://t2v-compbench-2025.github.io/ Code: https://github.com/KaiyueSun98/T2V-CompBench/tree/V2
专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
Comments This paper has been accepted by AAAI 2025
专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted by the WACV 2025, including supplementary material
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted by ACM MM 2024, code will be released in https://github.com/XiaominLi1997/MoTrans
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV
专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments 13 pages, 11 figures, WACV2025
专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV
Comments NeurIPS 2024
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments accepted by NeurIPS 2024
专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV
Comments Add more results using task tokens, expand the introduction and related work FIX: error in LLM-as-judge evaluation that was over-inflating the results
专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV
Comments Accepted by ACM MM 2024
专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV
Comments ECCV 2024; Project page at https://idea2img.github.io/
专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
Comments ECCV24
专题命中 多模态生成 :image-text(abstract);any-to-any(abstract);分类 cs.CV
Comments 14 pages, 9 figures
专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract);分类 cs.CV
Comments CVPR 2024
专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments We open-source our data, model, and code at: https://github.com/linzhiqiu/t2v_metrics ; Project page: https://linzhiqiu.github.io/papers/vqascore
专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CL
Comments Taiyi-Diffusion-XL Tech Report
专题命中 多模态生成 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV
Comments Codes and models: \url{https://github.com/FoundationVision/LlamaGen}