UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI
Comments Accpeted to CVPR 2025 workshop
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI
Comments Accpeted to CVPR 2025 workshop
专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI
专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI
专题命中 多模态生成 :cross-modal(title);分类 cs.CV、cs.AI
Comments The paper has been accepted by Medical Image Computing and Computer Assisted Intervention Society (MICCAI) 2024
专题命中 多模态生成 :multi-modal(title);分类 cs.CV、cs.CL
专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI
Comments Currently submitted to: Scientific Reports
专题命中 多模态生成 :audio-visual(title);分类 cs.MM、eess.AS
Comments 8 pages, 5 figures, to be published in the proceedings of the 27th International Conference on Digital Audio Effects (DAFx24), for additional image and video examples see https://dzluke.github.io/DAFX2024/
专题命中 多模态生成 :multimodal(title);分类 cs.CL、cs.AI
Comments Accepted to Chemical Science Journal. Models are publicly available via https://huggingface.co/insilicomedicine/nach0_base and https://huggingface.co/insilicomedicine/nach0_large
Journal ref Chemical Science, 15(22), 8380-8389, 2024
专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI
专题命中 多模态生成 :cross-modal(title);分类 cs.CL、cs.AI
Comments Accepted in EACL 2024 as a long paper. See accepted/#long-papers" target="_blank" rel="noopener">https://2024.eacl.org/program/findings-accepted/#long-papers . Note: this paper's ArXiv version includes additional discussion, analysis, and types of experiments compared to the EACL version. Changes introduced in V2 of ArXiv paper: only this comment metadata. V1 was initially submission on July 26th, 2023 - release was delayed by ArXiv for a few days
专题命中 多模态生成 :multi-modal(title);分类 cs.CV、cs.CL
专题命中 多模态生成 :cross-modal(title);分类 cs.CV、eess.AS
专题命中 多模态生成 :cross-modal(title);分类 cs.CV、cs.CL
Comments Accepted as a long paper to ACL 2020
专题命中 多模态生成 :cross-modal(title);分类 cs.CV、cs.AI
专题命中 多模态生成 :multi-modal(title);分类 cs.CV、cs.AI
专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL
Comments This is an undergraduate project report. Completed Dec. 2019 at the Cooper Union
专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI
Comments 21 pages, 15 figures, 34th Conference on Neural Information Processing Systems (NeurIPS 2020)
专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL
专题命中 多模态生成 :multimodal(title);分类 cs.CL、cs.AI
Comments 9 pages (incl refs), 7 figures, 3 tables, proceedings of LREC 2020 (postponed due to COVID-19)
专题命中 多模态生成 :multimodal(title);分类 cs.CL、cs.AI
Comments Published as a conference paper at ICLR 2019
专题命中 多模态生成 :multimodal(title,comments);分类 cs.CV
Comments 5th International Workshop on Multiscale Multimodal Medical Imaging (MICCAI 2024), Project page: https://sven-luepke.github.io/phy-ldm-mri/
CustomDance:基于以人为中心的粗到细交互式控制的定制化3D舞蹈生成
专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract)
AI总结 CustomDance是基于粗到细交互式控制的定制化3D舞蹈生成系统,通过三阶段流程实现AI辅助编舞,在定量和定性比较中优于基线,为用户提供全面控制。
AVTok: 用于整体音视频生成的1D统一分词化
机构 * The Hong Kong University of Science and Technology(香港科技大学)
专题命中 多模态生成 :multimodal(abstract);audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS
AI总结 提出AVTok统一分词器,采用双流Transformer架构和分层训练策略,将音视频对编码为紧凑的一维潜在表示,在音视频重建及下游生成任务中表现优异。
Comments ECCV 2026
UniversalRAG: 在多样模态和粒度的语料库上实现检索增强生成
机构 * KAIST(韩国科学技术院)
专题命中 多模态生成 :any-to-any(abstract,abstract_cn);分类 cs.CV、cs.CL、cs.AI
AI总结 本文提出UniversalRAG,一种能够处理多种模态和粒度的检索增强生成框架,通过动态路由机制和多粒度组织,提升跨模态知识检索的有效性,实验表明其在多个模态基准上的优越性。
Comments ACL 2026. Project page : https://universalrag.github.io
联合音视频生成的推理时缩放
机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院) ; Luma AI
专题命中 多模态生成 :multimodal(abstract);audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS
AI总结 针对联合音视频生成中多目标优化的挑战,提出多验证器框架与自适应奖励加权算法,在无需额外训练的情况下显著提升语义对齐、感知质量和音视频同步。
Comments Accepted by Transactions on Machine Learning Research (TMLR). Project page: https://jung-jaemin.github.io/ITS-AVGen-Proj/
ETCHR: 通过编辑来澄清和利用推理
机构 * The Chinese University of Hong Kong Shanghai AI Laboratory(香港中文大学上海人工智能实验室) ; Shanghai AI Laboratory(上海人工智能实验室) ; Shanghai Jiao Tong University(上海交通大学) ; Shanghai Innovation Institute(上海创新研究院)
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract_cn);分类 cs.CV、cs.CL、cs.AI
AI总结 针对多模态大语言模型在需要细粒度关注或视角变换的问题上纯文本思维链的瓶颈,提出了一种解耦的图像编辑模型ETCHR,通过两阶段训练(推理模仿和推理增强)弥合语言侧和生成侧差距,无需训练即可插入不同MLLM,在五个任务族上平均Pass@1提升4.61-5.47个百分点。
Comments Code, model and data are open-sourced at https://github.com/InternLM/ETCHR
OSCBench:文本到视频生成中对象状态变化的基准测试
机构 * National University of Singapore(新加坡国立大学) ; Singapore Management University(新加坡管理大学) ; Carnegie Mellon University(卡内基梅隆大学) ; Fudan University(复旦大学)
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 本文提出OSCBench基准,用于评估文本到视频模型在对象状态变化上的性能,揭示当前模型在处理新场景时的不足。
Comments ACL 2026 Main Conference, Project page: https://hanxjing.github.io/OSCBench
微小推理时间缩放与潜在验证器
机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) ; University of Pisa(比萨大学)
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.MM
AI总结 本文提出VHS验证器,直接在扩散变换器单步生成器的中间隐藏表示上操作,减少验证成本并提升性能,实现更高效的推理时间缩放。
Comments Findings of CVPR 2026 - Code at: https://aimagelab.github.io/VHS/
一个尺寸,多种适配:在大规模广告图像生成中对多样化群体点击偏好进行对齐
机构 * NLPR & MAIS, CASIA(中国科学院长春光学精密机械与物理研究所 & 中国科学院自动化所) ; School of AI, UCAS(中国科学院大学人工智能学院) ; HKUST(gz)(香港科技大学) ; PRLab, NJU(南京大学PRLab)
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.MM
AI总结 本文提出OSMF框架,通过自适应分组和群组感知多模态模型,解决广告图像生成中用户群体点击偏好多样性的优化问题。
TalkVerse:民主化分钟级音频驱动视频生成
机构 * The Chinese University of Hong Kong(香港中文大学) ; Snap Inc.(Snap公司)
专题命中 多模态生成 :MLLM(abstract);audio-visual(abstract);分类 cs.CV、cs.AI、cs.MM
AI总结 TalkVerse通过大规模开放数据和高效模型,实现了分钟级音频驱动视频生成,降低研究门槛,提升生成质量与效率。
Comments open-sourced single-person full-body talking video generation dataset, training code and checkpoints